Cache Coherency Protocol Overhead in Non-Uniform Memory Access Architectures: An Empirical Analysis of MESIF-to-Directory Transition Latencies Under Heterogeneous Workload Contention
Keywords:
NUMA architecture, cache coherency protocols, MESIF state machine, directory-based coherence, inter-socket latency, hardware performance counters, false sharing, HPC workload contention, NUMA-aware schedulingAbstract
Non-uniform memory access (NUMA) architectures increasingly underpin high-performance computing clusters, yet cache coherency protocol transitions—particularly MESIF-to-directory escalations—introduce measurable latency penalties under heterogeneous workload contention. This study presents a controlled empirical analysis of directory-based coherency overhead across a 128-core AMD EPYC 9654 dual-socket testbed, instrumenting hardware performance counters to isolate inter-node snoop traffic from intra-node false-sharing artifacts. Using a purpose-built micro-benchmark suite alongside production-grade HPC workloads (LINPACK, STREAM, NAS Parallel Benchmarks), we quantify transition latency distributions, cache-line invalidation rates, and directory saturation thresholds. Results demonstrate that MESIF Modified-to-Shared transitions account for up to 34.7% of total inter-socket communication latency under peak contention. Proposed heuristic-driven prefetch dampening reduces coherency stall cycles by 19.3% without degrading throughput. These findings carry direct implications for OS-level NUMA-aware scheduling and firmware-level coherency tuning.
References
Semeniuk, V. V. (2025). OPTIMIZATION OF LOCAL DEVELOPMENT PROCESS USING DOCKER PHP IMAGE THAT COMES WITH A FULL SET OF TOOLS OUT OF THE BOX–PERFORMANCE AND OPTIMIZATION EXTENSIONS. ІНФОРМАЦІЙНЕ ЗАБЕЗПЕЧЕННЯ БАГАТОІНДЕКСНОЇ ТРАНСПОРТНОЇ ЗАДАЧІ З НЕЧІТКИМИ ІНТЕРВАЛАМИ.