TY - GEN
T1 - TiNA
T2 - 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, ASPLOS 2026
AU - Agarwal, Siddharth
AU - Wang, Tianchen
AU - Huang, Jinghan
AU - Agarwal, Saksham
AU - Kim, Nam Sung
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2025/12/11
Y1 - 2025/12/11
N2 - To manufacture a large CPU cost-effectively, the industry has begun exploiting emerging packaging technologies that integrate multiple chiplets—each comprising a subset of cores and/or memory and I/O subsystems—into a single package. However, such a CPU experiences longer memory access latency with more pronounced variance, especially when its cores in one chiplet access LLC slices1 or DRAM controllers in other chiplets. This creates unique challenges in µs-scale networking, which is highly sensitive to memory access latency. In this work, we start by proposing exploiting a little-known mode, known as Sub-NUMA Clustering (SNC), in the latest chiplet-based CPUs. As it restricts receiving and processing packets to a particular chiplet unless explicitly specified otherwise, it offers shorter memory access latency and, consequently, lower networking latency than the default mode (non-SNC). Nonetheless, when receiving long bursts of packets2, SNC incurs higher networking latency than non-SNC, as it provides less LLC capacity for CPU cores processing the packets, making Direct Cache Access (DCA)—a commonly used CPU feature to reduce memory access latency for packet processing—ineffective. To address this drawback, we propose TiNA, a tiered network buffer architecture consisting of an enhanced NIC and networking stack, which opportunistically uses LLC slices in other chiplets for DCA only when receiving long bursts of packets. On average, TiNA reduces the mean (tail) latency by 25% (18%) and 28% (22%), compared to SNC and non-SNC, respectively, across diverse network applications and traces.
AB - To manufacture a large CPU cost-effectively, the industry has begun exploiting emerging packaging technologies that integrate multiple chiplets—each comprising a subset of cores and/or memory and I/O subsystems—into a single package. However, such a CPU experiences longer memory access latency with more pronounced variance, especially when its cores in one chiplet access LLC slices1 or DRAM controllers in other chiplets. This creates unique challenges in µs-scale networking, which is highly sensitive to memory access latency. In this work, we start by proposing exploiting a little-known mode, known as Sub-NUMA Clustering (SNC), in the latest chiplet-based CPUs. As it restricts receiving and processing packets to a particular chiplet unless explicitly specified otherwise, it offers shorter memory access latency and, consequently, lower networking latency than the default mode (non-SNC). Nonetheless, when receiving long bursts of packets2, SNC incurs higher networking latency than non-SNC, as it provides less LLC capacity for CPU cores processing the packets, making Direct Cache Access (DCA)—a commonly used CPU feature to reduce memory access latency for packet processing—ineffective. To address this drawback, we propose TiNA, a tiered network buffer architecture consisting of an enhanced NIC and networking stack, which opportunistically uses LLC slices in other chiplets for DCA only when receiving long bursts of packets. On average, TiNA reduces the mean (tail) latency by 25% (18%) and 28% (22%), compared to SNC and non-SNC, respectively, across diverse network applications and traces.
KW - Chiplet
KW - Direct Cache Access
KW - Direct Memory Access
KW - Non Uniform Cache Access
KW - Sub-NUMA Clustering
UR - https://www.scopus.com/pages/publications/105036327367
UR - https://www.scopus.com/pages/publications/105036327367#tab=citedBy
U2 - 10.1145/3760250.3762224
DO - 10.1145/3760250.3762224
M3 - Conference contribution
AN - SCOPUS:105036327367
T3 - International Conference on Architectural Support for Programming Languages and Operating Systems - ASPLOS
SP - 298
EP - 313
BT - ASPLOS 2026 - Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1
PB - Association for Computing Machinery
Y2 - 22 March 2026 through 26 March 2026
ER -