
核心结论通过 NVLink 进行 GPU-to-GPU peer memory 访问时本地 GPU 的 L2 Cache 不缓存从远端 GPU 读取的数据。这与 LTC Fabricdie 内跨分区访问的行为形成鲜明对比访问类型本地 L2 是否缓存数据缓存位置LTC Fabricdie 内跨 L2 分区✅是— 在本地 L2 复制 [23]本地 L2Shared 状态NVLinkGPU 间 peer memory❌否— 直接 bypass 本地 L2 [181][191]仅远端 GPU 的 L2证据来源1. Stanford Hazy Research Blog最明确[181][182] Stanford Hazy Research 关于多 GPU kernel 设计的系列文章明确写道“Remote L2 caching behavior: Because remote HBM access bypasses the local L2 cache, data is cached only on the remote peer’s L2. This design simplifies inter-GPU memory consistency but makes every remote access bottlenecked by NVLink bandwidth.”[182] 进一步解释了数据路径“The data path always flows through the source device’s L2 cache, then its crossbar, across NVLink, through the destination device’s crossbar, and finally to the SMs.This means peer data is never cached in the local L2, and peer memory access is always bottlenecked by NVLink bandwidth.”数据路径图远程 GPU (Source) 本地 GPU (Destination) ┌─────────────────┐ ┌─────────────────┐ │ HBM │ │ │ │ ↑ │ │ │ │ L2 Cache │ │ ❌ L2 Cache │ │ (命中后缓存) │ │ (不缓存peer │ │ ↓ │ │ 数据) │ │ Crossbar │ │ Crossbar │ │ ↓ │ │ ↓ │ │ NVLink 接口 │ ─── NVLink ───→ │ NVLink 接口 │ └─────────────────┘ │ ↓ │ │ SM (L1) │ └─────────────────┘2. “Spy in the GPU-box” 论文反向工程实证[191]Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU Systems通过微基准测试直接测量了缓存行为“Our experiment shows that this data accessed on the remote GPU is cached on the remote GPU,rather on the local L2 GPU. Of course,caching the data locally, would introduce cache coherence issuessince copies of the same data could exist in multiple L2 caches.”“In summary, our reverse engineering results demonstrate thatan access to the memory of a remote GPU through NVLink is cached on the L2 cache of remote GPU, but not L2 cache of local GPU.”他们测量了四种访问延迟本地 L2 命中 (~250 cycles)本地 HBM 访问远程 L2 命中数据在远程 GPU 的 L2 中远程 HBM 访问这说明访问远程 GPU 的数据时如果该数据在远程 GPU 的 L2 中你仍然可以命中远程 L2延迟比远程 HBM 低但绝不会在本地 L2 中留下副本。3. NVIDIA Developer Forums用户实测[192] NVIDIA 官方论坛上的讨论用户cudaMancpy问“Is there a way to see if it’s being cached? (I mean, I also want to check the hit ratio for the data that was accessed from peer.)”经过实验后用户自己得出结论“It seemsno caching totally. (Remote GPU memory’s total data amount is 256MB. Received User Bytes value is for about 268MB.)”“Peer Memory is enabled. In this state, when the local GPU reads the data from the remote GPU memory, I thought caching occurred by default with the local GPU L2 cache. However, it does not appear to be the result of the experiment.It seems to be a one-time read only data without caching with direct access.”NVIDIA 员工Robert_Crovella也确认了测试方法如果数据被缓存反复访问性能应该提升如果每次都走 NVLink性能维持在 NVLink 带宽水平。4. “Processing Large Data on GPUs with Fast Interconnects” 论文[193] 这篇 2020 年的论文明确写道“We observe that, in contrast to GPU memory, the small hash table isnot cached in the GPU’s L2 cache for NVLink 2.0.The L2 cache is memory-side, and cannot cache remote data.”这是关键洞察L2 Cache 是memory-side的——它只缓存属于本 GPU HBM 地址范围的数据。远程 GPU 的 HBM 数据不属于本地 GPU 的 memory space所以本地 L2 不缓存。5. Multi-GPU Programming Blog[206] 经典的多 GPU 编程教程也证实了这一点“The difference in L2 cache is quite interesting. The data is actually cached in L2.The only difference is that it’s L2 of a different GPU.”为什么 NVIDIA 这样设计如果本地 L2 缓存远端数据结果需要跨 GPU 缓存一致性协议复杂度高延迟大一个 cache line 可能在多个 GPU L2 中存在需要 Directory Invalidate 广播写操作需要失效所有副本严重影响性能GPU 间一致性粒度需精确到 cache line当前设计远端数据不缓存在本地 L2结果简化的弱一致性模型易于实现硬件成本低每次远程访问都走 NVLink可预测但性能受限软件层面通过同步原语如nvshmem_quiet保证一致性程序员显式控制GPU 间一致性粒度显式同步点而非隐式 cache line[204] 也提到“GPU caches not globally coherent across GPUs, only the CPU–GPU NVLink-C2C path is cache coherent.”这说明GPU 之间的缓存一致性不是硬件自动维护的至少不是 cache-line granularity 的而是通过软件层NVSHMEM/NCCL的显式同步来保证。一个细节什么是 “memory-side” L2[193] 提出的 “memory-side” 概念非常重要地址空间归属模型 GPU-A 的 L2 Cache 只缓存 GPU-A HBM 地址空间 的数据 GPU-B 的 L2 Cache 只缓存 GPU-B HBM 地址空间 的数据 即使 GPU-A 通过 NVLink 读取 GPU-B 的数据 该数据的地址仍属于 GPU-B 的地址空间 所以 GPU-A 的 L2 不会缓存它。 但 GPU-B 的 L2 会缓存它因为这是它的 memory space这与论文 [23] 中发现的 LTC Fabric 行为截然不同LTC Fabricdie 内通过地址哈希同一虚拟地址可能映射到不同物理 L2 分区。读取远程分区的数据时在本地 L2 建立一个 Shared 副本因为两个分区共享同一个虚拟地址空间都是本地的。NVLinkGPU 间两个 GPU 有各自独立的 HBM 地址空间即使通过 UVA/P2P 映射底层仍区分归属。本地 L2 严格按地址归属缓存不缓存属于其他 GPU的数据。结论NVLink 本地 GPU 不缓存远端数据这个说法是正确的。这是 NVIDIA 在 GPU-to-GPU NVLink 上的有意设计选择目的是简化多 GPU 缓存一致性。代价是每次远程访问都必须经过 NVLink无法通过本地 L2 的后续命中来加速。这与 die 内部的 LTC Fabric 行为跨分区访问在本地 L2 复制数据形成鲜明对比也解释了为什么跨 GPU 访问比 die 内跨分区访问慢得多。参考资料Stanford Hazy Research Blog [181][182]、Spy in the GPU-box 论文 [191]、NVIDIA Developer Forums [192]、Fast Interconnects 论文 [193]、Multi-GPU Programming Blog [206]、AI Infrastructure 文档 [204]。原文