在大模型推理工程实践中,很多人把注意力放在算力(TFLOPS)上,却忽略了另一个同等关键、甚至往往是真正瓶颈的资源——显存带宽。Llama 3 70B 这类模型在 Decode 阶段每生成一个 token 就要读取全部权重,70B 参数 × 2 字节 = 140 GB 的数据需要在几十毫秒内从 HBM 流过。如果 H100 的 3.35 TB/s 带宽被打满,理论极限也就约 24 token/s;实际还要扣除 KV Cache 读取、Activation 中间张量以及框架开销,能跑到 12-15 token/s 已经相当不错。理解显存占用构成、诊断带宽瓶颈、用 Roofline 模型定位优化方向,是 LLM 推理工程师的必备技能。本文从一个具体的显存测算公式出发,逐步拆解 VRAM 占用、讲解 HBM 带宽测量方法、演示如何用 nsys + Nsight Compute 定位访存密集型 kernel,最后给出一份带宽优化清单。

一、大模型推理显存占用的精确拆解
推理时 GPU 显存并非只存模型权重,而是由四大部分组成:模型权重、KV Cache、Activation(中间张量)以及 CUDA 上下文与临时缓冲。准确测算每一项是容量规划的前提,也是诊断 OOM 的第一步。
1. 模型权重是最容易算的部分。一个有 P 个参数的模型,在 FP16/BF16 下占 2P 字节,在 INT8 下占 P 字节,在 FP8 下同样是 P 字节(E4M3/E5M2 均为 1 字节)。以 Llama 3 70B 为例:FP16 权重 = 70 × 2 = 140 GB,单张 H100 80GB 放不下,必须张量并行切到 2 张卡(每卡 70 GB 权重,仍接近上限)。这也是为什么 70B 级别模型在生产中通常需要至少 2×80GB 或 4×40GB 的配置。
2. KV Cache是推理独有的显存大头,且随序列长度线性增长。精确公式如下:
1
2
3
4
5
6
7
8
9
10
11
12
13
14 # KV Cache 显存精确计算
# L = 层数, H = KV 头数, D = 头维度, S = 序列长度, B = batch, P = 精度字节数
def kv_cache_bytes(L, H, D, S, B, P=2):
# 每 token 每层: 2 (K+V) × H × D × P
per_token_per_layer = 2 * H * D * P
return L * per_token_per_layer * S * B
# Llama 3 70B: L=80, H_kv=8 (GQA), D=128
# 单条 4K 序列:
print(kv_cache_bytes(80, 8, 128, 4096, 1)) # ≈ 1.07 GB
# 32 条并发 4K 序列:
print(kv_cache_bytes(80, 8, 128, 4096, 32)) # ≈ 34.4 GB
# 32 条并发 32K 序列 (长上下文):
print(kv_cache_bytes(80, 8, 128, 32768, 32)) # ≈ 275 GB — 单卡放不下
可以看到,KV Cache 在长上下文场景下会迅速膨胀。Llama 3 70B 在 32K 上下文、batch=32 时 KV Cache 高达 275 GB,远超权重本身。这也是 PagedAttention、KV Cache 量化、MLA 等技术存在的根本原因。相比之下,3. Activation(中间张量)在推理时通常远小于训练,Decode 阶段单 token 的 Activation 只有几十 MB 量级,一般不是瓶颈;但 Prefill 阶段长 prompt 的 Activation 可以达到 GB 级,需要注意。
4. CUDA 上下文与临时缓冲是常被忽视的部分。每张卡的 CUDA Context 约 300-500 MB,NCCL 通信缓冲在多卡推理时占用 200 MB-1 GB,vLLM 的 CUDA Graph 还会为每个 batch size 预留一份图缓冲(每个图约 50-200 MB)。综合下来这块”隐性显存”通常占 1-2 GB,规划时要预留出来。
二、HBM 带宽:为什么 Decode 阶段是访存受限
理解了显存构成,接下来看带宽。GPU 的 HBM(High Bandwidth Memory)带宽是推理吞吐的硬上限之一。下表列出了主流推理 GPU 的关键规格:
| GPU | HBM 容量 | 带宽 (TB/s) | FP16 算力 (TFLOPS) | 算力/带宽比 |
|---|---|---|---|---|
| H100 SXM | 80 GB | 3.35 | 989 | 295 |
| H100 PCIe | 80 GB | 2.0 | 756 | 378 |
| A100 80GB | 80 GB | 2.0 | 312 | 156 |
| A100 40GB | 40 GB | 1.55 | 312 | 201 |
| L40S | 48 GB | 0.866 | 362 | 418 |
| RTX 4090 | 24 GB | 1.008 | 330 | 327 |
算力/带宽比越高,说明该卡越偏向”算力富余、带宽紧缺”,Decode 阶段越容易撞带宽墙。Decode 阶段每次只生成 1 个 token,计算量极小(主要是若干矩阵-向量乘),但每次都要读取全部权重。这种低计算强度(arithmetic intensity,FLOPs/Byte)的操作天然是访存受限的。
以 Llama 3 70B FP16 为例,Decode 单 token 的算术强度可以估算:每个参数参与约 2 次乘加(一次 forward),总 FLOPs ≈ 2 × 70B = 140 GFLOPS;读取字节数 = 140 GB。算术强度 = 140 / 140 = 1 FLOP/Byte。而 H100 的平衡点(peak compute / peak bandwidth)= 989 / 3.35 ≈ 295 FLOP/Byte。1 远小于 295,说明 Decode 是纯纯的访存受限,算力几乎被浪费——这正是带宽是瓶颈的数学证明。
三、Roofline 模型:可视化你的瓶颈在哪里

Roofline 模型把算力和带宽统一到一张图上:横轴是算术强度(FLOP/Byte),纵轴是可达算力(TFLOPS),曲线在低强度区是一条斜线(带宽受限,slope = bandwidth),在高强度区是一条水平线(算力受限,ceiling = peak compute)。两条线的交点就是该 GPU 的”脊点”(ridge point)。
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21 # Roofline 模型计算
def roofline(arithmetic_intensity, peak_bw_tbps, peak_compute_tflops):
bw_limited = arithmetic_intensity * peak_bw_tbps * 1000 # TB/s→GB/s→TFLOPS
return min(bw_limited, peak_compute_tflops)
# H100 SXM: 带宽 3.35 TB/s, FP16 算力 989 TFLOPS
ridge = 989 / 3.35 # ≈ 295 FLOP/Byte
print(f"Ridge point: {ridge:.1f} FLOP/Byte")
# Decode (Llama 70B): 算术强度 ≈ 1
ai_decode = 1
print(f"Decode achievable: {roofline(ai_decode, 3.35, 989):.1f} TFLOPS") # ≈ 3.35 TFLOPS
# 理论 token/s = 带宽 / 权重字节数
print(f"Theoretical tokens/s: {3.35e12 / 140e9:.1f}") # ≈ 23.9
# Prefill (Llama 70B, batch=32, seq=4096): 算术强度高得多
# FLOPs ≈ 2 × 70B × batch × seq = 2×70e9×32×4096 ≈ 18.4 PFLOPS
# Bytes read = 140 GB (权重只读一次)
ai_prefill = (2 * 70e9 * 32 * 4096) / (140e9)
print(f"Prefill AI: {ai_prefill:.1f} FLOP/Byte") # ≈ 131
print(f"Prefill achievable: {roofline(ai_prefill, 3.35, 989):.1f} TFLOPS")
这段代码清晰地展示了:Decode 阶段算术强度约 1,落在 Roofline 的斜线段,可达算力仅 3.35 TFLOPS(989 峰值的 0.3%);而 Prefill 阶段由于 batch 维度的计算复用,算术强度可达 100+,接近脊点,算力利用率大幅提升。这也是为什么 Continuous Batching 能提升吞吐——它通过增大 Decode 的有效 batch 来提升算术强度,让多个请求的 KV Cache 读取被权重读取”摊薄”。
四、实战:用 Nsight 工具链诊断带宽瓶颈
理论说够了,下面进入实操。诊断带宽瓶颈的标准流程是:nsys 定位哪段时间在做什么 → ncu 下钻到具体 kernel 的带宽利用率。
第一步:用 nsys 抓取推理 trace。启动 vLLM 并在推理过程中用 nsys 采集:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17 # 启动 vLLM 推理服务
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Meta-Llama-3-70B-Instruct \
--tensor-parallel-size 2 \
--port 8000 &
# 发送推理请求的同时采集 nsys profile
nsys profile -t cuda,nvtx,osrt,cudnn,cublas \
--delay 5 --duration 30 \
-o llama70b_decode_profile \
-f true \
-- curl -s http://localhost:8000/v1/completions \
-H "Content-Type: application/json" \
-d '{"model":"meta-llama/Meta-Llama-3-70B-Instruct","prompt":"Explain attention mechanism","max_tokens":256}'
# 生成 .qdrep 文件,用 Nsight Systems GUI 打开
nsys stats llama70b_decode_profile.qdrep # 命令行统计报告
第二步:在 nsys 时间线中识别 Decode 区间。Prefill 通常表现为一大块连续的 GEMM kernel(大计算量),Decode 则表现为周期性的小 kernel 序列(每生成一个 token 一轮)。用 NVTX 标记可以精确定位:
1
2
3
4
5
6
7
8
9
10 # 在推理代码中添加 NVTX 标记
import torch
from torch.cuda import nvtx
@torch.no_grad()
def decode_step(model, tokens, kv_cache):
nvtx.range_push("decode_step")
logits = model.forward(tokens, kv_cache=kv_cache)
nvtx.range_pop()
return logits
第三步:用 ncu 下钻到 kernel 级带宽指标。对 Decode 区间内的 kernel 单独 profiling:
1
2
3
4
5
6
7
8
9
10
11
12
13 # 对特定 kernel 做 detailed profiling
ncu --set full \
--kernel-name "gemm" \
--target-processes all \
--launch-skip 10 --launch-count 5 \
-o decode_kernel_profile \
python benchmark_decode.py
# 关注的关键指标:
# - dram__bytes_read.sum / dram__bytes_write.sum (实际 DRAM 读写量)
# - sm__throughput.avg.pct_of_peak_sustained_elapsed (SM 利用率)
# - dram__throughput.avg.pct_of_peak_sustained_elapsed (DRAM 带宽利用率)
# - gpu__time_duration.sum (kernel 耗时)
如果看到
|
1
|
dram__throughput.avg.pct_of_peak_sustained_elapsed
|
接近 80-90%,而
|
1
|
sm__throughput
|
只有 5-15%,就确认了该 kernel 是访存受限——这正是 Decode 阶段 GEMV 的典型特征。
五、带宽优化清单:从量化到算子融合
确认了带宽瓶颈后,优化手段大致分三个层次:减少要读的数据量、让读取更高效、提高计算复用。
层次一:减少数据量是最直接的手段。权重量化(INT8/FP8)可以把权重读取量减半甚至降到 1/4;KV Cache 量化(INT8/INT4)能减少 KV 读取;更激进的 W4A16(如 AWQ、GPTQ)把权重量化到 4-bit,理论带宽需求降到 FP16 的 1/4。下表对比了不同精度对带宽的影响:
| 量化方案 | 权重字节/参数 | Llama 70B 权重体积 | 理论 Decode token/s (H100 3.35TB/s) | 精度损失 |
|---|---|---|---|---|
| FP16/BF16 | 2 | 140 GB | ~23.9 | 基线 |
| FP8 (E4M3) | 1 | 70 GB | ~47.8 | <1% |
| INT8 | 1 | 70 GB | ~47.8 | 0.5-2% |
| W4A16 (AWQ/GPTQ) | 0.5 | 35 GB | ~95.7 | 1-3% |
| W4A8 | 0.5 | 35 GB | ~95.7 | 2-4% |
层次二:让读取更高效。这涉及访存模式优化:确保 GEMV 的权重矩阵按连续内存布局存储(行优先或列优先视 kernel 而定);用
|
1
|
torch.compile
|
或自定义 Triton kernel 做 kernel fusion,把多个小 kernel 合并成一个大 kernel,减少 kernel launch 开销和中间结果的 HBM 往返;使用 CUDA Graph 消除 Decode 阶段每步的 kernel launch 开销(vLLM 默认开启)。
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24 # vLLM 中开启 CUDA Graph 加速 Decode
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Meta-Llama-3-70B-Instruct \
--enforce-eager # 关闭则启用 CUDA Graph (默认)
# 不加 --enforce-eager 即启用 CUDA Graph
# 自定义 Triton kernel 融合示例: fused RMSNorm + QKV projection
import triton
import triton.language as tl
@triton.jit
def fused_rmsnorm_qkv_kernel(
x_ptr, weight_ptr, qkv_weight_ptr, out_ptr,
N: tl.constexpr, H: tl.constexpr, D: tl.constexpr
):
row = tl.program_id(0)
cols = tl.arange(0, D)
x = tl.load(x_ptr + row * D + cols)
# RMSNorm (减少一次中间写回 HBM)
rms = tl.sqrt(tl.mean(x * x) + 1e-6)
x_norm = x / rms * tl.load(weight_ptr + cols)
# QKV projection (融合减少一次 HBM 读写)
# ... 简化示意
tl.store(out_ptr + row * H * 3 + cols, x_norm)
层次三:提高计算复用。Continuous Batching 通过在同一 step 内 batch 多个 Decode 请求,让权重只读一次却服务多个请求,等效提升了算术强度。batch=8 时,Decode 算术强度从 1 提升到约 8,可达算力从 3.35 TFLOPS 提升到约 26.8 TFLOPS——这就是 batching 提升吞吐的本质。Speculative Decoding 则通过让小模型一次猜多个 token、大模型一次验证多个 token,提高了大模型每次 forward 的有效计算量,变相提升了算术强度。
六、监控与持续诊断:DCGM 指标体系
单次 profiling 解决一次性问题,生产环境需要持续的带宽监控。NVIDIA DCGM(Data Center GPU Manager)提供了关键的带宽相关指标,可以接入 Prometheus + Grafana:
<
pre>
|
1
2 3 4 5 6 7 8 |
# dcgm-exporter 配置 (重点关注带宽指标)
# /etc/dcgm-exporter/dcp-metrics-dcgm.csv DCGM_FI_DEV_GPU_UTIL, gauge, GPU 利用率 (SM active) DCGM_FI_DEV_MEM_COPY_UTIL, gauge, 显存带宽利用率 DCGM_FI_DEV_FB_USED, gauge, 已用显存 (MB) DCGM_FI_DEV_FB_FREE, gauge, 空闲显存 (MB) DCGM_FI_PROF_PIPE_TENSOR_OP_RATE, gauge, Tensor Core 利用率 DCGM_FI_PROF_DRAM_ACTIVE, gauge, DRAM 活跃度 (带宽利用率)</pre> |
1
2
3
4
5
6 # 快速查看当前 GPU 带宽利用率
dcgmi dmon -e 1004,1005,1009 # mem_copy_util, gpu_util, fb_used
# 或用 nvidia-smi 查询带宽相关指标
nvidia-smi dmon -s m # 显示显存利用率
nvidia-smi query-gpu=memory.used,memory.free,utilization.memory,utilization.gpu --format=csv -l 1
在 Grafana 看板中,如果
|
1
|
DCGM_FI_PROF_DRAM_ACTIVE
|
长期 > 85% 而
|
1
|
DCGM_FI_DEV_GPU_UTIL
|
< 20%,基本可以判定你的推理服务是带宽受限的,此时再堆算力(换更高 TFLOPS 的卡)收益有限,应该优先考虑量化、batching 或更高带宽的卡(如 H100 → H200 的 4.8 TB/s)。
总结
大模型推理的性能优化是一场在算力和带宽之间的权衡游戏。本文从显存占用的精确拆解出发,给出了一套完整的带宽诊断方法论:先算清 VRAM 四部分(权重、KV Cache、Activation、上下文)确认容量可行;再用算术强度和 Roofline 模型判断是否带宽受限;然后用 nsys + ncu 定位到具体的访存热点 kernel;最后从减数据(量化)、提效率(kernel fusion、CUDA Graph)、增复用(batching、speculative decoding)三个层次施策。生产环境则用 DCGM 做持续监控,及时发现带宽墙。记住一个经验法则:Decode 阶段几乎总是访存受限的,优化重点在带宽而非算力;Prefill 阶段更偏算力受限,优化重点在 Tensor Core 利用率和并行度。把这两个阶段分开诊断、分别优化,才能把 GPU 的每一份资源都榨干。
汤不热吧