原文作者:Michael Carilli。官方页面记录:创建于 2020 年 9 月 15 日,更新于 2025 年 1 月 30 日,最近一次官方核验日期为 2024 年 11 月 5 日。
自动混合精度工具让一些操作使用 torch.float32(float),另一些操作使用 torch.float16(half)。线性层、卷积等操作在 float16 或 bfloat16 下通常更快;归约等操作往往需要 float32 的动态范围。混合精度为每个操作选择适合的数据类型,从而可能减少网络运行时间和内存占用。
通常,自动混合精度训练将 torch.autocast 与 GradScaler 配合使用。原文文字链接中沿用 torch.cuda.amp.GradScaler 的历史名称,以下示例代码实际使用 torch.amp.GradScaler。
本配方先测量一个简单网络在默认精度下的性能,再逐步加入 autocast 和 GradScaler,让同一个网络使用混合精度。
原文提供独立 Python 脚本,并将历史要求写为 PyTorch 1.6 或更新版本以及支持 CUDA 的 GPU。当前示例包含 torch.set_default_device、torch.amp.GradScaler 等 API,运行前应按实际 PyTorch 版本核对;不能据历史最低版本描述认定整段当前代码兼容所有旧版本。
混合精度主要有利于支持 Tensor Core 的架构,例如 Volta、Turing、Ampere。原文预期这些架构上可以出现约 2 至 3 倍加速;较早的 Kepler、Maxwell、Pascal 架构可能只有较小收益。可运行 nvidia-smi 查看 GPU 信息。本稿没有运行配方或测量加速,实际收益取决于设备、算子和负载。
import torch, time, gc
# Timing utilities
start_time = None
def start_timer():
global start_time
gc.collect()
torch.cuda.empty_cache()
torch.cuda.reset_max_memory_allocated()
torch.cuda.synchronize()
start_time = time.time()
def end_timer_and_print(local_msg):
torch.cuda.synchronize()
end_time = time.time()
print("\n" + local_msg)
print("Total execution time = {:.3f} sec".format(end_time - start_time))
print("Max memory used by tensors = {} bytes".format(torch.cuda.max_memory_allocated()))
一个简单网络
下面由线性层和 ReLU 组成的网络,是原文用来观察混合精度加速的例子。
def make_model(in_size, out_size, num_layers):
layers = []
for _ in range(num_layers - 1):
layers.append(torch.nn.Linear(in_size, in_size))
layers.append(torch.nn.ReLU())
layers.append(torch.nn.Linear(in_size, out_size))
return torch.nn.Sequential(*tuple(layers)).cuda()
batch_size、in_size、out_size 和 num_layers 被选得足够大,以便给 GPU 足够多的工作。通常,GPU 饱和时混合精度收益最大;小网络可能受 CPU 限制,此时混合精度不一定改善性能。
线性层参与计算的维度还被设成 8 的倍数,以便在支持 Tensor Core 的 GPU 上使用它们。练习:改变这些尺寸,观察混合精度加速如何变化。
batch_size = 512 # Try, for example, 128, 256, 513.
in_size = 4096
out_size = 4096
num_layers = 3
num_batches = 50
epochs = 3
device = 'cuda' if torch.cuda.is_available() else 'cpu'
torch.set_default_device(device)
# Creates data in default precision.
# The same data is used for both default and mixed precision trials below.
# You don't need to manually change inputs' ``dtype`` when enabling mixed precision.
data = [torch.randn(batch_size, in_size) for _ in range(num_batches)]
targets = [torch.randn(batch_size, out_size) for _ in range(num_batches)]
loss_fn = torch.nn.MSELoss().cuda()
数据以默认精度创建,后面的默认精度和混合精度试验使用同一份数据;启用混合精度时,不需要手动改变输入的 dtype。
静态核验提醒:虽然 device 有 CPU 回退分支,计时函数仍调用 CUDA,make_model 和损失函数也无条件使用 .cuda()。因此这份配方不能原样当作纯 CPU 脚本运行。
默认精度
不使用 AMP 时,下面的简单网络以默认精度 torch.float32 执行所有操作。
net = make_model(in_size, out_size, num_layers)
opt = torch.optim.SGD(net.parameters(), lr=0.001)
start_timer()
for epoch in range(epochs):
for input, target in zip(data, targets):
output = net(input)
loss = loss_fn(output, target)
loss.backward()
opt.step()
opt.zero_grad() # set_to_none=True here can modestly improve performance
end_timer_and_print("Default precision:")
加入 torch.autocast
torch.autocast 实例是上下文管理器,让脚本指定区域使用混合精度。在该区域内,CUDA 操作采用 autocast 选择的 dtype,以兼顾性能和准确性。每个操作采用什么精度、何时选择该精度,请参阅Autocast 算子参考。
for epoch in range(0): # 0 epochs, this section is for illustration only
for input, target in zip(data, targets):
# Runs the forward pass under ``autocast``.
with torch.autocast(device_type=device, dtype=torch.float16):
output = net(input)
# output is float16 because linear layers ``autocast`` to float16.
assert output.dtype is torch.float16
loss = loss_fn(output, target)
# loss is float32 because ``mse_loss`` layers ``autocast`` to float32.
assert loss.dtype is torch.float32
# Exits ``autocast`` before backward().
# Backward passes under ``autocast`` are not recommended.
# Backward ops run in the same ``dtype`` ``autocast`` chose for corresponding forward ops.
loss.backward()
opt.step()
opt.zero_grad() # set_to_none=True here can modestly improve performance
此例中的线性层输出为 float16,mse_loss 的输出为 float32。执行 backward() 前先退出 autocast:不建议把反向过程包在该上下文中;反向操作使用对应前向操作所选择的数据类型。range(0) 表示这段仅用于说明,不会执行训练轮次。
加入 GradScaler
梯度缩放有助于防止混合精度训练中很小的梯度下溢为零。GradScaler 可以方便地完成这些步骤。
在一次完整收敛训练开始时,使用默认参数构造一个 scaler。整个训练过程使用同一个实例;同一脚本若进行多次独立收敛训练,每次都应使用新的专属实例。它们是轻量对象。若默认参数导致网络无法收敛,原文建议提交问题。
# Constructs a ``scaler`` once, at the beginning of the convergence run, using default arguments.
# If your network fails to converge with default ``GradScaler`` arguments, please file an issue.
# The same ``GradScaler`` instance should be used for the entire convergence run.
# If you perform multiple convergence runs in the same script, each run should use
# a dedicated fresh ``GradScaler`` instance. ``GradScaler`` instances are lightweight.
scaler = torch.amp.GradScaler("cuda")
for epoch in range(0): # 0 epochs, this section is for illustration only
for input, target in zip(data, targets):
with torch.autocast(device_type=device, dtype=torch.float16):
output = net(input)
loss = loss_fn(output, target)
# Scales loss. Calls ``backward()`` on scaled loss to create scaled gradients.
scaler.scale(loss).backward()
# ``scaler.step()`` first unscales the gradients of the optimizer's assigned parameters.
# If these gradients do not contain ``inf``s or ``NaN``s, optimizer.step() is then called,
# otherwise, optimizer.step() is skipped.
scaler.step(opt)
# Updates the scale for next iteration.
scaler.update()
opt.zero_grad() # set_to_none=True here can modestly improve performance
scaler.scale(loss).backward() 先缩放损失,并生成缩放后的梯度。scaler.step(opt) 会先反缩放优化器对应参数的梯度;如果没有 inf 或 NaN,才调用优化器的 step(),否则跳过。scaler.update() 更新下一轮的缩放因子。
组合起来:自动混合精度
下面还展示 autocast 和 GradScaler 的可选参数 enabled。设为 False 后,它们的调用不再进行混合精度或缩放工作,因此可以不写额外的 if/else,切换默认精度和混合精度。
use_amp = True
net = make_model(in_size, out_size, num_layers)
opt = torch.optim.SGD(net.parameters(), lr=0.001)
scaler = torch.amp.GradScaler("cuda" ,enabled=use_amp)
start_timer()
for epoch in range(epochs):
for input, target in zip(data, targets):
with torch.autocast(device_type=device, dtype=torch.float16, enabled=use_amp):
output = net(input)
loss = loss_fn(output, target)
scaler.scale(loss).backward()
scaler.step(opt)
scaler.update()
opt.zero_grad() # set_to_none=True here can modestly improve performance
end_timer_and_print("Mixed precision:")
检查或修改梯度,例如裁剪
scaler.scale(loss).backward() 产生的所有梯度都被缩放。如果要在 backward() 和 scaler.step(optimizer) 之间检查或修改参数的 .grad,应先调用 scaler.unscale_(optimizer) 反缩放。
for epoch in range(0): # 0 epochs, this section is for illustration only
for input, target in zip(data, targets):
with torch.autocast(device_type=device, dtype=torch.float16):
output = net(input)
loss = loss_fn(output, target)
scaler.scale(loss).backward()
# Unscales the gradients of optimizer's assigned parameters in-place
scaler.unscale_(opt)
# Since the gradients of optimizer's assigned parameters are now unscaled, clips as usual.
# You may use the same value for max_norm here as you would without gradient scaling.
torch.nn.utils.clip_grad_norm_(net.parameters(), max_norm=0.1)
scaler.step(opt)
scaler.update()
opt.zero_grad() # set_to_none=True here can modestly improve performance
反缩放是原地进行的。随后可以像普通训练那样裁剪梯度,max_norm 可以使用与未进行梯度缩放时相同的数值。
保存与恢复
要以位级精度保存并恢复启用 AMP 的训练,使用 scaler.state_dict() 和 scaler.load_state_dict()。
保存时,把 scaler 状态与模型、优化器状态一起保存。可在一次迭代开始、尚未进行前向计算时保存,也可以在迭代结束、调用 scaler.update() 后保存。
checkpoint = {"model": net.state_dict(),
"optimizer": opt.state_dict(),
"scaler": scaler.state_dict()}
# Write checkpoint as desired, e.g.,
# torch.save(checkpoint, "filename")
恢复时,把 scaler 状态与模型、优化器状态一起加载。按需读取检查点,例如:
dev = torch.cuda.current_device()
checkpoint = torch.load("filename",
map_location = lambda storage, loc: storage.cuda(dev))
net.load_state_dict(checkpoint["model"])
opt.load_state_dict(checkpoint["optimizer"])
scaler.load_state_dict(checkpoint["scaler"])
如果检查点来自未使用 AMP 的训练,而你希望使用 AMP 继续训练,就正常加载模型和优化器状态;检查点中没有 scaler 状态,因此应创建新的 GradScaler。
如果检查点来自使用 AMP 的训练,而你希望不使用 AMP 继续训练,就正常加载模型和优化器状态,并忽略保存的 scaler 状态。
推理与评估
可以单独使用 autocast 包住推理或评估的前向过程,不需要 GradScaler。
进阶主题
自动混合精度示例包含以下进阶用法:
- 梯度累积。
- 梯度惩罚与二次反向传播。
- 包含多个模型、优化器或损失的网络。
- 多 GPU,例如
torch.nn.DataParallel或torch.nn.parallel.DistributedDataParallel。 - 自定义 autograd 函数,即
torch.autograd.Function的子类。
如果在同一脚本中进行多次独立收敛训练,每次都应使用新的专属 GradScaler。如果要在 dispatcher 中注册自定义 C++ 算子,请查看 dispatcher 教程的 autocast 小节。
排查问题
AMP 加速很小
- 网络可能没有给 GPU 足够多的工作,瓶颈在 CPU,因此 AMP 对 GPU 的改善不起作用。可在不耗尽显存的前提下增大批量或网络;避免过多 CPU 与 GPU 同步,例如
.item()或打印 CUDA 张量值;尽量把许多小 CUDA 操作合并成少量大操作。 - 网络可能受 GPU 计算能力限制,例如包含大量矩阵乘法或卷积,但 GPU 没有 Tensor Core,此时加速较小是预期情况。
- 矩阵乘法尺寸可能不适合 Tensor Core,应确保参与维度是 8 的倍数。对带编码器或解码器的 NLP 模型,这一点可能不容易发现。卷积过去也有类似限制,但原文说明 CuDNN 7.3 及以后不再有这类限制;可查看 NVIDIA Apex 项目中的说明。
损失为 inf 或 NaN
先检查网络是否属于前述进阶用法,并阅读优先使用 binary_cross_entropy_with_logits 而非 binary_cross_entropy的说明。
如果确信 AMP 用法正确,可能需要提交问题;提交前可以收集以下信息:
- 分别给
autocast或GradScaler设置enabled=False,观察inf或NaN是否仍出现。 - 如果怀疑网络某一部分,例如复杂损失函数发生溢出,让该前向区域在
float32下运行,再观察问题是否仍存在。autocast 文档的最后一个示例展示了如何局部禁用 autocast,并转换子区域输入,使其在float32下执行。
类型不匹配错误,可能表现为 CUDNN_STATUS_BAD_PARAM
autocast 会尽量覆盖需要或有利于类型转换的算子。显式覆盖的算子根据数值特性和实践经验选定。如果启用了 autocast 的前向区域或其后的反向过程发生类型不匹配,可能是某个算子没有被正确覆盖。
原文建议提交问题并附错误回溯。在运行脚本前设置 export TORCH_SHOW_CPP_STACKTRACES=1,可以提供更详细的后端算子错误信息。
完整独立配方可在官方页面底部下载 Python 源码、Jupyter Notebook 或 ZIP 文件。











暂无评论内容