PyTorch 自动混合精度

原文作者:Michael Carilli。官方页面记录:创建于 2020 年 9 月 15 日,更新于 2025 年 1 月 30 日,最近一次官方核验日期为 2024 年 11 月 5 日。

自动混合精度工具让一些操作使用 torch.float32(float),另一些操作使用 torch.float16(half)。线性层、卷积等操作在 float16 或 bfloat16 下通常更快;归约等操作往往需要 float32 的动态范围。混合精度为每个操作选择适合的数据类型,从而可能减少网络运行时间和内存占用。

通常,自动混合精度训练将 torch.autocast 与 GradScaler 配合使用。原文文字链接中沿用 torch.cuda.amp.GradScaler 的历史名称,以下示例代码实际使用 torch.amp.GradScaler。

本配方先测量一个简单网络在默认精度下的性能,再逐步加入 autocast 和 GradScaler,让同一个网络使用混合精度。

原文提供独立 Python 脚本,并将历史要求写为 PyTorch 1.6 或更新版本以及支持 CUDA 的 GPU。当前示例包含 torch.set_default_device、torch.amp.GradScaler 等 API,运行前应按实际 PyTorch 版本核对;不能据历史最低版本描述认定整段当前代码兼容所有旧版本。

混合精度主要有利于支持 Tensor Core 的架构,例如 Volta、Turing、Ampere。原文预期这些架构上可以出现约 2 至 3 倍加速;较早的 Kepler、Maxwell、Pascal 架构可能只有较小收益。可运行 nvidia-smi 查看 GPU 信息。本稿没有运行配方或测量加速,实际收益取决于设备、算子和负载。

import torch, time, gc

# Timing utilities
start_time = None

def start_timer():
    global start_time
    gc.collect()
    torch.cuda.empty_cache()
    torch.cuda.reset_max_memory_allocated()
    torch.cuda.synchronize()
    start_time = time.time()

def end_timer_and_print(local_msg):
    torch.cuda.synchronize()
    end_time = time.time()
    print("\n" + local_msg)
    print("Total execution time = {:.3f} sec".format(end_time - start_time))
    print("Max memory used by tensors = {} bytes".format(torch.cuda.max_memory_allocated()))

一个简单网络

下面由线性层和 ReLU 组成的网络,是原文用来观察混合精度加速的例子。

def make_model(in_size, out_size, num_layers):
    layers = []
    for _ in range(num_layers - 1):
        layers.append(torch.nn.Linear(in_size, in_size))
        layers.append(torch.nn.ReLU())
    layers.append(torch.nn.Linear(in_size, out_size))
    return torch.nn.Sequential(*tuple(layers)).cuda()

batch_size、in_size、out_size 和 num_layers 被选得足够大,以便给 GPU 足够多的工作。通常,GPU 饱和时混合精度收益最大;小网络可能受 CPU 限制,此时混合精度不一定改善性能。

线性层参与计算的维度还被设成 8 的倍数,以便在支持 Tensor Core 的 GPU 上使用它们。练习:改变这些尺寸,观察混合精度加速如何变化。

batch_size = 512 # Try, for example, 128, 256, 513.
in_size = 4096
out_size = 4096
num_layers = 3
num_batches = 50
epochs = 3

device = 'cuda' if torch.cuda.is_available() else 'cpu'
torch.set_default_device(device)

# Creates data in default precision.
# The same data is used for both default and mixed precision trials below.
# You don't need to manually change inputs' ``dtype`` when enabling mixed precision.
data = [torch.randn(batch_size, in_size) for _ in range(num_batches)]
targets = [torch.randn(batch_size, out_size) for _ in range(num_batches)]

loss_fn = torch.nn.MSELoss().cuda()

数据以默认精度创建,后面的默认精度和混合精度试验使用同一份数据;启用混合精度时,不需要手动改变输入的 dtype。

静态核验提醒:虽然 device 有 CPU 回退分支,计时函数仍调用 CUDA,make_model 和损失函数也无条件使用 .cuda()。因此这份配方不能原样当作纯 CPU 脚本运行。

默认精度

不使用 AMP 时,下面的简单网络以默认精度 torch.float32 执行所有操作。

net = make_model(in_size, out_size, num_layers)
opt = torch.optim.SGD(net.parameters(), lr=0.001)

start_timer()
for epoch in range(epochs):
    for input, target in zip(data, targets):
        output = net(input)
        loss = loss_fn(output, target)
        loss.backward()
        opt.step()
        opt.zero_grad() # set_to_none=True here can modestly improve performance
end_timer_and_print("Default precision:")

加入 torch.autocast

torch.autocast 实例是上下文管理器,让脚本指定区域使用混合精度。在该区域内,CUDA 操作采用 autocast 选择的 dtype,以兼顾性能和准确性。每个操作采用什么精度、何时选择该精度,请参阅Autocast 算子参考。

for epoch in range(0): # 0 epochs, this section is for illustration only
    for input, target in zip(data, targets):
        # Runs the forward pass under ``autocast``.
        with torch.autocast(device_type=device, dtype=torch.float16):
            output = net(input)
            # output is float16 because linear layers ``autocast`` to float16.
            assert output.dtype is torch.float16

            loss = loss_fn(output, target)
            # loss is float32 because ``mse_loss`` layers ``autocast`` to float32.
            assert loss.dtype is torch.float32

        # Exits ``autocast`` before backward().
        # Backward passes under ``autocast`` are not recommended.
        # Backward ops run in the same ``dtype`` ``autocast`` chose for corresponding forward ops.
        loss.backward()
        opt.step()
        opt.zero_grad() # set_to_none=True here can modestly improve performance

此例中的线性层输出为 float16,mse_loss 的输出为 float32。执行 backward() 前先退出 autocast:不建议把反向过程包在该上下文中;反向操作使用对应前向操作所选择的数据类型。range(0) 表示这段仅用于说明,不会执行训练轮次。

加入 GradScaler

梯度缩放有助于防止混合精度训练中很小的梯度下溢为零。GradScaler 可以方便地完成这些步骤。

在一次完整收敛训练开始时,使用默认参数构造一个 scaler。整个训练过程使用同一个实例;同一脚本若进行多次独立收敛训练,每次都应使用新的专属实例。它们是轻量对象。若默认参数导致网络无法收敛,原文建议提交问题。

# Constructs a ``scaler`` once, at the beginning of the convergence run, using default arguments.
# If your network fails to converge with default ``GradScaler`` arguments, please file an issue.
# The same ``GradScaler`` instance should be used for the entire convergence run.
# If you perform multiple convergence runs in the same script, each run should use
# a dedicated fresh ``GradScaler`` instance. ``GradScaler`` instances are lightweight.
scaler = torch.amp.GradScaler("cuda")

for epoch in range(0): # 0 epochs, this section is for illustration only
    for input, target in zip(data, targets):
        with torch.autocast(device_type=device, dtype=torch.float16):
            output = net(input)
            loss = loss_fn(output, target)

        # Scales loss. Calls ``backward()`` on scaled loss to create scaled gradients.
        scaler.scale(loss).backward()

        # ``scaler.step()`` first unscales the gradients of the optimizer's assigned parameters.
        # If these gradients do not contain ``inf``s or ``NaN``s, optimizer.step() is then called,
        # otherwise, optimizer.step() is skipped.
        scaler.step(opt)

        # Updates the scale for next iteration.
        scaler.update()

        opt.zero_grad() # set_to_none=True here can modestly improve performance

scaler.scale(loss).backward() 先缩放损失,并生成缩放后的梯度。scaler.step(opt) 会先反缩放优化器对应参数的梯度;如果没有 inf 或 NaN,才调用优化器的 step(),否则跳过。scaler.update() 更新下一轮的缩放因子。

组合起来:自动混合精度

下面还展示 autocast 和 GradScaler 的可选参数 enabled。设为 False 后,它们的调用不再进行混合精度或缩放工作,因此可以不写额外的 if/else,切换默认精度和混合精度。

use_amp = True

net = make_model(in_size, out_size, num_layers)
opt = torch.optim.SGD(net.parameters(), lr=0.001)
scaler = torch.amp.GradScaler("cuda" ,enabled=use_amp)

start_timer()
for epoch in range(epochs):
    for input, target in zip(data, targets):
        with torch.autocast(device_type=device, dtype=torch.float16, enabled=use_amp):
            output = net(input)
            loss = loss_fn(output, target)
        scaler.scale(loss).backward()
        scaler.step(opt)
        scaler.update()
        opt.zero_grad() # set_to_none=True here can modestly improve performance
end_timer_and_print("Mixed precision:")

检查或修改梯度,例如裁剪

scaler.scale(loss).backward() 产生的所有梯度都被缩放。如果要在 backward() 和 scaler.step(optimizer) 之间检查或修改参数的 .grad,应先调用 scaler.unscale_(optimizer) 反缩放。

for epoch in range(0): # 0 epochs, this section is for illustration only
    for input, target in zip(data, targets):
        with torch.autocast(device_type=device, dtype=torch.float16):
            output = net(input)
            loss = loss_fn(output, target)
        scaler.scale(loss).backward()

        # Unscales the gradients of optimizer's assigned parameters in-place
        scaler.unscale_(opt)

        # Since the gradients of optimizer's assigned parameters are now unscaled, clips as usual.
        # You may use the same value for max_norm here as you would without gradient scaling.
        torch.nn.utils.clip_grad_norm_(net.parameters(), max_norm=0.1)

        scaler.step(opt)
        scaler.update()
        opt.zero_grad() # set_to_none=True here can modestly improve performance

反缩放是原地进行的。随后可以像普通训练那样裁剪梯度,max_norm 可以使用与未进行梯度缩放时相同的数值。

保存与恢复

要以位级精度保存并恢复启用 AMP 的训练,使用 scaler.state_dict() 和 scaler.load_state_dict()。

保存时,把 scaler 状态与模型、优化器状态一起保存。可在一次迭代开始、尚未进行前向计算时保存,也可以在迭代结束、调用 scaler.update() 后保存。

checkpoint = {"model": net.state_dict(),
              "optimizer": opt.state_dict(),
              "scaler": scaler.state_dict()}
# Write checkpoint as desired, e.g.,
# torch.save(checkpoint, "filename")

恢复时,把 scaler 状态与模型、优化器状态一起加载。按需读取检查点,例如:

dev = torch.cuda.current_device()
checkpoint = torch.load("filename",
                        map_location = lambda storage, loc: storage.cuda(dev))
net.load_state_dict(checkpoint["model"])
opt.load_state_dict(checkpoint["optimizer"])
scaler.load_state_dict(checkpoint["scaler"])

如果检查点来自未使用 AMP 的训练,而你希望使用 AMP 继续训练,就正常加载模型和优化器状态;检查点中没有 scaler 状态,因此应创建新的 GradScaler。

如果检查点来自使用 AMP 的训练,而你希望不使用 AMP 继续训练,就正常加载模型和优化器状态,并忽略保存的 scaler 状态。

推理与评估

可以单独使用 autocast 包住推理或评估的前向过程,不需要 GradScaler。

进阶主题

自动混合精度示例包含以下进阶用法:

  • 梯度累积。
  • 梯度惩罚与二次反向传播。
  • 包含多个模型、优化器或损失的网络。
  • 多 GPU,例如 torch.nn.DataParallel 或 torch.nn.parallel.DistributedDataParallel。
  • 自定义 autograd 函数,即 torch.autograd.Function 的子类。

如果在同一脚本中进行多次独立收敛训练,每次都应使用新的专属 GradScaler。如果要在 dispatcher 中注册自定义 C++ 算子,请查看 dispatcher 教程的 autocast 小节。

排查问题

AMP 加速很小

  1. 网络可能没有给 GPU 足够多的工作,瓶颈在 CPU,因此 AMP 对 GPU 的改善不起作用。可在不耗尽显存的前提下增大批量或网络;避免过多 CPU 与 GPU 同步,例如 .item() 或打印 CUDA 张量值;尽量把许多小 CUDA 操作合并成少量大操作。
  2. 网络可能受 GPU 计算能力限制,例如包含大量矩阵乘法或卷积,但 GPU 没有 Tensor Core,此时加速较小是预期情况。
  3. 矩阵乘法尺寸可能不适合 Tensor Core,应确保参与维度是 8 的倍数。对带编码器或解码器的 NLP 模型,这一点可能不容易发现。卷积过去也有类似限制,但原文说明 CuDNN 7.3 及以后不再有这类限制;可查看 NVIDIA Apex 项目中的说明。

损失为 inf 或 NaN

先检查网络是否属于前述进阶用法,并阅读优先使用 binary_cross_entropy_with_logits 而非 binary_cross_entropy的说明。

如果确信 AMP 用法正确,可能需要提交问题;提交前可以收集以下信息:

  1. 分别给 autocast 或 GradScaler 设置 enabled=False,观察 inf 或 NaN 是否仍出现。
  2. 如果怀疑网络某一部分,例如复杂损失函数发生溢出,让该前向区域在 float32 下运行,再观察问题是否仍存在。autocast 文档的最后一个示例展示了如何局部禁用 autocast,并转换子区域输入,使其在 float32 下执行。

类型不匹配错误,可能表现为 CUDNN_STATUS_BAD_PARAM

autocast 会尽量覆盖需要或有利于类型转换的算子。显式覆盖的算子根据数值特性和实践经验选定。如果启用了 autocast 的前向区域或其后的反向过程发生类型不匹配,可能是某个算子没有被正确覆盖。

原文建议提交问题并附错误回溯。在运行脚本前设置 export TORCH_SHOW_CPP_STACKTRACES=1,可以提供更详细的后端算子错误信息。

完整独立配方可在官方页面底部下载 Python 源码、Jupyter Notebook 或 ZIP 文件。


来源:Michael Carilli,Automatic Mixed Precision。官方源文件:amp_recipe.py。按 BSD 3-Clause License许可使用。此页作中文翻译和版式整理,补充历史版本、CPU 回退与性能数值边界;保留十一段原文代码,包括原文注释。未执行训练或生成性能测量。以下保留完整许可。

BSD 3-Clause License

Copyright (c) 2017-2022, Pytorch contributors All rights reserved.

Redistribution and use in source and binary forms, with or without modification, are permitted provided that the following conditions are met:

* Redistributions of source code must retain the above copyright notice, this list of conditions and the following disclaimer.

* Redistributions in binary form must reproduce the above copyright notice, this list of conditions and the following disclaimer in the documentation and/or other materials provided with the distribution.

* Neither the name of the copyright holder nor the names of its contributors may be used to endorse or promote products derived from this software without specific prior written permission.

THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS “AS IS” AND ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT HOLDER OR CONTRIBUTORS BE LIABLE FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.

© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容