原文:Hugging Face Transformers 文档贡献者,Continuous batching。原始文档文件署名 Copyright 2025 The HuggingFace Team,采用 Apache License 2.0;完整许可证见文末。中文翻译、技术整理与示意图:未完纪。
连续批处理会在每个生成步骤重新安排批次。已经完成的请求退出,等待中的新请求及时补入,无须等同一批里的最长请求结束。这种调度方式旨在提高 GPU 利用率;具体吞吐和延迟仍取决于模型、注意力内核、输入长度、并发与硬件,不能仅凭启用某个开关保证提升。
本文根据 2026-10-05 读取的官方完整正文整理。所有模型加载和 GPU 示例均未执行,仅作静态审查。文档面向 Python 推理管理;官方对生产服务另指向 transformers serve。下面的队列代码本身没有实现 HTTP 鉴权、租户隔离或限流。

一、先用 generate_batch 处理一组提示
generate_batch() 接收已经分词的提示列表,内部负责连续调度,并阻塞到所有请求完成,然后返回各请求结果。需要实时提交或流式返回时,再直接使用管理器。
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from transformers.generation import ContinuousBatchingConfig, GenerationConfig
model_id = "Qwen/Qwen3-4B"
model = AutoModelForCausalLM.from_pretrained(
model_id,
attn_implementation="flash_attention_2",
device_map="auto",
dtype=torch.bfloat16,
)
tokenizer = AutoTokenizer.from_pretrained(model_id)
prompts = [
"What's up?",
"Name a cat breed.",
"Write a detailed history of quantum mechanics.",
]
inputs = [tokenizer.encode(prompt) for prompt in prompts]
generation_config = GenerationConfig(
max_new_tokens=64,
eos_token_id=tokenizer.eos_token_id,
)
outputs = model.generate_batch(
inputs=inputs, generation_config=generation_config
)
for request_id, output in outputs.items():
text = tokenizer.decode(output.generated_tokens, skip_special_tokens=True)
print(f"[{request_id}] {text}")
模型和 tokenizer 的下载需要相应网络访问、磁盘和显存;FlashAttention 还需兼容的设备与依赖。这里保留原文模型和精度选择,它们不是适用于任意设备的默认值。部署时应固定库版本和模型 revision,并先核对模型本身的许可与访问条件。
二、用管理器分别控制提交与取回
ContinuousBatchingManager 在后台线程运行,每一步查看哪些请求结束、哪些请求可以进入当前批次。上下文管理器负责开始与结束生命周期,适合把资源清理与代码作用域绑定。
with model.continuous_batching_context_manager(
generation_config=generation_config
) as manager:
manager.add_request(
input_ids=tokenizer.encode("Write a detailed history of quantum mechanics."),
request_id="long",
max_new_tokens=512,
)
manager.add_request(
input_ids=tokenizer.encode("What's up?"),
request_id="short_0",
max_new_tokens=32,
)
manager.add_request(
input_ids=tokenizer.encode("Name a cat breed."),
request_id="short_1",
max_new_tokens=32,
)
for result in manager:
text = tokenizer.decode(result.generated_tokens, skip_special_tokens=True)
print(f"[{result.request_id}] {text}")
短请求完成后可以释放批次中的位置,长请求继续生成。结果按照完成情况到达,不应假定迭代顺序等于输入顺序;用 request_id 关联输入和输出。add_request() 可接受显式 ID,也可让管理器自动生成。add_requests(inputs=inputs) 一次提交多项,启用块共享时会自动排序以改善前缀缓存命中。取消则使用 manager.cancel_request(request_id="my_request")。
若不用上下文管理器,可显式初始化与启动;下例把资源释放放在 finally 中,是针对原文片段所作的生命周期补充。
manager = model.init_continuous_batching(
generation_config=generation_config
)
manager.start()
try:
manager.add_request(
input_ids=tokenizer.encode("Name a cat breed."),
request_id="one",
)
for result in manager:
print(result.request_id, result.generated_tokens)
finally:
manager.destroy()
停止与销毁是两件事
manager.stop() 默认停止接受新提交,等待排队和活动请求完成后退出后台线程。传入 hard_stop=True 会立即放弃待处理工作:排队和活动请求以 RuntimeError 失败,而不是正常完成。
调用 stop() 后,新的 add_request() 或 add_requests() 会被丢弃并记录警告;同一个管理器仍可再次 start()。destroy() 用于释放分布式资源:若尚在运行会先停止,但销毁后不能重启。上下文退出时会调用 stop(),通常也会 destroy();设置 persistent_manager=True 才会把管理器缓存在模型上供下一会话复用。
编者补充:服务关闭时应先关闭业务入口,再按所选策略排空或取消任务。不能把“调用提交函数没有抛异常”当成请求已经被接收的充分证据,因为停止阶段的新请求可能只是被丢弃并记日志。
每个请求使用不同采样参数
per_request_processors=True 允许同一前向计算中的请求分别使用 temperature、top_k 和 top_p。仅创建配置对象还不够,必须把它传给管理器;同时,在基础 GenerationConfig 中将需要的处理器设为非默认值,以便运行时创建对应 logits processor。例如 temperature 应不为 None 或 1,随后仍可提交 temperature 为 1 的请求。
下面是补全后的示例,与原文分散片段的差异是:显式传入两种配置,开启采样,并为三个参数设置非默认初值。
sampling = GenerationConfig(
do_sample=True,
max_new_tokens=64,
eos_token_id=tokenizer.eos_token_id,
temperature=0.8,
top_p=0.95,
top_k=20,
)
cb_config = ContinuousBatchingConfig(per_request_processors=True)
with model.continuous_batching_context_manager(
generation_config=sampling,
continuous_batching_config=cb_config,
) as manager:
manager.add_request(
input_ids=tokenizer.encode("Invent a cat name."),
request_id="creative", temperature=0.9, top_p=0.95,
)
manager.add_request(
input_ids=tokenizer.encode("Name a cat breed."),
request_id="precise", temperature=0.1, top_k=10,
)
for result in manager:
print(result.request_id, result.generated_tokens)
普通取回与流式取回
直接迭代管理器即可依次接收结果。get_result() 从输出队列取下一项;传入 request_id 时,如果队首结果并不匹配,该结果会重新入队,方法返回 None。因此,None 不等同于目标请求失败或已经完成;高并发服务也不宜围绕它无休止忙轮询。
# 在已启动的 manager 作用域中使用:
result = manager.get_result()
result_for_one = manager.get_result(request_id="my_request")
要边生成边返回,提交时设置 streaming=True,再使用 request_id_iter() 接收该请求的部分输出。原文演示只解码每次结果中的最后一个 token:
from transformers.generation.continuous_batching import RequestStatus
manager.add_request(
input_ids=tokenizer.encode("Tell a short story."),
request_id="streamed",
streaming=True,
)
for chunk in manager.request_id_iter(request_id="streamed"):
token = tokenizer.decode(
chunk.generated_tokens[-1:], skip_special_tokens=True
)
print(token, end="", flush=True)
if chunk.status == RequestStatus.FINISHED:
break
这段要放在已运行的管理器作用域内。编者提示:逐个 token 独立解码是原文的最小演示,实际界面应核对所用 tokenizer 的增量解码行为,处理多字节字符、空白和错误/取消状态,而不是把该片段当作完整网络流协议。
三、生成配置和调度配置各管一层
GenerationConfig 管理采样、停止等生成语义;ContinuousBatchingConfig 管理 KV 缓存、批次预算、调度和硬件执行。应通过独立参数 continuous_batching_config 传入。原文明确:把它放入 GenerationConfig 的旧用法已经弃用、会产生 FutureWarning,计划在 v5.19 移除。
cb_config = ContinuousBatchingConfig(
max_memory_percent=0.8,
block_size=256,
scheduler_type="fifo",
)
outputs = model.generate_batch(
inputs=inputs,
generation_config=generation_config,
continuous_batching_config=cb_config,
)
max_memory_percent=0.8 表示将空闲 GPU 内存的一定比例分配给 KV 缓存,并非模型总显存占用只会达到 80%。调度器可选 "fifo" 或 "prefill_first",它们对新请求预填充与活动请求解码的优先关系不同。
缓存块大小与 prefill 预算
block_size 是一个 KV 块可容纳的 token 数,至少为 4,低于此值 generate_batch 会抛 ValueError。缓存还会保留两个块用于填充和哨兵记账,块太小会让固定开销变得明显;块太大则可能浪费每个序列最后一块未填满的空间。文档默认值为 256,与解码快速内核的配置相配。设置示例为 ContinuousBatchingConfig(block_size=128),不是建议把它压到最低。
max_batch_tokens 限制单次前向计算处理的查询 token 总数。更高预算能装入更多 prompt token,适用于 prefill 为瓶颈的请求;超出预算的长提示会分块分步处理。页面给出的默认值是 8192,会按可用显存缩小,但不低于 256。
输入缓冲区与 KV 块使用同一显存预算:把 max_batch_tokens 调大,会留给 num_blocks 更少空间,进而影响长请求和并发。ContinuousBatchingConfig(max_batch_tokens=16384) 只是增大预算的写法,需按实际资源验证。
批次请求数与活动请求安全余量
max_requests_per_batch 限制一次前向计算的请求数。词表较大时,logits 张量尤其昂贵,采样转换到 fp32 还会扩大占用;减少请求数上限可以限制这一部分缓冲区。文档默认按提交的提示数设置,后备值为 1024,再受 token 与块预算限制。例如可写 ContinuousBatchingConfig(max_requests_per_batch=256)。
safety_margin 取值为 0 到 1。当空闲 KV 块低于 safety_margin * num_blocks,调度器暂停接纳新的 prefill,先推进现有请求的解码,避免新工作挤占全部空间。0 表示关闭余量。FIFO 默认 0.15,prefill_first 默认 0.0;这是调度策略,不是防止所有 OOM 的保证。
返回 token 的对数概率
设置 ContinuousBatchingConfig(return_logprobs=True) 后,结果会包含生成 token 的 logprobs,可为某些强化学习流程提供输入。应确保该配置确实传给本次生成,而不只是声明变量。
logprob_config = ContinuousBatchingConfig(return_logprobs=True)
outputs = model.generate_batch(
inputs=inputs,
generation_config=generation_config,
continuous_batching_config=logprob_config,
)
for request_id, output in outputs.items():
for token_id, log_prob in zip(output.generated_tokens, output.logprobs):
print(tokenizer.decode([token_id]), log_prob)
四、用更多资源换取执行效率
| 机制 | 主要作用 | 需要考虑的代价 |
|---|---|---|
| CUDA graphs | 复用相同形状的 GPU 执行图,减少 CPU 分发开销 | 记录图需要显存,形状填充可能增加计算 |
| 异步批处理 | 把下一批 CPU 调度与当前批 GPU 计算重叠 | 需要 CUDA graphs,输入张量显存约翻倍,不是整个模型显存翻倍 |
| 编译 | 通过 torch.compile 优化前向路径 | 预热与自动调优时间,实际收益需测量 |
| 解码快速路径 | 通过块表直接读写分页 KV 缓存 | 每请求块表占空间,有内核与模型限制 |
| CPU 卸载 | 保留被挤出的 KV 块,恢复时避免重新 prefill | 固定页 CPU 内存和设备间传输开销 |
| 前缀共享 | 多个请求复用共同前缀的 KV 块 | 只在满足注意力条件时可用 |
CUDA graphs 与异步批处理
cb_config = ContinuousBatchingConfig(
use_cuda_graph=True,
q_padding_interval_size=64,
kv_padding_interval_size=16384,
use_async_batching=True,
)
管理器将 query 与 KV 长度填充到固定间隔,增加形状重复的机会。间隔越小,填充的浪费可能越少,但不同形状增多,需要录制和存储更多图。应同时考虑图缓存、输入缓冲和 KV 缓存,而不是只观察一项开销。
编译级别
default_compile_level 为没有显式配置的 varlen 与 decode 路径提供默认编译策略:
| 级别 | mode | dynamic | 含义 |
|---|---|---|---|
| 0 | — | — | 默认,不编译,启动快 |
| 1 | default | True | 较轻优化,较短预热 |
| 2 | max-autotune-no-cudagraphs | True | 更多调优,预热增加 |
| 3 | max-autotune-no-cudagraphs | False | 更积极的固定形状优化,预热更长 |
如 ContinuousBatchingConfig(default_compile_level=1)。显式 varlen_compile_config 与 decode_compile_config 优先。原文说明 FlashAttention 的 varlen 路径因序列长度参数容易引发重复编译而跳过编译,所以该级别在此情形主要影响 decode。短测试可从低级别开始,长时间服务再评估预热是否能被摊薄;表述为预期权衡,并非本稿的性能测试结果。
解码快速路径
当一个批次只有 decode 请求,即每个序列仅一个 query token,管理器可使用 flash_attn_with_kvcache,通过块表原位更新分页 KV 缓存。max_blocks_per_request 决定每请求块表尺寸;管理器知道 max_prompt_length 和 max_generated_length 时可推断尺寸,否则采用后备 32 块。可显式设为 64;设为 0 则关闭快速路径。
原文列出的支持组合是 CUDA 配合 FlashAttention 2/3,以及 XPU 配合相应 FlashAttention 2 内核。不可用时回退到 varlen 路径;显式指定块表大小却无法使用时会记录警告。滑动窗口层不支持这套块表,因此会强制关闭快速路径,即使手动设成非零也会被覆盖。
CPU 卸载、分布式超时与前缀缓存
GPU KV 缓存满时,CPU 卸载把被驱逐的块复制到预分配的固定页缓冲区。GPU 重新有空间后再复制回来并继续请求,避免重算整个 prompt 和已生成部分。cpu_offload_space 单位为 GiB,默认 0.0 关闭;例如设为 8.0。安装 psutil 时,默认安全阈值 0.8 将请求空间限制为可用系统内存的 80%;设置为 None 则由阈值推断大小。CPU 内存也不是无限的。
张量并行下,管理器会建立 CPU 通信组,协调提交、取消和关闭。cpu_group_timeout 限制 collective 最长等待时间;例如 600.0 秒。某一 rank 停滞时,这让其他 rank 不致永远等待。可按稀疏通信的工作负载适当加长,设为 None 则禁用超时,也意味着失去该退出边界。
allow_block_sharing=True 默认启用前缀缓存。拥有相同系统提示的请求可复用共同前缀块,减少重复 prefill。该机制要求全部模型层使用 full attention,滑动窗口模型会自动关闭。
五、注意力后端:先核对安装版本
这里存在一个需要显式保留的源文差异:读取的在线页面使用 "paged|flash_attention_2"、"paged|sdpa" 与 "paged|eager",并说明普通后端进入连续批处理时自动加前缀;显式指定 paged 的 eager/SDPA 可阻止自动替换为 FlashAttention。
同日读取的 GitHub main 文档文件 已改为普通 "flash_attention_2" 和 "sdpa",仍保留 "paged|eager";并通过 ContinuousBatchingConfig(auto_switch_to_flash=False) 阻止自动切换。这个差异说明页面与开发分支并非同一版本快照。本文主示例采用两者共同支持的 FlashAttention 写法;其他后端请按安装版本的文档选择,不能把两个快照的控制方式混装。
FlashAttention 需要 flash-attn 包或受支持 kernels 内核;原生 SDPA 不需要额外的 FlashAttention 包。自动切换的动机是减少注意力 mask 等开销,仍需按模型结构和平台核验兼容性。
六、张量并行与滑动窗口
单卡放不下模型时,可使用张量并行分片权重。原文示例指定四个设备,分页 KV 缓存也会按模型上的并行配置分片。以下保留其结构,注意它需要支持的模型架构与实际四卡环境。
from transformers import AutoModelForCausalLM, AutoTokenizer, DistributedConfig
from transformers.generation import GenerationConfig
distributed_config = DistributedConfig(tp_size=4)
model = AutoModelForCausalLM.from_pretrained(
"Qwen/Qwen3-32B",
attn_implementation="flash_attention_2",
distributed_config=distributed_config,
)
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-32B")
inputs = [tokenizer.encode(p) for p in ["What's up?", "Name a cat breed."]]
generation_config = GenerationConfig(
max_new_tokens=64, eos_token_id=tokenizer.eos_token_id
)
outputs = model.generate_batch(
inputs=inputs, generation_config=generation_config
)
torchrun --nproc-per-node 4 cb_tp.py
并行规模必须整除模型的 num_key_value_heads,否则分页缓存会在启动时报错。不要同时指定 device_map 和 distributed_config:前者把完整模块放到指定设备,后者在设备间分片同一组参数,语义冲突。
Mistral、Gemma 2 等含滑动窗口注意力的模型也可使用连续批处理。原教程展示了在加载前改 config.sliding_window = 4096 的实验方式。下面按 main 的 SDPA 写法补全导入与关闭自动切换;若安装版本仍采用在线页的旧写法,应将后端设为 "paged|sdpa",不要照搬不支持的配置字段。
import torch
from transformers import AutoConfig, AutoModelForCausalLM
from transformers.generation import ContinuousBatchingConfig
config = AutoConfig.from_pretrained("google/gemma-2-2b")
config.sliding_window = 4096
model = AutoModelForCausalLM.from_pretrained(
"google/gemma-2-2b",
config=config,
attn_implementation="sdpa",
device_map="auto",
dtype=torch.bfloat16,
)
cb_config = ContinuousBatchingConfig(auto_switch_to_flash=False)
# 调用 generate_batch 或初始化管理器时传入 cb_config。
这属于自定义实验配置,改变窗口大小不等于已经验证模型质量。启用滑动窗口时,前缀缓存和 decode 快速路径都会自动关闭。
使用前的实际检查
静态审查没有发现示例中硬编码访问令牌或 shell 注入;但模型加载会下载外部内容,需锁定来源与版本,避免为不可信模型随意开启远程代码信任。请求入口应设置输入长度、生成长度、排队数和超时边界,取消操作必须限定到请求所有者;否则队列可耗尽资源,或出现跨请求干扰。日志、前缀缓存和返回结果也应按数据敏感性管理。
原文进一步阅读包括 连续批处理介绍与架构文档。它们可帮助理解 chunked prefill、KV 管理和调度,但本文没有借用其中的基准数字作自身测试结论。
本文基于 Apache 2.0 文档作中文翻译与代码整理,保留原始归属并标明改动;许可证见 官方 LICENSE及文末完整文本。
原始版权与许可
Copyright 2018- The Hugging Face team. All rights reserved.
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding those notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
To apply the Apache License to your work, attach the following
boilerplate notice, with the fields enclosed by brackets "[]"
replaced with your own identifying information. (Don't include
the brackets!) The text should be enclosed in the appropriate
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright [yyyy] [name of copyright owner]
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.












暂无评论内容