SPEED-Bench 服务端基准测试

这是一个轻量级 SPEED-Bench 客户端,通过兼容 OpenAI 的 API 对已经运行的 llama-server 进行基准测试。它主要用于评估推测解码(草稿模型、n-gram、MTP、EAGLE3 等),报告各类别的吞吐量、延迟和草稿接受率。

数据集处理方式遵循 aiperf SPEED-Bench 教程,该教程还更详细地说明了数据集布局。

安装

pip install -r tools/server/bench/speed-bench/requirements.txt

启动服务器

客户端不会启动服务器,因此请先自行启动 llama-server。如果关注吞吐量数字,请将客户端 --concurrency 设置为服务器的槽位数(--np):

llama-server \
  -m target.gguf \
  -c 8192 \
  --port 8080 \
  -ngl 99 -fa on \
  --np 1 \
  --jinja

使用推测解码时,请为你的配置选择适当的标志启动服务器,例如用 -md 指定草稿模型,或使用 --spec-type ngram-mod。详情见推测解码文档。

运行

python tools/server/bench/speed-bench/speed_bench.py \
  --url localhost:8080 \
  --bench qualitative \
  --category coding \
  --osl 1024 \
  --concurrency 1

选项

选项 默认值 说明
--url localhost:8080 服务器 URL。协议和 /v1 均可省略,也允许末尾带斜杠,因此 localhost:8080 和 http://localhost:8080/v1/ 都可以使用。
--model 无 每个请求中发送的可选 model 字段。
--bench qualitative SPEED-Bench 配置,例如 qualitative、throughput_1k。请参阅可用的数据集变体。
--category all 基准中的类别筛选器,使用逗号分隔列表或 all。对于 qualitative,类别为 coding、humanities、math、multilingual、qa、rag、reasoning、roleplay、stem、summarization、writing。对于 throughput_{ISL} 划分,类别为 high_entropy、low_entropy、mixed。
--osl 1024 输出序列长度,映射到 max_tokens。
--extra-inputs {"temperature":0} 以 JSON 对象形式提供的额外请求字段。
--concurrency 1 并发客户端请求数,通常与 --np 一致。
--limit 无 每个类别的最大样本数,适合冒烟测试。
--timeout 600 每个请求的超时时间,单位为秒。
--output 无 将原始的逐请求结果及摘要保存为 JSON。

几个常用选项:

  • --category all 运行基准中的每个类别。
  • --category coding,math 只运行这两个类别。
  • --bench throughput_8k 运行输入长度固定的吞吐量划分。
  • --limit 8 每个类别最多保留 8 个样本,足以进行快速检查。

throughput_{ISL} 划分使用固定输入长度(1k–32k),适合长上下文测试,也适合在已知长度的提示上比较不同 llama-server 批处理设置,例如遍历 -ub / --ubatch-size。确保服务器的 -c 对所选划分足够大。提高 -ub 时,也将 -b 至少提高到相同值,因为物理 ubatch 不能超过逻辑 batch。

指定 --output 时,JSON 文件包含运行 config、selected_samples / completed_samples / failed_samples 计数、各类别的 summary 行,以及逐样本的 results。

指标

摘要为每个类别输出一行,另有一行 overall:

  • samples:成功完成的样本数量。
  • avg_prompt_t/s:llama.cpp 的预填充吞吐量(timings.prompt_per_second),对类别中的样本求平均。
  • avg_pred_t/s:llama.cpp 的解码吞吐量(timings.predicted_per_second),对类别中的样本求平均。
  • avg_latency:客户端观测到的平均端到端请求延迟。
  • accept_rate:类别范围内的 accepted / draft_n;没有生成草稿时(draft_n == 0)为 n/a。

基线与推测解码的比较

用 --output 保存各服务器的一次运行结果,然后用 speed_bench_compare.py 比较两个 JSON 文件。

首先,启动普通 llama-server(不启用推测解码),并保存基线:

python tools/server/bench/speed-bench/speed_bench.py \
  --url localhost:8080 \
  --bench qualitative \
  --category all \
  --osl 1024 \
  --concurrency 1 \
  --output baseline.json

然后重启 llama-server,启用推测解码,再保存一次运行:

python tools/server/bench/speed-bench/speed_bench.py \
  --url localhost:8080 \
  --bench qualitative \
  --category all \
  --osl 1024 \
  --concurrency 1 \
  --output spec.json

最后比较两者:

python tools/server/bench/speed-bench/speed_bench_compare.py \
  --baseline baseline.json \
  --speculative spec.json

比较表会增加:

  • decode_speedup = spec_avg_pred_t/s / base_avg_pred_t/s
  • latency_speedup = base_avg_latency / spec_avg_latency

两次运行的 --bench、--category、--osl 和 --limit 应保持一致,否则它们使用的提示就不相同。

来源与许可

原文:SPEED-Bench server benchmark。作者及权利人:ggml-org / llama.cpp 项目贡献者。中文翻译基于用户提供的完整原文缓存,代码和示例输出保留原文;本文未实际执行示例。

许可依据:MIT License。

许可文本

MIT License

Copyright (c) 2023-2026 The ggml authors

Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the “Software”), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:

The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.

THE SOFTWARE IS PROVIDED “AS IS”, WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.

© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容