稀疏编码器训练总览:SPLADE、静态查询路由与 CSR

原文作者:Sentence Transformers 文档贡献者(页面无个人署名);项目版权声明含 Nils Reimers。中文翻译与技术整理:未完纪。核验日期:2026-10-05。

稀疏编码器训练流程:选择 SPLADE、静态查询路由或 CSR,组合排序损失与稀疏正则,联合观察检索质量与激活维度。
未完纪编辑整理;根据已核验来源自绘,非源站截图

稀疏编码器训练的目标不只是把相似文本的分数拉近,还要控制向量中有多少维被激活。非零维度太多,索引会更大、检索会更慢;压得过稀,又可能损失召回和排序质量。本文完整整理 Sentence Transformers 的 Sparse Encoder Training Overview,覆盖 SPLADE、查询侧免 Transformer 推理的 SPLADE、CSR、数据与损失、评估、Trainer、回调及多数据集训练。

本文基于 2026-10-05 读取的官方文档及对应 Markdown 快照。示例中所有训练、模型下载、数据下载和评估均未执行,不含本站实测分数。模型与数据集的版本、许可证和资源需求,应由实际训练者另行固定与核对。

为什么还需要微调

“相似”取决于任务。以“苹果发布新 iPad”和“英伟达准备推出下一代 GPU”两条新闻为例:如果目标是把新闻分为经济、体育、科技、政治等类别,两者都属于科技,表示应较接近;如果目标是语义文本相似度,两句话讲的是不同事件,表示应有明显差异;如果目标是语义搜索,则核心关系是查询与文档,而非两个文档彼此是否相似。

因此,微调往往能改善特定任务上的表现,但改善幅度不能预先保证。官方页还提供 Training Examples 的真实应用脚本。页首提到,使用 AI 编码代理时可安装 train-sentence-transformers skill;这是可选工具提示。安装第三方技能前也应审查来源和权限。

# 原文可选工具提示,方括号表示可选参数,不是字面命令
hf skills add train-sentence-transformers [--claude] [--global]

训练由哪些组件组成

一个稀疏编码器训练流程有四到六个组件:模型、数据集、损失函数、Trainer 是基本组成;训练参数和 Evaluator 可选,但通常值得明确设置。训练参数决定学习率、批大小、保存和评估节奏;Evaluator 用可解释的任务指标补充验证损失。

模型一:传统 SPLADE

SparseEncoder 由一系列通用模块、稀疏编码器专用模块或自定义模块组成。如果已有包含 modules.json 的 SparseEncoder 检查点,可以直接继续微调,无需重新拼装其模块。例如:

from sentence_transformers import SparseEncoder
model = SparseEncoder("naver/splade-cocondenser-ensembledistil")

从预训练掩码语言模型开始时,SPLADE 通常由配置 transformer_task="fill-mask" 的 Transformer 和 SpladePooling 组成。BERT、RoBERTa、DistilBERT、ModernBERT 等 MLM 的输出经池化,生成维度与词表大小相同的稀疏向量;每个维度可对应词表中的一个 token。原文显式构造方式如下,内存允许时偏好用 FP32 加载训练模型:

from sentence_transformers import SparseEncoder
from sentence_transformers.sparse_encoder.modules import Transformer, SpladePooling

mlm_transformer = Transformer(
    "google-bert/bert-base-uncased",
    transformer_task="fill-mask",
    model_kwargs={"torch_dtype": "float32"},
)
splade_pooling = SpladePooling(pooling_strategy="max")
model = SparseEncoder(modules=[mlm_transformer, splade_pooling])

给 SparseEncoder 提供 fill-mask 模型结构时,这是默认构造方式,因此也可简写为 SparseEncoder("google-bert/bert-base-uncased", model_kwargs={"torch_dtype": "float32"})。不要把同一个名字的基础 Transformer、带 MLM head 的结构和已训练的稀疏检索模型当作完全等价的检查点。

模型二:查询侧免 Transformer 推理的 SPLADE

Inference-free SPLADE 使用 Router,让查询和文档走不同路径。文档仍经过 MLM Transformer 与 SpladePooling;查询经过 SparseStaticEmbedding,根据每个查询 token 取预计算分值。这些静态权重可训练,可理解为轻量的线性权重。复杂计算主要移到可离线完成的文档索引阶段,因此适合对查询延迟敏感的搜索场景。

这里的 inference-free 指查询不必做完整 Transformer 前向推理,不是完全没有分词、查表、聚合和检索计算。原文结构示例如下:

from sentence_transformers import SparseEncoder
from sentence_transformers.sparse_encoder.modules import (
    Router, SparseStaticEmbedding, SpladePooling, Transformer,
)

doc_encoder = Transformer(
    "google-bert/bert-base-uncased", transformer_task="fill-mask"
)
router = Router.for_query_document(
    query_modules=[
        SparseStaticEmbedding(tokenizer=doc_encoder.tokenizer, frozen=False)
    ],
    document_modules=[doc_encoder, SpladePooling("max")],
)
model = SparseEncoder(modules=[router], similarity_fn_name="dot")

训练带 Router 的模型,必须用 SparseEncoderTrainingArguments.router_mapping 把数据集列名送入正确路径。数据列若叫 question、answer,就分别映射到 query、document。这个要求与普通损失主要依赖输入列顺序的规则同时存在,不能混淆。

router_mapping={"question": "query", "answer": "document"}

SparseStaticEmbedding 通常需要比其他模块更高的学习率。原文用基础学习率 2e-5,并通过 learning_rate_mapping 把静态嵌入模块提高到 1e-3。实际匹配表达式必须覆盖模型中正确的参数名称;后面的端到端路由示例使用 r"SparseStaticEmbedding\.weight"。模块或版本改动后,应重新核对参数映射,不能只看数值已经写进配置。

模型三:CSR 稀疏表示

Contrastive Sparse Representation(CSR)在已经训练好的稠密 Sentence Transformer 上叠加 SparseAutoEncoder。常见结构是 Transformer → Pooling → SparseAutoEncoder,而不是把 MLM 词表输出直接池化。原文展示的从模块构建方式如下:

from sentence_transformers import SparseEncoder
from sentence_transformers.sparse_encoder.modules import SparseAutoEncoder, Transformer
from sentence_transformers.sentence_transformer.modules import Pooling

transformer = Transformer(
    "google-bert/bert-base-uncased",
    model_kwargs={"torch_dtype": "float32"},
)
pooling = Pooling(transformer.get_embedding_dimension(), pooling_mode="mean")
sae = SparseAutoEncoder(
    input_dim=transformer.get_embedding_dimension(),
    hidden_dim=4 * transformer.get_embedding_dimension(),
    k=256,
    k_aux=512,
)
model = SparseEncoder(modules=[transformer, pooling, sae])

k 是保留的 top 值数量,k_aux 用于辅助损失。给定一个稠密 Sentence Transformer 或非 MLM Transformer,也可用 SparseEncoder 自动初始化 CSR,例如加载 mixedbread-ai/mxbai-embed-large-v1。原文展示的该模型结构为 1024 维稠密表示、4096 维隐藏层、k=256、k_aux=512;这些是该例的结构,不是所有模型的固定尺寸。

与 SPLADE 不同,CSR 的稀疏维度不等于基础模型的词表大小,不能把每个激活维度直接解释成一个词。原文认为 CSR 更适合 1024–4096 等高维稠密表示,但本文未验证某个具体模型的效果。

准备数据:格式、标签与列顺序

SparseEncoderTrainer 接受 datasets.Dataset,或用于多个数据集的 DatasetDict。Hugging Face Hub 数据通过 load_dataset 加载;部分数据集需要指定子集。例如 all-nli 有 pair、pair-class、pair-score、triplet 四种格式:

from datasets import load_dataset
train_dataset = load_dataset(
    "sentence-transformers/all-nli", "triplet", split="train"
)
eval_dataset = load_dataset(
    "sentence-transformers/all-nli", "triplet", split="dev"
)

原文此例显示 anchor、positive、negative 三列及 557850 条训练行。这是文档中的数据展示,并非本次实际下载后统计;上游数据变化时应重新确认。Hub 上带 sentence-transformers 标签的数据集可作为寻找兼容训练材料的入口,但标签不能替代许可证与数据质量检查。

本地 CSV、JSON 等格式同样可由 datasets 加载;需要预处理的数据,可先生成等长列表,再由 Dataset.from_dict 构造。每个字典键会成为一个列:

from datasets import Dataset, load_dataset

csv_dataset = load_dataset("csv", data_files="my_file.csv")
json_dataset = load_dataset("json", data_files="my_file.json")

# 以下两个列表应由已清洗并核验的数据生成
anchors = []
positives = []
dataset = Dataset.from_dict({"anchor": anchors, "positive": positives})

空列表示例仅展示构造接口,不能直接当作有效训练集。数据格式必须匹配损失函数:需要标签的损失,必须有名为 label 或 score 的列,该列会自动成为监督标签。其余列都视为模型输入;在普通输入匹配中列名不重要,列的数量与顺序才重要。

例如 text1、text2、label 三列,label 是 0–1 的浮点相似度分数,就符合 SparseCoSENTLoss、SparseAnglELoss、SparseCosineSimilarityLoss 的两段输入加标签要求。三元组列若依次是 good_answer、bad_answer、question,Trainer 会把 good_answer 错当 anchor、bad_answer 当 positive、question 当 negative;即使代码能运行,训练语义也错了。

dataset = dataset.select_columns(["question", "good_answer", "bad_answer"])

上面是针对该三元组例子的重排示意。sample_id、metadata、source、type 等无关列要用 remove_columns 删除,或用 select_columns 只保留需要的列,否则它们也会被当成输入。路由模型还要再核对这些列对应 query/document 路径。

损失函数:排序目标外,还需要稀疏约束

损失度量模型在一批数据上的表现,优化器据此更新权重。没有适合所有任务的单一最佳损失,选择取决于数据形式和目标任务。训练 SparseEncoder 时,通常必须用 SpladeLoss 或 CSRLoss 包裹主损失,为任务损失加上稀疏正则。例外是 SparseMSELoss:它做嵌入级蒸馏,直接学习教师的稀疏表示,可单独使用。

from sentence_transformers.sparse_encoder.losses import (
    SpladeLoss, SparseMultipleNegativesRankingLoss,
)
loss = SpladeLoss(
    model=model,
    loss=SparseMultipleNegativesRankingLoss(model=model),
    query_regularizer_weight=5e-5,
    document_regularizer_weight=3e-5,
)

此主损失可用相关文本对或三元组,并利用批内负例。原文配套使用 sentence-transformers/natural-questions,显示 query、answer 两列,100231 行;同样是文档展示数据,不是本文实测。查询与文档的正则权重分别控制两侧稀疏约束,不能只盯一个总体损失值判断模型是否适合上线。

训练参数:性能和观察分别设置

用途 关键参数
优化与时长 learning_rate、lr_scheduler_type、warmup_steps、num_train_epochs、max_steps、optim。
批次与内存 per_device_train_batch_size、per_device_eval_batch_size、auto_find_batch_size、gradient_accumulation_steps、gradient_checkpointing、eval_accumulation_steps。
精度 fp16、bf16;取决于设备支持及数值稳定性,不能无条件照抄。
最佳模型 load_best_model_at_end、metric_for_best_model。
采样和路由 batch_sampler、multi_dataset_batch_sampler、router_mapping、learning_rate_mapping。
评估与保存 eval_strategy、eval_steps、save_strategy、save_steps、save_total_limit。
日志与追踪 report_to、run_name、log_level、logging_steps。
Hub 发布 push_to_hub、hub_model_id、hub_strategy、hub_private_repo;属于外部发布配置,应单独决定。

output_dir 是必要参数。原文常规 SPLADE 示例采用 1 个 epoch、每设备训练/验证批大小 16、学习率 2e-5、warmup_steps=0.1、fp16=True、bf16=False,并用 BatchSamplers.NO_DUPLICATES 避免批内重复样本影响负例。eval_steps 与 save_steps 为 0.1,logging_steps 为 0.01,保留两个检查点。这里保留原参数写法;浮点步数的比例语义和可用参数应以所安装版本的训练参数实现为准。

原文说明安装 wandb 后 run_name 会用于 W&B 记录。编辑版端到端示例显式 report_to="none"、push_to_hub=False,并移除末尾自动上传调用,避免把试验元数据或模型意外发送到外部服务。我们也把默认示例 fp16 改为 False;训练者确认设备支持后再启用。上述改动影响日志和运行条件,没有替代任务本身的数据与训练审核。

评估:验证损失之外,也看检索与稀疏度

eval_dataset 用于计算验证损失;Evaluator 提供更直观的任务指标。两者可以同时使用、只选其一,也可以都不提供,运行节奏由 eval_strategy 与 eval_steps 决定。可用 Evaluator 及其数据要求如下:

Evaluator 需要的数据
SparseBinaryClassificationEvaluator 带类别标签的文本对。
SparseEmbeddingSimilarityEvaluator 带相似度分数的文本对。
SparseInformationRetrievalEvaluator qid→query、cid→document,以及 qid→相关 cid 集合。
SparseNanoBEIREvaluator 无需手工提供数据;内部从 Hub 获取相应基准,不代表离线无数据。
SparseMSEEvaluator 教师编码的源文本及学生编码的目标文本;两组可以相同。
SparseRerankingEvaluator 由 query、positive 列表、negative 列表组成的字典列表。
SparseTranslationEvaluator 两种语言的配对句子。
SparseTripletEvaluator anchor、positive、negative 三元组。

多个评估器可由 SequentialEvaluator 合成后传给 Trainer。没有自己整理的评估数据时,可用 SparseNanoBEIREvaluator() 加载公共小型基准;频繁评估仍有计算、内存和下载成本。

原文 STSb 示例使用余弦相似度,与多数稀疏检索采用点积并不冲突:应按评估任务说明选择。完整构造如下:

from datasets import load_dataset
from sentence_transformers.sentence_transformer.evaluation import SimilarityFunction
from sentence_transformers.sparse_encoder.evaluation import (
    SparseEmbeddingSimilarityEvaluator, SparseTripletEvaluator,
)

stsb = load_dataset("sentence-transformers/stsb", split="validation")
sts_evaluator = SparseEmbeddingSimilarityEvaluator(
    sentences1=stsb["sentence1"],
    sentences2=stsb["sentence2"],
    scores=stsb["score"],
    main_similarity=SimilarityFunction.COSINE,
    name="sts-dev",
)

triplets = load_dataset(
    "sentence-transformers/all-nli", "triplet", split="dev[:1000]"
)
triplet_evaluator = SparseTripletEvaluator(
    anchors=triplets["anchor"],
    positives=triplets["positive"],
    negatives=triplets["negative"],
    main_distance_function=SimilarityFunction.DOT,
    name="all-nli-dev",
)

如果 eval_steps 很小,文档建议把训练期验证集做小,以降低评估开销;例如 90% 训练、1% 中间验证、9% 最终测试是一种可考虑的划分。训练后,trainer.evaluate(test_dataset) 给出测试损失;test_evaluator(model) 提供具体测试指标。如果先做最终评估、再保存模型,自动生成的 model card 会包含相应测试结果。分布式训练时,Evaluator 只在第一台设备上运行,而训练集与验证集由多设备分担,资源估算不能直接按设备数均摊。

端到端示例:标准 SPLADE

下面把组件连成一个完整训练脚本。它仍会下载公共模型和数据、训练并写入本地模型目录;仅仅没有自动外传日志或推送模型,不代表无需资源与数据审核。训练前应在独立环境固定版本和输入,确认磁盘、内存、GPU 与网络条件。本文没有执行它。

import logging
from datasets import load_dataset
from sentence_transformers import (
    SparseEncoder, SparseEncoderModelCardData,
    SparseEncoderTrainer, SparseEncoderTrainingArguments,
)
from sentence_transformers.sparse_encoder.evaluation import SparseNanoBEIREvaluator
from sentence_transformers.sparse_encoder.losses import (
    SparseMultipleNegativesRankingLoss, SpladeLoss,
)
from sentence_transformers.sentence_transformer.training_args import BatchSamplers

logging.basicConfig(
    format="%(asctime)s - %(message)s",
    datefmt="%Y-%m-%d %H:%M:%S",
    level=logging.INFO,
)

model = SparseEncoder(
    "distilbert/distilbert-base-uncased",
    model_card_data=SparseEncoderModelCardData(
        language="en", license="apache-2.0",
        model_name="DistilBERT base trained on Natural-Questions tuples",
    ),
    model_kwargs={"torch_dtype": "float32"},
    similarity_fn_name="dot",  # 编辑版显式指定检索点积
)

full_dataset = load_dataset(
    "sentence-transformers/natural-questions", split="train"
).select(range(100_000))
dataset_dict = full_dataset.train_test_split(test_size=1_000, seed=12)
train_dataset = dataset_dict["train"]
eval_dataset = dataset_dict["test"]

loss = SpladeLoss(
    model=model,
    loss=SparseMultipleNegativesRankingLoss(model=model),
    query_regularizer_weight=5e-5,
    document_regularizer_weight=3e-5,
)

run_name = "splade-distilbert-base-uncased-nq"
args = SparseEncoderTrainingArguments(
    output_dir=f"models/{run_name}",
    num_train_epochs=1,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=16,
    learning_rate=2e-5,
    warmup_steps=0.1,
    fp16=False,  # 编辑版:确认硬件支持后再启用;原文为 True
    bf16=False,
    batch_sampler=BatchSamplers.NO_DUPLICATES,
    eval_strategy="steps",
    eval_steps=0.1,
    save_strategy="steps",
    save_steps=0.1,
    save_total_limit=2,
    logging_steps=0.01,
    run_name=run_name,
    report_to="none",  # 编辑版:不自动连接日志服务
    push_to_hub=False,  # 编辑版:不自动发布模型
)

dev_evaluator = SparseNanoBEIREvaluator(
    dataset_names=["msmarco", "nfcorpus", "nq"], batch_size=16
)
# 不在未经稀疏训练的基础模型上预先跑完整评估
trainer = SparseEncoderTrainer(
    model=model, args=args, train_dataset=train_dataset,
    eval_dataset=eval_dataset, loss=loss, evaluator=dev_evaluator,
)
trainer.train()
dev_evaluator(model)
model.save_pretrained(f"models/{run_name}/final")
# 原文最后的 model.push_to_hub(run_name) 已移除

脚本中的 model_card_data 记录语言、名称与许可证元数据。填入 apache-2.0 并不会自动为底座模型、数据集或派生产物授予该许可;发布前仍需核对来源许可。select(range(100_000)) 也假设当前数据至少有 100000 行,换数据源后要确认长度。这个例子把名为 test 的拆分当中间 eval_dataset 使用,因此它不是严格独立、从未参与调参的最终测试集;如需可信最终指标,应另留测试集。这两点是编辑补充。

端到端变体:查询侧静态嵌入

原文第二个完整脚本保留相同 Natural Questions 数据拆分、训练步数、评估集和保存流程,但替换模型、正则权重及路由学习率。为避免两份脚本的公共部分失去同步,本文把准确差异列成可替换的代码块;按下述三处替换上一脚本,公共的 imports、数据加载、Evaluator、Trainer、训练与保存保持一致。先额外导入对应模块并替换 model 构造:

from sentence_transformers.sparse_encoder.modules import (
    Router, SparseStaticEmbedding, SpladePooling, Transformer,
)

mlm_transformer = Transformer(
    "distilbert/distilbert-base-uncased",
    transformer_task="fill-mask",
    tokenizer_args={"model_max_length": 512},
    model_kwargs={"torch_dtype": "float32"},
)
splade_pooling = SpladePooling(
    pooling_strategy="max",
    embedding_dimension=mlm_transformer.get_embedding_dimension(),
)
router = Router.for_query_document(
    query_modules=[
        SparseStaticEmbedding(tokenizer=mlm_transformer.tokenizer, frozen=False)
    ],
    document_modules=[mlm_transformer, splade_pooling],
)
model = SparseEncoder(
    modules=[router],
    similarity_fn_name="dot",
    model_card_data=SparseEncoderModelCardData(
        language="en", license="apache-2.0",
        model_name="Inference-free SPLADE trained on Natural-Questions tuples",
    ),
)

再把 loss 改成查询正则为 0、文档正则为 3e-4;查询侧已是静态词权重,不照搬传统 SPLADE 的同一组正则值:

loss = SpladeLoss(
    model=model,
    loss=SparseMultipleNegativesRankingLoss(model=model),
    query_regularizer_weight=0,
    document_regularizer_weight=3e-4,
)

最后替换训练参数块。列 query 走查询路径,answer 走文档路径;静态 embedding 参数使用 1e-3,其他参数仍用 2e-5。与前例相同,编辑版关闭默认混合精度与外部日志/上传,数值选择本身没有经过本次训练验证。

run_name = "inference-free-splade-distilbert-base-uncased-nq"
args = SparseEncoderTrainingArguments(
    output_dir=f"models/{run_name}",
    num_train_epochs=1,
    per_device_train_batch_size=16,
    per_device_eval_batch_size=16,
    learning_rate=2e-5,
    learning_rate_mapping={r"SparseStaticEmbedding\.weight": 1e-3},
    warmup_steps=0.1,
    fp16=False,
    bf16=False,
    batch_sampler=BatchSamplers.NO_DUPLICATES,
    router_mapping={"query": "query", "answer": "document"},
    eval_strategy="steps",
    eval_steps=0.1,
    save_strategy="steps",
    save_steps=0.1,
    save_total_limit=2,
    logging_steps=0.01,
    run_name=run_name,
    report_to="none",
    push_to_hub=False,
)

原文在第二个例子中还打印了训练集和第一条样本。本文省略这两条打印,避免读者换成私有语料后把原始文本直接送进日志。需要检查列名、样本数或脱敏样本时,应在受控环境中有选择地查看。

回调与多数据集训练

Trainer 支持多种 transformers.TrainerCallback 子类。SpladeRegularizerWeightSchedulerCallback 可在训练中调度 SpladeLoss 的 lambda 正则参数;安装 wandb 时,WandbCallback 可自动记录到 W&B;有 TensorBoard 时,TensorBoardCallback 可记录训练指标;安装 codecarbon 时,CodeCarbonCallback 可估算训练碳排放,并将其写入自动模型卡。这些集成各有数据路径和环境条件,不能因为代码没有显式写上传语句就假定不会记录到服务。

多数据集训练不要求先把所有数据转成完全相同格式。向 train_dataset(以及可选 eval_dataset)传入 Dataset 字典或 DatasetDict;若各数据集使用不同损失,则传入以相同数据集名称为键的 loss 字典。每个训练或验证 batch 只来自其中一个数据集。不同数据集之间按 multi_dataset_batch_sampler 的策略选择 batch:

策略 取样行为
MultiDatasetBatchSamplers.ROUND_ROBIN 轮流从各数据集取批次,直到其中一个耗尽;各数据集被平等取样,但较大数据集很可能不能全部使用。
MultiDatasetBatchSamplers.PROPORTIONAL(默认) 按数据集大小比例取样;所有样本都得到使用,较大数据集出现得更频繁。

原文指出高表现模型常同时使用多个数据集,但多源本身不保证更强。每份数据的格式、列顺序、标签、路由与损失映射都要分别匹配,才能避免把多个错误组合在同一次训练里。

同时调检索质量、稀疏度与运行成本

评价稀疏编码器不能只看任务得分。低稀疏度意味着大量维度非零,存储和检索成本可能很高。各 Evaluator 会输出 active_dims 和 sparsity_ratio,可以与检索、相似度等指标一起观察。SPLADE 的 query_regularizer_weight、document_regularizer_weight,以及 CSRLoss 的 beta、gamma,应在质量与稀疏程度之间调节。

不要在尚未训练的基础模型上先跑全量 Evaluator:其输出可能非常不稀疏,内存占用会出乎意料地高。前述脚本只构造评估器,没有在 trainer.train() 前实际调用;注释“评估基线”不应被误读成必须执行的步骤。

原文还指出,更强的稀疏编码器多依靠来自更强教师模型(例如 CrossEncoder)的蒸馏,而不只是文本对或三元组直接训练。SPLADE-v3 论文给出了 SparseDistillKLDivLoss 和 SparseMarginMSELoss 的相关思路。这一观察不意味着更换一个损失就能复制论文成绩,需要连同教师、数据与训练条件一起评估。

多数稠密表示常以余弦相似度训练和检索;SparseEncoder 则通常使用点积。某些损失允许指定相似度函数,可传 model.similarity 或 model.similarity_pairwise,但必须与训练目标和索引检索方式一致。本文明确保留 STSb 例子中的余弦,标准检索与路由示例显式采用点积。

官方页最后链接到 Training and Finetuning Sparse Embedding Models 的端到端文章,用真实任务训练并与基础检查点及其他稀疏检索模型比较。该链接是进一步案例,本文没有把其中未读取的性能数据引入结论。

本次静态审查发现与未测试部分

实际影响包括模型/数据网络下载、本地检查点写入、可能较大的 CPU/GPU/内存消耗,以及原文可自动接入的日志服务和 push_to_hub 上传。没有发现硬编码秘密、执行用户输入的 eval/exec 或 shell 拼接示例;数据内容也未被当作可执行指令。代码没有启用 trust_remote_code,但仍应只使用经过审阅的模型、数据及固定依赖。

本稿对每段代码只进行了静态阅读和语法层面的格式检查,没有训练模型、访问训练凭证、生成基准结果或执行原文命令。report_to="none" 与移除上传调用是明确的编辑变更,不应被误称为原文默认行为。静态审核没有发现其他问题不代表依赖或运行时没有漏洞。

来源、署名与许可

来源项目采用 Apache License 2.0,Copyright 2019 Nils Reimers;完整 LICENSE 保留在本文下方。本文为中文翻译整理,代码差异已标注;模型/数据各自许可证并不由此自动覆盖。

保留作者与适用许可证。本文调整与安全补充均已说明,未执行原文应用代码或命令。

Apache License 2.0

                                Apache License
                           Version 2.0, January 2004
                        http://www.apache.org/licenses/

   TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION

   1. Definitions.

      "License" shall mean the terms and conditions for use, reproduction,
      and distribution as defined by Sections 1 through 9 of this document.

      "Licensor" shall mean the copyright owner or entity authorized by
      the copyright owner that is granting the License.

      "Legal Entity" shall mean the union of the acting entity and all
      other entities that control, are controlled by, or are under common
      control with that entity. For the purposes of this definition,
      "control" means (i) the power, direct or indirect, to cause the
      direction or management of such entity, whether by contract or
      otherwise, or (ii) ownership of fifty percent (50%) or more of the
      outstanding shares, or (iii) beneficial ownership of such entity.

      "You" (or "Your") shall mean an individual or Legal Entity
      exercising permissions granted by this License.

      "Source" form shall mean the preferred form for making modifications,
      including but not limited to software source code, documentation
      source, and configuration files.

      "Object" form shall mean any form resulting from mechanical
      transformation or translation of a Source form, including but
      not limited to compiled object code, generated documentation,
      and conversions to other media types.

      "Work" shall mean the work of authorship, whether in Source or
      Object form, made available under the License, as indicated by a
      copyright notice that is included in or attached to the work
      (an example is provided in the Appendix below).

      "Derivative Works" shall mean any work, whether in Source or Object
      form, that is based on (or derived from) the Work and for which the
      editorial revisions, annotations, elaborations, or other modifications
      represent, as a whole, an original work of authorship. For the purposes
      of this License, Derivative Works shall not include works that remain
      separable from, or merely link (or bind by name) to the interfaces of,
      the Work and Derivative Works thereof.

      "Contribution" shall mean any work of authorship, including
      the original version of the Work and any modifications or additions
      to that Work or Derivative Works thereof, that is intentionally
      submitted to Licensor for inclusion in the Work by the copyright owner
      or by an individual or Legal Entity authorized to submit on behalf of
      the copyright owner. For the purposes of this definition, "submitted"
      means any form of electronic, verbal, or written communication sent
      to the Licensor or its representatives, including but not limited to
      communication on electronic mailing lists, source code control systems,
      and issue tracking systems that are managed by, or on behalf of, the
      Licensor for the purpose of discussing and improving the Work, but
      excluding communication that is conspicuously marked or otherwise
      designated in writing by the copyright owner as "Not a Contribution."

      "Contributor" shall mean Licensor and any individual or Legal Entity
      on behalf of whom a Contribution has been received by Licensor and
      subsequently incorporated within the Work.

   2. Grant of Copyright License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      copyright license to reproduce, prepare Derivative Works of,
      publicly display, publicly perform, sublicense, and distribute the
      Work and such Derivative Works in Source or Object form.

   3. Grant of Patent License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      (except as stated in this section) patent license to make, have made,
      use, offer to sell, sell, import, and otherwise transfer the Work,
      where such license applies only to those patent claims licensable
      by such Contributor that are necessarily infringed by their
      Contribution(s) alone or by combination of their Contribution(s)
      with the Work to which such Contribution(s) was submitted. If You
      institute patent litigation against any entity (including a
      cross-claim or counterclaim in a lawsuit) alleging that the Work
      or a Contribution incorporated within the Work constitutes direct
      or contributory patent infringement, then any patent licenses
      granted to You under this License for that Work shall terminate
      as of the date such litigation is filed.

   4. Redistribution. You may reproduce and distribute copies of the
      Work or Derivative Works thereof in any medium, with or without
      modifications, and in Source or Object form, provided that You
      meet the following conditions:

      (a) You must give any other recipients of the Work or
          Derivative Works a copy of this License; and

      (b) You must cause any modified files to carry prominent notices
          stating that You changed the files; and

      (c) You must retain, in the Source form of any Derivative Works
          that You distribute, all copyright, patent, trademark, and
          attribution notices from the Source form of the Work,
          excluding those notices that do not pertain to any part of
          the Derivative Works; and

      (d) If the Work includes a "NOTICE" text file as part of its
          distribution, then any Derivative Works that You distribute must
          include a readable copy of the attribution notices contained
          within such NOTICE file, excluding those notices that do not
          pertain to any part of the Derivative Works, in at least one
          of the following places: within a NOTICE text file distributed
          as part of the Derivative Works; within the Source form or
          documentation, if provided along with the Derivative Works; or,
          within a display generated by the Derivative Works, if and
          wherever such third-party notices normally appear. The contents
          of the NOTICE file are for informational purposes only and
          do not modify the License. You may add Your own attribution
          notices within Derivative Works that You distribute, alongside
          or as an addendum to the NOTICE text from the Work, provided
          that such additional attribution notices cannot be construed
          as modifying the License.

      You may add Your own copyright statement to Your modifications and
      may provide additional or different license terms and conditions
      for use, reproduction, or distribution of Your modifications, or
      for any such Derivative Works as a whole, provided Your use,
      reproduction, and distribution of the Work otherwise complies with
      the conditions stated in this License.

   5. Submission of Contributions. Unless You explicitly state otherwise,
      any Contribution intentionally submitted for inclusion in the Work
      by You to the Licensor shall be under the terms and conditions of
      this License, without any additional terms or conditions.
      Notwithstanding the above, nothing herein shall supersede or modify
      the terms of any separate license agreement you may have executed
      with Licensor regarding such Contributions.

   6. Trademarks. This License does not grant permission to use the trade
      names, trademarks, service marks, or product names of the Licensor,
      except as required for reasonable and customary use in describing the
      origin of the Work and reproducing the content of the NOTICE file.

   7. Disclaimer of Warranty. Unless required by applicable law or
      agreed to in writing, Licensor provides the Work (and each
      Contributor provides its Contributions) on an "AS IS" BASIS,
      WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
      implied, including, without limitation, any warranties or conditions
      of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
      PARTICULAR PURPOSE. You are solely responsible for determining the
      appropriateness of using or redistributing the Work and assume any
      risks associated with Your exercise of permissions under this License.

   8. Limitation of Liability. In no event and under no legal theory,
      whether in tort (including negligence), contract, or otherwise,
      unless required by applicable law (such as deliberate and grossly
      negligent acts) or agreed to in writing, shall any Contributor be
      liable to You for damages, including any direct, indirect, special,
      incidental, or consequential damages of any character arising as a
      result of this License or out of the use or inability to use the
      Work (including but not limited to damages for loss of goodwill,
      work stoppage, computer failure or malfunction, or any and all
      other commercial damages or losses), even if such Contributor
      has been advised of the possibility of such damages.

   9. Accepting Warranty or Additional Liability. While redistributing
      the Work or Derivative Works thereof, You may choose to offer,
      and charge a fee for, acceptance of support, warranty, indemnity,
      or other liability obligations and/or rights consistent with this
      License. However, in accepting such obligations, You may act only
      on Your own behalf and on Your sole responsibility, not on behalf
      of any other Contributor, and only if You agree to indemnify,
      defend, and hold each Contributor harmless for any liability
      incurred by, or claims asserted against, such Contributor by reason
      of your accepting any such warranty or additional liability.

   END OF TERMS AND CONDITIONS

   APPENDIX: How to apply the Apache License to your work.

      To apply the Apache License to your work, attach the following
      boilerplate notice, with the fields enclosed by brackets "{}"
      replaced with your own identifying information. (Don't include
      the brackets!)  The text should be enclosed in the appropriate
      comment syntax for the file format. We also recommend that a
      file or class name and description of purpose be included on the
      same "printed page" as the copyright notice for easier
      identification within third-party archives.

   Copyright 2019 Nils Reimers

   Licensed under the Apache License, Version 2.0 (the "License");
   you may not use this file except in compliance with the License.
   You may obtain a copy of the License at

       http://www.apache.org/licenses/LICENSE-2.0

   Unless required by applicable law or agreed to in writing, software
   distributed under the License is distributed on an "AS IS" BASIS,
   WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
   See the License for the specific language governing permissions and
limitations under the License.

Sentence Transformers 上游通知

以下保留所引用上游代码的官方 NOTICE.txt,与原有完整 Apache License 2.0 一起提供;其范围不扩张到第三方模型或数据集。

-------------------------------------------------------------------------------
Sentence Transformers

Copyright 2019-2025
Ubiquitous Knowledge Processing (UKP) Lab
Technische Universität Darmstadt

Copyright 2025-present
Hugging Face, Inc.
-------------------------------------------------------------------------------
© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容