嵌入向量量化

原文:Sentence Transformers 文档贡献者,Embedding Quantization。本文依据 2026-10-05 读取的完整文档,并连读官方 semantic_search_recommended.py 全文作中文翻译整理。项目许可为 Apache License 2.0,保留 Copyright 2019 Nils Reimers 及 Sentence Transformers 贡献者归属;文末完整脚本仍属原作者代码,中文说明和风险注释为本次修改。

大规模向量检索的成本很大一部分来自表示本身。许多嵌入模型输出 1024 维向量,每维 float32 占 4 字节。5000 万个这样的向量仅原始数据就需要 204.8 GB(十进制),原文概括为约 200 GB,尚未计算索引和服务进程的开销。量化把每个维度转换为更小的数值表示,以减少内存、磁盘和计算需求。

这里量化的是模型输出的嵌入向量。它与把模型权重转换成 INT8/INT4 是不同层次的优化,模型仍可按原方式生成向量。原文与 mixedbread.ai 的联合介绍讨论了检索质量与性能实验;本文保留这些结果的来源身份,不把它们当成对任意模型、数据或硬件的保证。

查询先生成 float32 向量并转为二进制,从 FAISS 召回40个候选,再按候选编号读取磁盘 INT8 向量,用原始查询重评分并选出前10项。
原文推荐的两阶段检索流程,图中 40 和 10 是 top_k=10、rescore_multiplier=4 的示例。未完纪原创技术示意图。

1. 二进制量化:每个维度只保留一个位

二进制量化以 0 为阈值:归一化向量的某维大于 0 就编码为 1,否则为 0。每维从 32 位缩为 1 位,原始向量表示的存储量减少至三十二分之一。两个二进制向量可用汉明距离比较,也就是对应位置上不同位的数量;距离越小,表示越接近。

位运算让汉明距离的计算非常高效。原文用“2 个 CPU 周期”强调底层操作的速度;这不能理解为任意长度向量或一次完整检索只需两个周期,实际耗时还受向量长度、指令集、内存访问、索引以及硬件影响。

Yamada 等人(2021)提出了一个重评分步骤(论文称 rerank):先用二进制查询对二进制语料检索 rescore_multiplier × top_k 个候选,再用 float32 查询和这些二进制文档向量计算点积。也就是说,压缩后的快速第一阶段负责缩小集合,第二阶段恢复一部分排序信息。

原文报告这种重评分可保留“最高约 96%”的检索性能,表示存储减少 32 倍,检索速度最高提升 32 倍。这些是引用实验结果和表示层面的比例,不是每一份工作负载都能达到的端到端加速。

2. 在 Sentence Transformers 中得到二进制向量

1024 个位通常不会逐位单独存储,而是通过 np.packbits 打包到 128 个字节。输出数组虽然 dtype 是 int8 或 uint8,其每个元素实际承载 8 个维度的符号位,不能把形状中的 128 误认为嵌入模型改成了 128 维。

from sentence_transformers import SentenceTransformer
from sentence_transformers.quantization import quantize_embeddings

model = SentenceTransformer("mixedbread-ai/mxbai-embed-large-v1")

# 方法一:在 encode 时指定量化精度
binary_embeddings = model.encode(
    ["I am driving to the lake.", "It is a beautiful day."],
    precision="binary",
)

# 方法二:先生成普通向量,再量化
embeddings = model.encode(
    ["I am driving to the lake.", "It is a beautiful day."]
)
binary_embeddings = quantize_embeddings(embeddings, precision="binary")

原文的两条输入、1024 维示例给出的数组属性如下。这是原文展示和类型大小推导,并非本次运行输出:

表示 shape dtype nbytes
普通向量 (2, 1024) float32 8192
打包二进制 (2, 128) int8 256

若向量库要求无符号字节,使用 precision="ubinary",得到 uint8;例如下文的 FAISS 二进制索引使用 ubinary。这里 binary 与 ubinary 是打包字节的符号类型区别,不能与普通的 INT8 标量量化混淆。

3. 标量 INT8 量化与校准范围

标量量化把 float32 的连续值映射到 int8 的 256 个离散等级(-128 到 127)。通常从一批较大的校准向量中,按维计算最小值和最大值,再据此划分量化桶。每个维度占 1 字节,1024 维对应 1024 字节,原始表示比 float32 小 4 倍。也可以选择 uint8,取决于向量库支持情况。

校准数据会决定每一维的量化桶,对检索质量有直接影响。原文推荐三种输入方式:一次对较大向量集合量化;直接提供每维的 min/max 范围;或传入较大的 calibration_embeddings,让函数从中计算范围。若仅凭两条向量建立 INT8 桶,库会给出稳定性警告。

from sentence_transformers import SentenceTransformer
from sentence_transformers.quantization import quantize_embeddings
from datasets import load_dataset

model = SentenceTransformer("mixedbread-ai/mxbai-embed-large-v1")
corpus = load_dataset(
    "google-research-datasets/nq_open", split="train[:1000]"
)["question"]
calibration_embeddings = model.encode(corpus)

embeddings = model.encode(
    ["I am driving to the lake.", "It is a beautiful day."]
)
int8_embeddings = quantize_embeddings(
    embeddings,
    precision="int8",
    calibration_embeddings=calibration_embeddings,
)

在原文示例中,float32 数组是 (2,1024)、8192 字节;量化后的 int8 数组形状仍为 (2,1024),占 2048 字节。标量量化也可以搭配重评分。校准集应代表未来数据分布;不能把任意两条查询量化得到的标尺与另一个语料索引混用。

4. 把二进制速度与 INT8 排序组合起来

官方展示了一个约 4100 万条 Wikipedia 文本的组合方案:

  1. 用 mixedbread-ai/mxbai-embed-large-v1 生成 float32 查询向量。
  2. 用 quantize_embeddings 将查询转成二进制。
  3. 在约 5.2 GB 的二进制索引中检索前 40 个候选。
  4. 按这 40 个候选的编号,从磁盘上的 INT8 索引按需加载向量。
  5. 使用 float32 查询与 INT8 文档向量重评分,选出前 10 项。
  6. 按分数排序并显示文本。

原文给出的 INT8 索引磁盘体积为 47.5 GB,两个索引合计约 52 GB,并称检索常驻二进制索引约占 5.2 GB 内存;演示部分把规模近似写为 5 GB 内存和 50 GB 磁盘,相对于其约 200 GB 的 float32 对照方案。这里是原文系统及索引实现的结果,索引开销不能只用维数乘字节数来替代。

编辑补充:原文将磁盘映射 INT8 索引写为“0 bytes of memory”,推荐脚本也称 view 不消耗内存。准确理解应是无需预先把整个 INT8 索引复制到进程堆;读取的页面仍会占用操作系统页缓存或驻留内存,候选数组及重评分也需要内存。本文不把磁盘映射描述为物理内存开销为零。

量化还可与 Matryoshka Embeddings 配合,或在检索后继续用 Cross-Encoder 重排序。量化检索减少候选筛选成本,并不排斥后续更精细的相关性模型。

5. 连读官方推荐脚本

下文完整保留本次读取的官方推荐脚本,便于看清数据、索引和结果之间的关系。脚本的下载、建索引和交互循环没有在本次编辑环境中执行。其依赖包括 faiss、numpy、datasets、usearch 和 sentence-transformers,实际安装版本需要自行固定并验证。

它使用 Quora 重复问题数据,把每对 sentence1、sentence2 交错展开成 text,移除原列,然后取前 100000 条。因为语料与查询都是问题,模型统一使用 query 提示词;源码特别提醒,一般文档检索不应机械地给文档也加查询提示词。默认问题为 “How do I become a good programmer?”。

建库分支先生成归一化的完整语料向量;转换成 ubinary 后放入 1024 维 FAISS IndexBinaryFlat,并写入磁盘。再对同一批向量量化为 int8,用 USearch 的 1024 维、内积 ip、i8 类型索引保存。以 0 到 N-1 作为键,所以 FAISS 返回的编号、USearch 键和 corpus 的位置必须严格一致。最后删除创建时的 USearch 对象,再以 view=True 恢复。

search 函数按六个阶段计时:嵌入、查询量化、二进制检索、候选加载、重评分、排序。它先取 top_k × rescore_multiplier 个候选,读取对应 INT8 向量并转换为普通整数数组,再计算 query_embedding @ int8_embeddings.T,按降序保留 top_k。返回分数、语料编号及耗时字典。Total Retrieval Time 不包含 Embed Time,这一点在比较端到端延迟时尤其要注意。

脚本最后以无限循环显示耗时和结果,再读取下一条问题。其没有专门的退出命令或 EOF 处理,不宜原封不动改成服务接口。下列代码保持上游正文,注释也保留原样;中文解释位于前后段落:

"""
This script showcases a recommended approach to perform semantic search using quantized embeddings with FAISS and usearch.
In particular, it uses binary search with int8 rescoring. The binary search is highly efficient, and its index can be kept
in memory even for massive datasets: it takes (num_dimensions * num_documents / 8) bytes, i.e. 1.19GB for 10 million embeddings.
"""

import json
import os
import time
import faiss
import numpy as np
from datasets import load_dataset
from usearch.index import Index

from sentence_transformers import SentenceTransformer
from sentence_transformers.util.quantization import quantize_embeddings
# We use usearch as it can efficiently load int8 vectors from disk.

# Load the model
# NOTE: Because we are only comparing questions here, we will use the "query" prompt for everything.
# Normally you don't use this prompt for documents, but only for the queries
model = SentenceTransformer(
    "mixedbread-ai/mxbai-embed-large-v1",
    prompts={"query": "Represent this sentence for searching relevant passages: "},
    default_prompt_name="query",
)
# Load a corpus with texts
dataset = load_dataset("sentence-transformers/quora-duplicates", "pair-class", split="train").map(
    lambda batch: {"text": [question for pair in zip(batch["sentence1"], batch["sentence2"]) for question in pair]},
    batched=True,
    remove_columns=["sentence1", "sentence2", "label"],
)
max_corpus_size = 100_000
corpus = dataset["text"][:max_corpus_size]

# Apply some default query
query = "How do I become a good programmer?"
# Try to load the precomputed binary and int8 indices
if os.path.exists("quora_faiss_ubinary.index"):
    binary_index: faiss.IndexBinaryFlat = faiss.read_index_binary("quora_faiss_ubinary.index")
    int8_view = Index.restore("quora_usearch_int8.index", view=True)

else:
    # Encode the corpus using the full precision
    full_corpus_embeddings = model.encode(corpus, normalize_embeddings=True, show_progress_bar=True)
    # Convert the embeddings to "ubinary" for efficient FAISS search
    ubinary_embeddings = quantize_embeddings(full_corpus_embeddings, "ubinary")
    binary_index = faiss.IndexBinaryFlat(1024)
    binary_index.add(ubinary_embeddings)
    faiss.write_index_binary(binary_index, "quora_faiss_ubinary.index")
    # Convert the embeddings to "int8" for efficiently loading int8 indices with usearch
    int8_embeddings = quantize_embeddings(full_corpus_embeddings, "int8")
    index = Index(ndim=1024, metric="ip", dtype="i8")
    index.add(np.arange(len(int8_embeddings)), int8_embeddings)
    index.save("quora_usearch_int8.index")
    del index

    # Load the int8 index as a view, which does not cost any memory
    int8_view = Index.restore("quora_usearch_int8.index", view=True)

def search(query, top_k: int = 10, rescore_multiplier: int = 4):
    # 1. Embed the query as float32
    start_time = time.time()
    query_embedding = model.encode(query)
    embed_time = time.time() - start_time

    # 2. Quantize the query to ubinary
    start_time = time.time()
    query_embedding_ubinary = quantize_embeddings(query_embedding.reshape(1, -1), "ubinary")
    quantize_time = time.time() - start_time
    # 3. Search the binary index
    start_time = time.time()
    _scores, binary_ids = binary_index.search(query_embedding_ubinary, top_k * rescore_multiplier)
    binary_ids = binary_ids[0]
    search_time = time.time() - start_time

    # 4. Load the corresponding int8 embeddings
    start_time = time.time()
    int8_embeddings = int8_view[binary_ids].astype(int)
    load_time = time.time() - start_time
    # 5. Rescore the top_k * rescore_multiplier using the float32 query embedding and the int8 document embeddings
    start_time = time.time()
    scores = query_embedding @ int8_embeddings.T
    rescore_time = time.time() - start_time

    # 6. Sort the scores and return the top_k
    start_time = time.time()
    indices = (-scores).argsort()[:top_k]
    top_k_indices = binary_ids[indices]
    top_k_scores = scores[indices]
    sort_time = time.time() - start_time
    return (
        top_k_scores.tolist(),
        top_k_indices.tolist(),
        {
            "Embed Time": f"{embed_time:.4f} s",
            "Quantize Time": f"{quantize_time:.4f} s",
            "Search Time": f"{search_time:.4f} s",
            "Load Time": f"{load_time:.4f} s",
            "Rescore Time": f"{rescore_time:.4f} s",
            "Sort Time": f"{sort_time:.4f} s",
            "Total Retrieval Time": f"{quantize_time + search_time + load_time + rescore_time + sort_time:.4f} s",
        },
    )

while True:
    scores, indices, timings = search(query)

    # Output the results
    print(f"Timings:\n{json.dumps(timings, indent=2)}")
    print(f"Query: {query}")
    for score, index in zip(scores, indices):
        print(f"(Score: {score:.4f}) {corpus[index]}")
    print("")

    # 10. Prompt for more queries
    query = input("Please enter a question: ")

6. 这份脚本的边界和静态检查发现

导入路径存在版本差异。文档示例使用 sentence_transformers.quantization;本次读取的 main 分支推荐脚本使用 sentence_transformers.util.quantization。本文分别保留读取到的写法,没有未经验证地统一路径。部署时应固定 sentence-transformers 版本,按对应版本 API 确认导入;main 是可变分支,不能当成稳定版本声明。

1024 维是示例硬编码。FAISS 和 USearch 都使用 1024,模型换成其他维度后必须同步调整。脚本注释说 1000 万条向量约 1.19 GB;按 1024 位计算为 1.28 GB(十进制)或约 1.19 GiB,应注意单位。建立索引时还暂存 float32、二进制和 INT8 数组,峰值内存高于只读检索阶段。

索引复用检查不足。现有分支只检查 quora_faiss_ubinary.index 是否存在,随即恢复两个文件;若另一个文件缺失、写入中断,或两者来自不同的语料顺序、模型版本、提示词、维度或量化设置,就可能报错或产生错配。代码也重新加载数据集,未将 corpus 与两个索引作为同一个版本化产物保存。实际系统应把两种索引、语料 ID 映射、模型 revision、提示词、归一化选项、校准参数和文件校验值作为同一构建批次发布,核对后再开放查询,并用临时文件加原子发布避免半成品。

候选数需要边界检查。原文默认语料远多于 40 条,因此示例配置合理;把函数推广到小索引或接受外部 top_k 时,应要求正整数,将召回数量限制在索引大小以内,并处理 FAISS 未命中时可能返回的无效编号,避免将 -1 当作有效数组索引。此项是静态使用条件分析,不是运行复现。

重评分不是原始余弦相似度。脚本直接把 float32 查询与未反量化的 INT8 编码点积,正是其近似重评分方案;没有恢复每维量化尺度和偏移。因此分数不等同 float32 语料上的余弦值,也不适合直接沿用原始余弦阈值。语料调用 normalize_embeddings=True,而查询没有显式指定;同一个查询的正比例缩放不改变点积候选排序,但会改变绝对分数,不应跨查询直接当统一置信度。

来源与资源同样需要审核。脚本会从模型、数据集服务下载内容,并通过本机原生库读取索引文件。应只加载可信构建的索引和模型,固定来源版本、控制文件体积和内存预算,不把外部上传的任意索引交给服务解析。源码没有 subprocess、Shell 拼接或实际硬编码秘密;这不等于依赖或文件解析没有漏洞。由于正文是检索示例,未提供鉴权、并发控制、请求限额或服务监控,也不能据此把它当生产服务。

7. 原文提供的其他实验入口

推荐检索脚本组合二进制与 INT8,适合先理解完整数据路径。原文还提供两类使用示例:semantic_search_faiss.py 用 semantic_search_faiss 辅助函数演示 FAISS 上的二进制、标量量化及重评分;semantic_search_usearch.py 展示 USearch 对应使用方式。

另有 FAISS 基准脚本比较 float32 检索、二进制加重评分、标量加重评分,原文称 ubinary 尤其明显;USearch 基准脚本在较新硬件上显示加速,尤其是 int8。这些链接作为原文进一步实验资源保留;本稿连读并审查的是推荐脚本,没有把另外四份脚本写成已逐行审核或已实测。

本次交付只完成全文、推荐脚本、静态风险及配图核对;没有下载模型、执行建库、查询或性能测试。读者评估时,应同时比较召回率或任务指标、峰值内存、索引磁盘体积、端到端延迟和近似分数校准,而不是仅凭压缩比判断结果。

示例代码 Apache License 2.0

以下许可适用于引用的Sentence Transformers示例代码。文档与原文署名另见篇首。

Apache License
                           Version 2.0, January 2004
                        http://www.apache.org/licenses/

   TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION

   1. Definitions.

      "License" shall mean the terms and conditions for use, reproduction,
      and distribution as defined by Sections 1 through 9 of this document.

      "Licensor" shall mean the copyright owner or entity authorized by
      the copyright owner that is granting the License.
      "Legal Entity" shall mean the union of the acting entity and all
      other entities that control, are controlled by, or are under common
      control with that entity. For the purposes of this definition,
      "control" means (i) the power, direct or indirect, to cause the
      direction or management of such entity, whether by contract or
      otherwise, or (ii) ownership of fifty percent (50%) or more of the
      outstanding shares, or (iii) beneficial ownership of such entity.
      "You" (or "Your") shall mean an individual or Legal Entity
      exercising permissions granted by this License.

      "Source" form shall mean the preferred form for making modifications,
      including but not limited to software source code, documentation
      source, and configuration files.
      "Object" form shall mean any form resulting from mechanical
      transformation or translation of a Source form, including but
      not limited to compiled object code, generated documentation,
      and conversions to other media types.

      "Work" shall mean the work of authorship, whether in Source or
      Object form, made available under the License, as indicated by a
      copyright notice that is included in or attached to the work
      (an example is provided in the Appendix below).
      "Derivative Works" shall mean any work, whether in Source or Object
      form, that is based on (or derived from) the Work and for which the
      editorial revisions, annotations, elaborations, or other modifications
      represent, as a whole, an original work of authorship. For the purposes
      of this License, Derivative Works shall not include works that remain
      separable from, or merely link (or bind by name) to the interfaces of,
      the Work and Derivative Works thereof.
      "Contribution" shall mean any work of authorship, including
      the original version of the Work and any modifications or additions
      to that Work or Derivative Works thereof, that is intentionally
      submitted to Licensor for inclusion in the Work by the copyright owner
      or by an individual or Legal Entity authorized to submit on behalf of
      the copyright owner. For the purposes of this definition, "submitted"
      means any form of electronic, verbal, or written communication sent
      to the Licensor or its representatives, including but not limited to
      communication on electronic mailing lists, source code control systems,
      and issue tracking systems that are managed by, or on behalf of, the
      Licensor for the purpose of discussing and improving the Work, but
      excluding communication that is conspicuously marked or otherwise
      designated in writing by the copyright owner as "Not a Contribution."
      "Contributor" shall mean Licensor and any individual or Legal Entity
      on behalf of whom a Contribution has been received by Licensor and
      subsequently incorporated within the Work.
   2. Grant of Copyright License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      copyright license to reproduce, prepare Derivative Works of,
      publicly display, publicly perform, sublicense, and distribute the
      Work and such Derivative Works in Source or Object form.
   3. Grant of Patent License. Subject to the terms and conditions of
      this License, each Contributor hereby grants to You a perpetual,
      worldwide, non-exclusive, no-charge, royalty-free, irrevocable
      (except as stated in this section) patent license to make, have made,
      use, offer to sell, sell, import, and otherwise transfer the Work,
      where such license applies only to those patent claims licensable
      by such Contributor that are necessarily infringed by their
      Contribution(s) alone or by combination of their Contribution(s)
      with the Work to which such Contribution(s) was submitted. If You
      institute patent litigation against any entity (including a
      cross-claim or counterclaim in a lawsuit) alleging that the Work
      or a Contribution incorporated within the Work constitutes direct
      or contributory patent infringement, then any patent licenses
      granted to You under this License for that Work shall terminate
      as of the date such litigation is filed.
   4. Redistribution. You may reproduce and distribute copies of the
      Work or Derivative Works thereof in any medium, with or without
      modifications, and in Source or Object form, provided that You
      meet the following conditions:

      (a) You must give any other recipients of the Work or
          Derivative Works a copy of this License; and

      (b) You must cause any modified files to carry prominent notices
          stating that You changed the files; and
      (c) You must retain, in the Source form of any Derivative Works
          that You distribute, all copyright, patent, trademark, and
          attribution notices from the Source form of the Work,
          excluding those notices that do not pertain to any part of
          the Derivative Works; and
      (d) If the Work includes a "NOTICE" text file as part of its
          distribution, then any Derivative Works that You distribute must
          include a readable copy of the attribution notices contained
          within such NOTICE file, excluding those notices that do not
          pertain to any part of the Derivative Works, in at least one
          of the following places: within a NOTICE text file distributed
          as part of the Derivative Works; within the Source form or
          documentation, if provided along with the Derivative Works; or,
          within a display generated by the Derivative Works, if and
          wherever such third-party notices normally appear. The contents
          of the NOTICE file are for informational purposes only and
          do not modify the License. You may add Your own attribution
          notices within Derivative Works that You distribute, alongside
          or as an addendum to the NOTICE text from the Work, provided
          that such additional attribution notices cannot be construed
          as modifying the License.
      You may add Your own copyright statement to Your modifications and
      may provide additional or different license terms and conditions
      for use, reproduction, or distribution of Your modifications, or
      for any such Derivative Works as a whole, provided Your use,
      reproduction, and distribution of the Work otherwise complies with
      the conditions stated in this License.
   5. Submission of Contributions. Unless You explicitly state otherwise,
      any Contribution intentionally submitted for inclusion in the Work
      by You to the Licensor shall be under the terms and conditions of
      this License, without any additional terms or conditions.
      Notwithstanding the above, nothing herein shall supersede or modify
      the terms of any separate license agreement you may have executed
      with Licensor regarding such Contributions.
   6. Trademarks. This License does not grant permission to use the trade
      names, trademarks, service marks, or product names of the Licensor,
      except as required for reasonable and customary use in describing the
      origin of the Work and reproducing the content of the NOTICE file.
   7. Disclaimer of Warranty. Unless required by applicable law or
      agreed to in writing, Licensor provides the Work (and each
      Contributor provides its Contributions) on an "AS IS" BASIS,
      WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
      implied, including, without limitation, any warranties or conditions
      of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
      PARTICULAR PURPOSE. You are solely responsible for determining the
      appropriateness of using or redistributing the Work and assume any
      risks associated with Your exercise of permissions under this License.
   8. Limitation of Liability. In no event and under no legal theory,
      whether in tort (including negligence), contract, or otherwise,
      unless required by applicable law (such as deliberate and grossly
      negligent acts) or agreed to in writing, shall any Contributor be
      liable to You for damages, including any direct, indirect, special,
      incidental, or consequential damages of any character arising as a
      result of this License or out of the use or inability to use the
      Work (including but not limited to damages for loss of goodwill,
      work stoppage, computer failure or malfunction, or any and all
      other commercial damages or losses), even if such Contributor
      has been advised of the possibility of such damages.
   9. Accepting Warranty or Additional Liability. While redistributing
      the Work or Derivative Works thereof, You may choose to offer,
      and charge a fee for, acceptance of support, warranty, indemnity,
      or other liability obligations and/or rights consistent with this
      License. However, in accepting such obligations, You may act only
      on Your own behalf and on Your sole responsibility, not on behalf
      of any other Contributor, and only if You agree to indemnify,
      defend, and hold each Contributor harmless for any liability
      incurred by, or claims asserted against, such Contributor by reason
      of your accepting any such warranty or additional liability.
   END OF TERMS AND CONDITIONS

   APPENDIX: How to apply the Apache License to your work.
      To apply the Apache License to your work, attach the following
      boilerplate notice, with the fields enclosed by brackets "{}"
      replaced with your own identifying information. (Don't include
      the brackets!)  The text should be enclosed in the appropriate
      comment syntax for the file format. We also recommend that a
      file or class name and description of purpose be included on the
      same "printed page" as the copyright notice for easier
      identification within third-party archives.
   Copyright 2019 Nils Reimers

   Licensed under the Apache License, Version 2.0 (the "License");
   you may not use this file except in compliance with the License.
   You may obtain a copy of the License at

       http://www.apache.org/licenses/LICENSE-2.0
   Unless required by applicable law or agreed to in writing, software
   distributed under the License is distributed on an "AS IS" BASIS,
   WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
   See the License for the specific language governing permissions and
limitations under the License.

Sentence Transformers 上游通知

以下保留官方源码仓库的 NOTICE.txt,与上文完整 Apache License 2.0 一起适用于所引用的上游代码;不据此扩张第三方模型、数据集或文档的许可范围。

-------------------------------------------------------------------------------
Sentence Transformers

Copyright 2019-2025
Ubiquitous Knowledge Processing (UKP) Lab
Technische Universität Darmstadt

Copyright 2025-present
Hugging Face, Inc.
-------------------------------------------------------------------------------
© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容