作者:Liam Thompson Liam Thompson
本笔记本介绍如何把 Hugging Face 模型上传到 Elasticsearch 集群,从而在 Elasticsearch 中实现语义重排序。我们将使用 retriever 抽象:一种更简洁的 Elasticsearch 语法,用来构建查询并组合不同搜索操作。
你将学习:
- 从 Hugging Face 选择一个交叉编码器(cross-encoder)模型,执行语义重排序。
- 使用 Eland——面向 Elasticsearch 机器学习的 Python 客户端——将模型上传到 Elasticsearch 部署。 Eland
- 创建推理端点,管理 rerank 任务。
- 使用 text_similarity_rerank 检索器查询数据。
准备条件
本例需要:
- Elastic 部署。非 Serverless 部署要求版本为 8.15.0 或更高。 free trial deployment options
- 本例使用 Elastic Cloud;它提供免费试用。 free trial
- 也可以查看其他部署选项。 deployment options
- 找到部署的 Cloud ID,并创建 API 密钥。详情参见链接。 Learn more
安装并导入软件包
eland 的安装需要几分钟。
!pip install -qU elasticsearch
!pip install eland[pytorch]
from elasticsearch import Elasticsearch, helpers
初始化 Elasticsearch Python 客户端
首先连接 Elasticsearch 实例。
>>> from getpass import getpass
>>> # https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud#finding-your-cloud-id
>>> ELASTIC_CLOUD_ID = getpass("Elastic Cloud ID: ")
>>> # https://www.elastic.co/search-labs/tutorials/install-elasticsearch/elastic-cloud#creating-an-api-key
>>> ELASTIC_API_KEY = getpass("Elastic Api Key: ")
>>> # Create the client instance
>>> client = Elasticsearch(
... # For local development
... # hosts=["http://localhost:9200"]
... cloud_id=ELASTIC_CLOUD_ID,
... api_key=ELASTIC_API_KEY,
... )
Elastic Cloud ID: ··········
Elastic Api Key: ··········
测试连接
运行下面的测试,确认 Python 客户端已经连接到 Elasticsearch 实例。
print(client.info())
本例使用一个小型电影数据集。
>>> from urllib.request import urlopen
>>> import json
>>> import time
>>> url = "https://huggingface.co/datasets/leemthompo/small-movies/raw/main/small-movies.json"
>>> response = urlopen(url)
>>> # Load the response data into a JSON object
>>> data_json = json.loads(response.read())
>>> # Prepare the documents to be indexed
>>> documents = []
>>> for doc in data_json:
... documents.append(
... {
... "_index": "movies",
... "_source": doc,
... }
... )
>>> # Use helpers.bulk to index
>>> helpers.bulk(client, documents)
>>> print("Done indexing documents into `movies` index!")
>>> time.sleep(3)
Done indexing documents into `movies` index!
使用 Eland 上传 Hugging Face 模型
使用 Eland 的 eland_import_hub_model 命令,将模型上传到 Elasticsearch。本例选择 cross-encoder/ms-marco-MiniLM-L-6-v2 文本相似度模型。
>>> !eland_import_hub_model \
... --cloud-id $ELASTIC_CLOUD_ID \
... --es-api-key $ELASTIC_API_KEY \
... --hub-model-id cross-encoder/ms-marco-MiniLM-L-6-v2 \
... --task-type text_similarity \
... --clear-previous \
... --start
2024-08-13 17:04:12,386 INFO : Establishing connection to Elasticsearch
2024-08-13 17:04:12,567 INFO : Connected to serverless cluster 'bd8c004c050e4654ad32fb86ab159889'
2024-08-13 17:04:12,568 INFO : Loading HuggingFace transformer tokenizer and model 'cross-encoder/ms-marco-MiniLM-L-6-v2'
/usr/local/lib/python3.10/dist-packages/huggingface_hub/file_download.py:1132: FutureWarning: `resume_download` is deprecated and will be removed in version 1.0.0. Downloads always resume when possible. If you want to force a new download, use `force_download=True`.
warnings.warn(
tokenizer_config.json: 100% 316/316 [00:00<00:00, 1.81MB/s]
config.json: 100% 794/794 [00:00<00:00, 4.09MB/s]
vocab.txt: 100% 232k/232k [00:00<00:00, 2.37MB/s]
special_tokens_map.json: 100% 112/112 [00:00<00:00, 549kB/s]
pytorch_model.bin: 100% 90.9M/90.9M [00:00<00:00, 135MB/s]
STAGE:2024-08-13 17:04:15 1454:1454 ActivityProfilerController.cpp:312] Completed Stage: Warm Up
STAGE:2024-08-13 17:04:15 1454:1454 ActivityProfilerController.cpp:318] Completed Stage: Collection
STAGE:2024-08-13 17:04:15 1454:1454 ActivityProfilerController.cpp:322] Completed Stage: Post Processing
2024-08-13 17:04:18,789 INFO : Creating model with id 'cross-encoder__ms-marco-minilm-l-6-v2'
2024-08-13 17:04:21,123 INFO : Uploading model definition
100% 87/87 [00:55<00:00, 1.57 parts/s]
2024-08-13 17:05:16,416 INFO : Uploading model vocabulary
2024-08-13 17:05:16,987 INFO : Starting model deployment
2024-08-13 17:05:18,238 INFO : Model successfully imported with id 'cross-encoder__ms-marco-minilm-l-6-v2'
创建推理端点
接下来为 rerank 任务创建推理端点,部署和管理模型,并在必要时于后台启动所需的机器学习资源。
client.inference.put(
task_type="rerank",
inference_id="my-msmarco-minilm-model",
inference_config={
"service": "elasticsearch",
"service_settings": {
"model_id": "cross-encoder__ms-marco-minilm-l-6-v2",
"num_allocations": 1,
"num_threads": 1,
},
},
)
运行下面的命令,确认推理端点已部署。
client.inference.get()
部署模型时,可能需要在 Kibana 或 Serverless 界面中同步机器学习保存对象。进入 Trained Models(已训练模型),选择 Synchronize saved objects(同步保存对象)。
词法查询
先用 standard 检索器测试词法搜索(全文搜索),再比较叠加语义重排序后的改善效果。
使用 query_string 进行词法匹配
假设我们模糊地记得有一部著名电影,讲的是一个会吃掉受害者的杀手。为了举例,假设我们暂时忘了 cannibal(食人者)这个词。
使用 query_string 查询,在 Elasticsearch 文档的 plot 字段中搜索短语 flesh-eating bad guy(吃人肉的坏人)。 query_string query
>>> resp = client.search(
... index="movies",
... retriever={
... "standard": {
... "query": {
... "query_string": {
... "query": "flesh-eating bad guy",
... "default_field": "plot",
... }
... }
... }
... },
... )
>>> if resp["hits"]["hits"]:
... for hit in resp["hits"]["hits"]:
... title = hit["_source"]["title"]
... plot = hit["_source"]["plot"]
... print(f"Title: {title}\nPlot: {plot}\n")
>>> else:
... print("No search results found")
No search results found
没有结果!数据中没有与 flesh-eating bad guy 近乎精确匹配的内容。因为不知道 Elasticsearch 数据中更准确的措辞,我们需要扩大搜索范围。
简单的 multi_match 查询
这个词法查询在 Elasticsearch 文档的 plot 和 genre 字段中,对 crime 进行标准关键词搜索。
>>> resp = client.search(
... index="movies",
... retriever={"standard": {"query": {"multi_match": {"query": "crime", "fields": ["plot", "genre"]}}}},
... )
>>> for hit in resp["hits"]["hits"]:
... title = hit["_source"]["title"]
... plot = hit["_source"]["plot"]
... print(f"Title: {title}\nPlot: {plot}\n")
Title: The Godfather
Plot: An organized crime dynasty's aging patriarch transfers control of his clandestine empire to his reluctant son.
Title: Goodfellas
Plot: The story of Henry Hill and his life in the mob, covering his relationship with his wife Karen Hill and his mob partners Jimmy Conway and Tommy DeVito in the Italian-American crime syndicate.
Title: The Silence of the Lambs
Plot: A young F.B.I. cadet must receive the help of an incarcerated and manipulative cannibal killer to help catch another serial killer, a madman who skins his victims.
Title: Pulp Fiction
Plot: The lives of two mob hitmen, a boxer, a gangster and his wife, and a pair of diner bandits intertwine in four tales of violence and redemption.
Title: Se7en
Plot: Two detectives, a rookie and a veteran, hunt a serial killer who uses the seven deadly sins as his motives.
Title: The Departed
Plot: An undercover cop and a mole in the police attempt to identify each other while infiltrating an Irish gang in South Boston.
Title: The Usual Suspects
Plot: A sole survivor tells of the twisty events leading up to a horrific gun battle on a boat, which began when five criminals met at a seemingly random police lineup.
Title: The Dark Knight
Plot: When the menace known as the Joker wreaks havoc and chaos on the people of Gotham, Batman must accept one of the greatest psychological and physical tests of his ability to fight injustice.
这次好多了,至少有了结果。放宽搜索条件,提高了找到相关结果的机会。
但相对于最初的 flesh-eating bad guy 查询,这些结果还不够准确。通用 match 查询将《沉默的羔羊》(The Silence of the Lambs)排在结果列表中间。接下来看看语义重排序模型能否更接近搜索者的原始意图。
语义重排序器
在下面的 retriever 语法中,用 text_similarity_reranker 包裹标准查询检索器。这样,就能利用已经部署到 Elasticsearch 的 NLP 模型,按照 flesh-eating bad guy 这句话重新排列结果。
>>> resp = client.search(
... index="movies",
... retriever={
... "text_similarity_reranker": {
... "retriever": {"standard": {"query": {"multi_match": {"query": "crime", "fields": ["plot", "genre"]}}}},
... "field": "plot",
... "inference_id": "my-msmarco-minilm-model",
... "inference_text": "flesh-eating bad guy",
... }
... },
... )
>>> for hit in resp["hits"]["hits"]:
... title = hit["_source"]["title"]
... plot = hit["_source"]["plot"]
... print(f"Title: {title}\nPlot: {plot}\n")
Title: The Silence of the Lambs
Plot: A young F.B.I. cadet must receive the help of an incarcerated and manipulative cannibal killer to help catch another serial killer, a madman who skins his victims.
Title: Pulp Fiction
Plot: The lives of two mob hitmen, a boxer, a gangster and his wife, and a pair of diner bandits intertwine in four tales of violence and redemption.
Title: Se7en
Plot: Two detectives, a rookie and a veteran, hunt a serial killer who uses the seven deadly sins as his motives.
Title: Goodfellas
Plot: The story of Henry Hill and his life in the mob, covering his relationship with his wife Karen Hill and his mob partners Jimmy Conway and Tommy DeVito in the Italian-American crime syndicate.
Title: The Dark Knight
Plot: When the menace known as the Joker wreaks havoc and chaos on the people of Gotham, Batman must accept one of the greatest psychological and physical tests of his ability to fight injustice.
Title: The Godfather
Plot: An organized crime dynasty's aging patriarch transfers control of his clandestine empire to his reluctant son.
Title: The Departed
Plot: An undercover cop and a mole in the police attempt to identify each other while infiltrating an Irish gang in South Boston.
Title: The Usual Suspects
Plot: A sole survivor tells of the twisty events leading up to a horrific gun battle on a boat, which began when five criminals met at a seemingly random police lineup.
成功了!《沉默的羔羊》成为第一个结果。语义重排序通过解析自然语言查询,找到了最相关的结果,克服了更依赖精确匹配的词法搜索的局限。
语义重排序只需几个步骤就能实现语义搜索,无需生成和存储嵌入向量。在 Elasticsearch 集群中原生使用 Hugging Face 托管的开源模型,适合原型开发、测试和构建搜索体验。
进一步阅读
- 本例使用 cross-encoder/ms-marco-MiniLM-L-6-v2 文本相似度模型。Elasticsearch 支持的第三方文本相似度模型列表,参见 Elastic NLP 模型参考文档。 cross-encoder/ms-marco-MiniLM-L-6-v2 the Elastic NLP model reference
- 了解 Hugging Face 与 Elasticsearch 的集成方式。 integrating Hugging Face
- 查看 elasticsearch-labs 仓库中的 Elastic Python 笔记本目录。 elasticsearch-labs repo
- 了解 Elasticsearch 检索器与重排序。 retrievers and reranking in Elasticsearch











暂无评论内容