文本分类是一项常见的自然语言处理任务:为一段文本分配标签或类别。许多大型企业已将文本分类用于各种实际应用。情感分析是其中最常见的形式之一,它会给文本序列分配“🙂 正面”“🙁 负面”或“😐 中性”等标签。
下面依次完成两件事:在 IMDb 数据集上微调 DistilBERT,判断影评是正面还是负面;随后用微调后的模型进行推理。适合这项任务的全部架构和检查点,可在文本分类任务页查看。
对应的视频教程见 Hugging Face 文本分类视频,可运行的交互笔记本见 PyTorch 文本分类 Colab。
准备依赖
开始前,安装所需的库:
pip install transformers datasets evaluate accelerate
如果准备将模型上传到 Hugging Face Hub 并与社区分享,可以登录自己的 Hugging Face 账户。下面的调用会显示登录提示;令牌只应输入官方登录界面,不应写入代码或文章。
>>> from huggingface_hub import notebook_login
>>> notebook_login()
加载 IMDb 数据集
先通过 🤗 Datasets 库加载 IMDb:
>>> from datasets import load_dataset
>>> imdb = load_dataset("stanfordnlp/imdb")
查看一个样本:
>>> imdb["test"][0]
{
"label": 0,
"text": "I love sci-fi and am willing to put up with a lot. Sci-fi movies/TV are usually underfunded, under-appreciated and misunderstood. I tried to like this, I really did, but it is to good TV sci-fi as Babylon 5 is to Star Trek (the original). Silly prosthetics, cheap cardboard sets, stilted dialogues, CG that doesn't match the background, and painfully one-dimensional characters cannot be overcome with a 'sci-fi' setting. (I'm sure there are those of you out there who think Babylon 5 is good sci-fi TV. It's not. It's clichéd and uninspiring.) While US viewers might like emotion and character development, sci-fi is a genre that does not take itself seriously (cf. Star Trek). It may treat important issues, yet not as a serious philosophy. It's really difficult to care about the characters here as they are not simply foolish, just missing a spark of life. Their actions and reactions are wooden and predictable, often painful to watch. The makers of Earth KNOW it's rubbish as they have to always say \"Gene Roddenberry's Earth...\" otherwise people would not continue watching. Roddenberry's ashes must be turning in their orbit as this dull, cheap, poorly edited (watching it without advert breaks really brings this home) trudging Trabant of a show lumbers into space. Spoiler. So, kill off a main character. And then bring him back as another actor. Jeeez! Dallas all over again.",
}
数据集包含两个字段:
text:影评文本。上例保留英文原始数据,避免改变模型实际接收的样本。label:负面影评为0,正面影评为1。
预处理
加载 DistilBERT 的分词器,处理 text 字段:
>>> from transformers import AutoTokenizer
>>> tokenizer = AutoTokenizer.from_pretrained("distilbert/distilbert-base-uncased")
创建预处理函数,对 text 分词,并将序列截断到不超过 DistilBERT 的最大输入长度:
>>> def preprocess_function(examples):
... return tokenizer(examples["text"], truncation=True)
使用 Datasets 的 undefined,将预处理函数应用到整个数据集。设置 batched=True 后,map 会一次处理多个样本,通常能加快预处理:
tokenized_imdb = imdb.map(preprocess_function, batched=True)
接着用 undefined 整理一个批次中的样本。整理批次时,将句子动态补齐到该批次最长句子的长度,比把整个数据集预先补齐到模型允许的最大长度更高效。
>>> from transformers import DataCollatorWithPadding
>>> data_collator = DataCollatorWithPadding(tokenizer=tokenizer)
评估
在训练过程中加入指标,有助于评估模型表现。🤗 Evaluate 可以加载评估方法;这里使用准确率指标。加载和计算指标的更多说明见 Evaluate 入门。
>>> import evaluate
>>> accuracy = evaluate.load("accuracy")
创建一个函数,把预测值和真实标签传给 undefined,计算准确率:
>>> import numpy as np
>>> def compute_metrics(eval_pred):
... predictions, labels = eval_pred
... predictions = np.argmax(predictions, axis=1)
... return accuracy.compute(predictions=predictions, references=labels)
compute_metrics 已准备好,配置训练时会再次用到它。
训练
首先建立类别编号与标签名称之间的双向映射:
>>> id2label = {0: "NEGATIVE", 1: "POSITIVE"}
>>> label2id = {"NEGATIVE": 0, "POSITIVE": 1}
如果还不熟悉通过 undefined 微调模型,可以先阅读基础训练教程。
现在使用 undefined 加载 DistilBERT,并传入类别数量以及标签映射:
>>> from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer
>>> model = AutoModelForSequenceClassification.from_pretrained(
... "distilbert/distilbert-base-uncased", num_labels=2, id2label=id2label, label2id=label2id
... )
之后还剩三个步骤:
- 在 undefined 中定义超参数。原教程将
output_dir列为唯一必填参数,用它指定模型的保存位置。push_to_hub=True启用上传到 Hub,需要事先登录并拥有对应仓库的写入权限。每个 epoch 结束时,Trainer会评估准确率,并保存训练检查点。 - 把训练参数、模型、数据集、分词器、数据整理器以及
compute_metrics函数传给Trainer。 - 调用
Trainer.train微调模型。
>>> training_args = TrainingArguments(
... output_dir="my_awesome_model",
... learning_rate=2e-5,
... per_device_train_batch_size=16,
... per_device_eval_batch_size=16,
... num_train_epochs=2,
... weight_decay=0.01,
... eval_strategy="epoch",
... save_strategy="epoch",
... load_best_model_at_end=True,
... push_to_hub=True,
... )
>>> trainer = Trainer(
... model=model,
... args=training_args,
... train_dataset=tokenized_imdb["train"],
... eval_dataset=tokenized_imdb["test"],
... processing_class=tokenizer,
... data_collator=data_collator,
... compute_metrics=compute_metrics,
... )
>>> trainer.train()
此示例用 IMDb 的 test 切分进行每轮评估。正式比较模型时,应另外划出验证集调参,并保留从未参与调参的最终测试集;教程中的用法不应直接当作严格的最终评测设计。
示例使用 eval_strategy 和 processing_class 等接口名称,应与实际安装的 Transformers 版本核对。原文提示还提到,将分词器传给 Trainer 时,默认的数据整理器可进行动态补齐;本例已经显式传入 data_collator,因此批次补齐方式清楚可见。
训练完成后,可以调用 Trainer.push_to_hub 分享模型:
>>> trainer.push_to_hub()
如果只准备在本地训练,将 push_to_hub 设为 False,并跳过登录和上传调用。本文保留原示例的上传配置,未执行训练或上传,也没有创建任何令牌。更完整的微调示例见对应的 PyTorch 笔记本。
推理
微调完成后,就能用模型预测文本的类别。先准备一段待分析的文本:
>>> text = "This was a masterpiece. Not completely faithful to the books, but enthralling from beginning to end. Might be my favorite of the three."
这段英文影评大意是:“这是一部杰作。虽然并非完全忠于原著,但从头到尾都引人入胜。可能是三部中我最喜欢的一部。”输入文本保留英文,符合本例的英语模型和数据集。
最便捷的试用方式是 undefined。用模型创建情感分析流水线,再传入文本:
>>> from transformers import pipeline
>>> classifier = pipeline("sentiment-analysis", model="stevhliu/my_awesome_model")
>>> classifier(text)
[{'label': 'POSITIVE', 'score': 0.9994940757751465}]
stevhliu/my_awesome_model 是原教程作者的示例仓库。使用自己训练的模型时,需要换成自己的模型仓库名或本地保存目录;这个名字不会自动指向上一节保存的模型。0.9994940757751465 是原文给出的示例分数,并非本文实测,也不保证换用其他模型后得到同一结果。
也可以手动复现流水线的处理步骤。先对文本分词,返回 PyTorch 张量:
>>> from transformers import AutoTokenizer
>>> tokenizer = AutoTokenizer.from_pretrained("stevhliu/my_awesome_model")
>>> inputs = tokenizer(text, return_tensors="pt")
将输入传入模型,取得 logits:
>>> import torch
>>> from transformers import AutoModelForSequenceClassification
>>> model = AutoModelForSequenceClassification.from_pretrained("stevhliu/my_awesome_model")
>>> with torch.no_grad():
... logits = model(**inputs).logits
取分数最高的类别,再用模型配置中的 id2label 映射转换成文本标签:
>>> predicted_class_id = logits.argmax().item()
>>> model.config.id2label[predicted_class_id]
'POSITIVE'
这里输入只有一个样本,因此可直接把 argmax 结果转成单个编号。批量推理时应按类别维度逐条取最大值,而不是直接照搬这个单样本表达式。
来源与许可
原文:Text classification,Hugging Face Transformers 文档贡献者。正文依据官方仓库的 sequence_classification.md 整理为中文;代码和原文示例输出保留,另补充了版本、测试集和本地使用边界。
Copyright 2022 The HuggingFace Team. All rights reserved.
本文所据文档遵循 Apache License, Version 2.0。使用与分发须遵守该许可。相关材料按“AS IS”提供,不附带任何明示或默示担保;具体权限与限制以许可原文为准。交付目录同时附有完整的 LICENSE.txt。中文翻译及明确标出的补充说明属于对原文的修改。











暂无评论内容