因果语言建模:微调DistilGPT2并生成文本

因果语言建模:微调DistilGPT2并生成文本

语言建模分为因果语言建模与掩码语言建模两类。本指南介绍因果语言建模。因果语言模型常用于文本生成,可以支持互动文字冒险等创意应用,也可以用于Copilot或CodeParrot这样的智能编程助手。

观看原文视频:因果语言建模

因果语言建模预测一个词元序列中的下一个词元,模型只能关注当前位置左侧的词元,因此看不到未来的词元。GPT-2就是因果语言模型的一个例子。

本指南将介绍:

  1. 在ELI5数据集的r/askscience子集上微调DistilGPT2。
  2. 使用微调后的模型推理。

开始前,确保安装了所需的库:

pip install transformers datasets evaluate

建议登录Hugging Face账号,以便上传模型并与社区分享。在提示输入时,用令牌登录:

>>> from huggingface_hub import notebook_login

>>> notebook_login()

加载ELI5数据集

先用🤗 Datasets加载ELI5-Category数据集的前5000个样本。这样可以先进行实验,确认流程正常,再投入更多时间训练完整数据集。

>>> from datasets import load_dataset

>>> eli5 = load_dataset("dany0407/eli5_category", split="train[:5000]")

使用Dataset.train_test_split将train划分为训练集和测试集:

>>> eli5 = eli5.train_test_split(test_size=0.2)

查看一个样本:

>>> eli5["train"][0]
{'q_id': '7h191n',
 'title': 'What does the tax bill that was passed today mean? How will it affect Americans in each tax bracket?',
 'selftext': '',
 'category': 'Economics',
 'subreddit': 'explainlikeimfive',
 'answers': {'a_id': ['dqnds8l', 'dqnd1jl', 'dqng3i1', 'dqnku5x'],
  'text': ["The tax bill is 500 pages long and there were a lot of changes still going on right to the end. It's not just an adjustment to the income tax brackets, it's a whole bunch of changes. As such there is no good answer to your question. The big take aways are: - Big reduction in corporate income tax rate will make large companies very happy. - Pass through rate change will make certain styles of business (law firms, hedge funds) extremely happy - Income tax changes are moderate, and are set to expire (though it's the kind of thing that might just always get re-applied without being made permanent) - People in high tax states (California, New York) lose out, and many of them will end up with their taxes raised.",
   'None yet. It has to be reconciled with a vastly different house bill and then passed again.',
   'Also: does this apply to 2017 taxes? Or does it start with 2018 taxes?',
   'This article explains both the House and senate bills, including the proposed changes to your income taxes based on your income level. URL_0'],
  'score': [21, 19, 5, 3],
  'text_urls': [[],
   [],
   [],
   ['https://www.investopedia.com/news/trumps-tax-reform-what-can-be-done/']]},
 'title_urls': ['url'],
 'selftext_urls': ['url']}

虽然内容看起来很多,但这里真正需要的只有text字段。语言建模的一个便利之处是无需额外标签,也称为无监督任务,因为下一个词本身就是标签。

预处理

观看原文视频:数据预处理

加载DistilGPT2分词器,处理text子字段:

>>> from transformers import AutoTokenizer

>>> tokenizer = AutoTokenizer.from_pretrained("distilbert/distilgpt2")

上例中的text实际上嵌套在answers中,需要用flatten展开嵌套结构:

>>> eli5 = eli5.flatten()
>>> eli5["train"][0]
{'q_id': '7h191n',
 'title': 'What does the tax bill that was passed today mean? How will it affect Americans in each tax bracket?',
 'selftext': '',
 'category': 'Economics',
 'subreddit': 'explainlikeimfive',
 'answers.a_id': ['dqnds8l', 'dqnd1jl', 'dqng3i1', 'dqnku5x'],
 'answers.text': ["The tax bill is 500 pages long and there were a lot of changes still going on right to the end. It's not just an adjustment to the income tax brackets, it's a whole bunch of changes. As such there is no good answer to your question. The big take aways are: - Big reduction in corporate income tax rate will make large companies very happy. - Pass through rate change will make certain styles of business (law firms, hedge funds) extremely happy - Income tax changes are moderate, and are set to expire (though it's the kind of thing that might just always get re-applied without being made permanent) - People in high tax states (California, New York) lose out, and many of them will end up with their taxes raised.",
  'None yet. It has to be reconciled with a vastly different house bill and then passed again.',
  'Also: does this apply to 2017 taxes? Or does it start with 2018 taxes?',
  'This article explains both the House and senate bills, including the proposed changes to your income taxes based on your income level. URL_0'],
 'answers.score': [21, 19, 5, 3],
 'answers.text_urls': [[],
  [],
  [],
  ['https://www.investopedia.com/news/trumps-tax-reform-what-can-be-done/']],
 'title_urls': ['url'],
 'selftext_urls': ['url']}

现在每个子字段都成为独立的一列,并带有answers前缀;text字段是列表。与其分别对每句话分词,不如把列表转成字符串后一起分词。

第一个预处理函数把每个样本的字符串列表连接起来,再对结果分词:

>>> def preprocess_function(examples):
...     return tokenizer([" ".join(x) for x in examples["answers.text"]])

用🤗 Datasets的Dataset.map将函数应用于整个数据集。设置batched=True可以一次处理多个元素,增大num_proc可使用更多进程。移除不需要的列:

>>> tokenized_eli5 = eli5.map(
...     preprocess_function,
...     batched=True,
...     num_proc=4,
...     remove_columns=eli5["train"].column_names,
... )

所得数据集包含词元序列,但部分序列超过模型的最大输入长度。接下来用第二个预处理函数连接所有序列,再按block_size切成较短的块。块长既应小于模型最大输入长度,也应小到适合GPU显存。

>>> block_size = 128


>>> def group_texts(examples):
...     # Concatenate all texts.
...     concatenated_examples = {k: sum(examples[k], []) for k in examples.keys()}
...     total_length = len(concatenated_examples[list(examples.keys())[0]])
...     # We drop the small remainder, we could add padding if the model supported it instead of this drop, you can
...     # customize this part to your needs.
...     if total_length >= block_size:
...         total_length = (total_length // block_size) * block_size
...     # Split by chunks of block_size.
...     result = {
...         k: [t[i : i + block_size] for i in range(0, total_length, block_size)]
...         for k, t in concatenated_examples.items()
...     }
...     result["labels"] = result["input_ids"].copy()
...     return result

对整个数据集应用group_texts:

>>> lm_dataset = tokenized_eli5.map(group_texts, batched=True, num_proc=4)

使用DataCollatorForLanguageModeling整理批次。整理时将句子动态填充到当前批次的最长长度,比把整个数据集都填充到最大长度更高效。

将序列结束词元设为填充词元,并设置mlm=False。原文将其描述为使用输入作为向右移一位的标签:

>>> from transformers import DataCollatorForLanguageModeling

>>> tokenizer.pad_token = tokenizer.eos_token
>>> data_collator = DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False)

训练

准备开始训练。通过AutoModelForCausalLM加载DistilGPT2:

>>> from transformers import AutoModelForCausalLM, TrainingArguments, Trainer

>>> model = AutoModelForCausalLM.from_pretrained("distilbert/distilgpt2")

接下来只剩三步:

  1. 在TrainingArguments中定义训练超参数。原文指出,唯一必填参数是指定保存目录的output_dir。设置push_to_hub=True可将模型推送到Hub;上传前需要登录Hugging Face。
  2. 把训练参数、模型、数据集和数据整理器传给Trainer。
  3. 调用Trainer.train微调模型。
>>> training_args = TrainingArguments(
...     output_dir="my_awesome_eli5_clm-model",
...     eval_strategy="epoch",
...     learning_rate=2e-5,
...     weight_decay=0.01,
...     push_to_hub=True,
... )

>>> trainer = Trainer(
...     model=model,
...     args=training_args,
...     train_dataset=lm_dataset["train"],
...     eval_dataset=lm_dataset["test"],
...     data_collator=data_collator,
...     processing_class=tokenizer,
... )

>>> trainer.train()

训练完成后,通过Trainer.evaluate评估模型并计算困惑度:

>>> import math

>>> eval_results = trainer.evaluate()
>>> print(f"Perplexity: {math.exp(eval_results['eval_loss']):.2f}")
Perplexity: 49.61

再用Trainer.push_to_hub把模型分享到Hub,供其他人使用:

>>> trainer.push_to_hub()

推理

模型微调完成后,就可以用于推理。先准备一个希望模型据此续写的提示:

>>> prompt = "Somatic hypermutation allows the immune system to"

最简单的试用方式是pipeline。用模型实例化文本生成流水线,再传入文本:

>>> from transformers import pipeline

>>> generator = pipeline("text-generation", model="username/my_awesome_eli5_clm-model")
>>> generator(prompt)
[{'generated_text': "Somatic hypermutation allows the immune system to be able to effectively reverse the damage caused by an infection.\n\n\nThe damage caused by an infection is caused by the immune system's ability to perform its own self-correcting tasks."}]

也可以先对文本分词,将input_ids作为PyTorch张量返回:

>>> from transformers import AutoTokenizer

>>> tokenizer = AutoTokenizer.from_pretrained("username/my_awesome_eli5_clm-model")
>>> inputs = tokenizer(prompt, return_tensors="pt").input_ids

调用generate生成文本。不同生成策略和控制生成的参数见文本生成策略。

>>> from transformers import AutoModelForCausalLM

>>> model = AutoModelForCausalLM.from_pretrained("username/my_awesome_eli5_clm-model")
>>> outputs = model.generate(inputs, max_new_tokens=100, do_sample=True, top_k=50, top_p=0.95)

将生成的词元ID解码回文本:

>>> tokenizer.batch_decode(outputs, skip_special_tokens=True)
["Somatic hypermutation allows the immune system to react to drugs with the ability to adapt to a different environmental situation. In other words, a system of 'hypermutation' can help the immune system to adapt to a different environmental situation or in some cases even a single life. In contrast, researchers at the University of Massachusetts-Boston have found that 'hypermutation' is much stronger in mice than in humans but can be found in humans, and that it's not completely unknown to the immune system. A study on how the immune system"]

来源:Hugging Face Transformers文档,Causal language modeling。Copyright 2023 The HuggingFace Team. All rights reserved. 本译稿对原文进行中文翻译,并单列实现说明。示例输出为原文展示结果,未在本稿环境实测。

Licensed under the Apache License, Version 2.0 (the “License”); you may not use this file except in compliance with the License. You may obtain a copy of the License at http://www.apache.org/licenses/LICENSE-2.0.

Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on an “AS IS” BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the specific language governing permissions and limitations under the License.

© 版权声明
THE END
喜欢就支持一下吧
点赞0 分享
评论 抢沙发

请登录后发表评论

    暂无评论内容