第 1 章 认识 Hugging Face 与 Transformers 生态
学习目标
- 知道 Hugging Face 是什么、
Hub与各库的定位 - 能按官方文档(huggingface.co/docs)找到对应库的教程与 API 参考
- 用 uv 完成环境安装并核对版本
- 运行第一条
pipeline,完成情感分析、完形填空与文本生成
1.1 Hugging Face 是什么
Hugging Face 是 AI 开源社区的基础设施提供者,由两部分组成:
- Hub(huggingface.co):一个托管模型、数据集、度量指标的公共仓库。任何组织和个人都能上传、下载、搜索模型——就像 AI 界的 GitHub。
- 开源库生态:围绕 Hub 提供的一整套 Python 库,核心是 transformers。
transformers 提供统一的 Python 接口,让你用几乎相同的代码加载、训练和推理各种架构的模型(BERT、GPT、T5、Llama……),不必为每种模型重写代码。
本书围绕「模型训练与模型微调」这条主线,覆盖以下库:
| 库 | 职责 | 本书章节 |
|---|---|---|
| transformers | 模型加载/保存/训练/推理的库 | 全书主线 |
| tokenizers | 分词器:文本 ↔ token 编号 | 第 4 章 |
| datasets | 数据集加载、变换、流式处理 | 第 5~6 章 |
| accelerate | 底层训练循环、混合精度、多卡 | 第 9 章 |
| peft | LoRA 等参数高效微调 | 第 10 章 |
| trl | 监督微调与偏好对齐(SFT/DPO/GRPO) | 第 11 章 |
| evaluate | 指标计算与评估 | 第 12 章 |
| huggingface_hub | 登录、搜索、上传模型 | 第 12 章 |
1.2 官方文档怎么看
打开 https://huggingface.co/docs,每个库的文档页通常分三类:
- Tutorials(教程):最短路径上手,如 transformers 的 Quick tour。
- How-to guides(指南):针对具体任务,如「微调一个文本分类模型」。
- API reference(API 参考):每个类、每个参数的精确定义,如
TrainingArguments。
本书相当于把这三类内容按教学顺序重排:先讲核心对象(第 2~4 章),再讲数据(第 5~6 章),然后是训练与微调(第 7~11 章),最后是评估与分享(第 12 章)。阅读本书时,遇到不确定的参数,回到官方 API 参考核对——那是最终权威。
1.3 环境安装
本项目用 uv 管理 Python 环境。首次使用,在项目根目录:
$ uv sync然后核对已安装的版本:
import transformers, torch
print("transformers:", transformers.__version__)
print("torch:", torch.__version__, "| cuda:", torch.cuda.is_available())输出(本书写作时,你的版本可能更新):
transformers: 5.14.1
torch: 2.13.0+cu130 | cuda: True本书所有示例基于 transformers 5.x 与 torch 2.x。示例的运行方式统一为
uv run python 文件名.py。
1.4 第一条 pipeline:情感分析
pipeline 是 transformers 提供的最简推理入口:给它一个任务名和一段输入,它自动完成「选模型 → 加载 → 分词 → 前向计算 → 后处理」全部步骤。
from transformers import pipeline
classifier = pipeline("text-classification",
model="bhadresh-savani/bert-base-uncased-emotion")
for text in ["I love this course!",
"This is absolutely terrible.",
"I am so surprised by the news."]:
print(text, "->", classifier(text))输出(本书在本机真实运行):
I love this course! -> [{'label': 'joy', 'score': 0.9761354327201843}]
This is absolutely terrible. -> [{'label': 'sadness', 'score': 0.9911189675331116}]
I am so surprised by the news. -> [{'label': 'surprise', 'score': 0.9914474487304688}]第一次运行会从 Hub 下载模型到本地缓存,之后直接使用缓存。pipeline 返回一个列表,每个元素是 {'label': 类别名, 'score': 概率}。
1.5 第二条 pipeline:完形填空
fill-mask 用 [MASK] 遮住一个词,让模型预测最可能的词:
from transformers import pipeline
filler = pipeline("fill-mask", model="google-bert/bert-base-uncased")
print(filler("The capital of France is [MASK]."))输出(真实运行,省略下载进度条):
[{'score': 0.41678887605667114, 'token': 3000, 'token_str': 'paris',
'sequence': 'the capital of france is paris.'},
{'score': 0.07141659408807755, 'token': 22479, 'token_str': 'lille', ...},
{'score': 0.06339281797409058, 'token': 10241, 'token_str': 'lyon', ...},
{'score': 0.044447511434555054, 'token': 16766, 'token_str': 'marseille', ...},
{'score': 0.030297160148620605, 'token': 7562, 'token_str': 'tours', ...}]模型按概率从高到低给出候选,paris 排第一。注意小写:因为 bert-base-uncased 会把输入转成小写。
1.6 第三条 pipeline:文本生成
from transformers import pipeline
generator = pipeline("text-generation", model="openai-community/gpt2")
print(generator("I enjoy walking with my cute dog",
max_new_tokens=20, do_sample=False))输出(真实运行):
[{'generated_text': "I enjoy walking with my cute dog, but I'm not sure if I'll ever be able to walk with my dog. I'm"}]GPT-2 是 2019 年的模型,输出质量有限,但机制清晰:模型逐个预测下一个 token。运行时会打印若干 v5 的提示(如 max_new_tokens 与默认 max_length 同时存在),不影响结果。
1.7 在 Hub 上搜索模型
除了网页搜索,还可以用 huggingface_hub 的 HfApi 在代码里搜索:
from huggingface_hub import HfApi
api = HfApi()
models = api.list_models(search="bert", limit=3)
print([m.id for m in models])输出:
['google-bert/bert-base-uncased', 'google-bert/bert-base-chinese', 'akdeniz27/bert-base-turkish-cased-ner']search 参数按关键词搜索模型 id;limit 限制返回条数。
动手实践
- 修改 1.4 中的例句,输入一句中文影评,观察情感模型输出什么(提示:该模型用英文训练,对中文会怎样?)。
- 用
fill-mask让模型填空:"The [MASK] is raining outside.",比较前三个候选。 - 在 https://huggingface.co/models 用关键词
chinese sentiment搜索,找到至少一个中文情感模型,记下它的 id,下一章用它做 pipeline。
常见错误
错误 1:模型名拼写错误或任务名拼错。
ValueError: Unknown task text_classification, available tasks are ...任务名用连字符:text-classification,不是下划线。模型 id 大小写敏感,复制时不要手打。
错误 2:以为 pipeline 只能加载一个固定模型。
pipeline("text-classification") 不指定 model 时,会使用该任务的默认模型;指定 model= 后使用你选择的模型。两者都是合法用法,差别只在「用谁」。
章末练习
基础
- 用
pipeline("fill-mask", model="google-bert/bert-base-uncased")让模型完成 3 个英文填空句,打印每个句子的 top-1 候选。 - 用
HfApi().list_models(search="gpt2", limit=5)列出 5 个模型 id。
提高
- 把 1.4 的循环改成接收一个列表一次调用:
classifier([text1, text2, ...]),对比返回结构与单句调用。 - 查看
pipeline的device参数文档,说明在 CPU 与 GPU 上分别如何设置。
挑战
- 自己写一个小函数
run_task(task, text),内部调用pipeline(task),对text-classification与fill-mask两个任务都能工作;无法处理的输入给出明确报错。
章末自测
- transformers 生态中负责「数据集加载与变换」的库是?
pipeline("text-classification")中的text-classification是模型名还是任务名?- 第一次运行 pipeline 时,模型文件会被下载到哪里?
fill-mask任务中[MASK]的作用是什么?- 判断:所有 pipeline 任务都必须手动指定模型。
HfApi().list_models(search="bert")中的search参数作用是什么?
