Skip to content

第 1 章 认识 Hugging Face 与 Transformers 生态 ​

学习目标 ​

  • 知道 Hugging Face 是什么、Hub 与各库的定位
  • 能按官方文档(huggingface.co/docs)找到对应库的教程与 API 参考
  • 用 uv 完成环境安装并核对版本
  • 运行第一条 pipeline,完成情感分析、完形填空与文本生成

1.1 Hugging Face 是什么 ​

Hugging Face 是 AI 开源社区的基础设施提供者,由两部分组成:

  1. Hub(huggingface.co):一个托管模型、数据集、度量指标的公共仓库。任何组织和个人都能上传、下载、搜索模型——就像 AI 界的 GitHub。
  2. 开源库生态:围绕 Hub 提供的一整套 Python 库,核心是 transformers。

transformers 提供统一的 Python 接口,让你用几乎相同的代码加载、训练和推理各种架构的模型(BERT、GPT、T5、Llama……),不必为每种模型重写代码。

本书围绕「模型训练与模型微调」这条主线,覆盖以下库:

库职责本书章节
transformers模型加载/保存/训练/推理的库全书主线
tokenizers分词器:文本 ↔ token 编号第 4 章
datasets数据集加载、变换、流式处理第 5~6 章
accelerate底层训练循环、混合精度、多卡第 9 章
peftLoRA 等参数高效微调第 10 章
trl监督微调与偏好对齐(SFT/DPO/GRPO)第 11 章
evaluate指标计算与评估第 12 章
huggingface_hub登录、搜索、上传模型第 12 章

1.2 官方文档怎么看 ​

打开 https://huggingface.co/docs,每个库的文档页通常分三类:

  1. Tutorials(教程):最短路径上手,如 transformers 的 Quick tour。
  2. How-to guides(指南):针对具体任务,如「微调一个文本分类模型」。
  3. API reference(API 参考):每个类、每个参数的精确定义,如 TrainingArguments。

本书相当于把这三类内容按教学顺序重排:先讲核心对象(第 2~4 章),再讲数据(第 5~6 章),然后是训练与微调(第 7~11 章),最后是评估与分享(第 12 章)。阅读本书时,遇到不确定的参数,回到官方 API 参考核对——那是最终权威。

1.3 环境安装 ​

本项目用 uv 管理 Python 环境。首次使用,在项目根目录:

$ uv sync

然后核对已安装的版本:

python
import transformers, torch
print("transformers:", transformers.__version__)
print("torch:", torch.__version__, "| cuda:", torch.cuda.is_available())

输出(本书写作时,你的版本可能更新):

transformers: 5.14.1
torch: 2.13.0+cu130 | cuda: True

本书所有示例基于 transformers 5.x 与 torch 2.x。示例的运行方式统一为 uv run python 文件名.py。

1.4 第一条 pipeline:情感分析 ​

pipeline 是 transformers 提供的最简推理入口:给它一个任务名和一段输入,它自动完成「选模型 → 加载 → 分词 → 前向计算 → 后处理」全部步骤。

python
from transformers import pipeline

classifier = pipeline("text-classification",
                      model="bhadresh-savani/bert-base-uncased-emotion")

for text in ["I love this course!",
             "This is absolutely terrible.",
             "I am so surprised by the news."]:
    print(text, "->", classifier(text))

输出(本书在本机真实运行):

I love this course! -> [{'label': 'joy', 'score': 0.9761354327201843}]
This is absolutely terrible. -> [{'label': 'sadness', 'score': 0.9911189675331116}]
I am so surprised by the news. -> [{'label': 'surprise', 'score': 0.9914474487304688}]

第一次运行会从 Hub 下载模型到本地缓存,之后直接使用缓存。pipeline 返回一个列表,每个元素是 {'label': 类别名, 'score': 概率}。

1.5 第二条 pipeline:完形填空 ​

fill-mask 用 [MASK] 遮住一个词,让模型预测最可能的词:

python
from transformers import pipeline

filler = pipeline("fill-mask", model="google-bert/bert-base-uncased")
print(filler("The capital of France is [MASK]."))

输出(真实运行,省略下载进度条):

[{'score': 0.41678887605667114, 'token': 3000, 'token_str': 'paris',
  'sequence': 'the capital of france is paris.'},
 {'score': 0.07141659408807755, 'token': 22479, 'token_str': 'lille', ...},
 {'score': 0.06339281797409058, 'token': 10241, 'token_str': 'lyon', ...},
 {'score': 0.044447511434555054, 'token': 16766, 'token_str': 'marseille', ...},
 {'score': 0.030297160148620605, 'token': 7562, 'token_str': 'tours', ...}]

模型按概率从高到低给出候选,paris 排第一。注意小写:因为 bert-base-uncased 会把输入转成小写。

1.6 第三条 pipeline:文本生成 ​

python
from transformers import pipeline

generator = pipeline("text-generation", model="openai-community/gpt2")
print(generator("I enjoy walking with my cute dog",
                max_new_tokens=20, do_sample=False))

输出(真实运行):

[{'generated_text': "I enjoy walking with my cute dog, but I'm not sure if I'll ever be able to walk with my dog. I'm"}]

GPT-2 是 2019 年的模型,输出质量有限,但机制清晰:模型逐个预测下一个 token。运行时会打印若干 v5 的提示(如 max_new_tokens 与默认 max_length 同时存在),不影响结果。

1.7 在 Hub 上搜索模型 ​

除了网页搜索,还可以用 huggingface_hub 的 HfApi 在代码里搜索:

python
from huggingface_hub import HfApi

api = HfApi()
models = api.list_models(search="bert", limit=3)
print([m.id for m in models])

输出:

['google-bert/bert-base-uncased', 'google-bert/bert-base-chinese', 'akdeniz27/bert-base-turkish-cased-ner']

search 参数按关键词搜索模型 id;limit 限制返回条数。

动手实践 ​

  1. 修改 1.4 中的例句,输入一句中文影评,观察情感模型输出什么(提示:该模型用英文训练,对中文会怎样?)。
  2. 用 fill-mask 让模型填空:"The [MASK] is raining outside.",比较前三个候选。
  3. 在 https://huggingface.co/models 用关键词 chinese sentiment 搜索,找到至少一个中文情感模型,记下它的 id,下一章用它做 pipeline。

常见错误 ​

错误 1:模型名拼写错误或任务名拼错。

ValueError: Unknown task text_classification, available tasks are ...

任务名用连字符:text-classification,不是下划线。模型 id 大小写敏感,复制时不要手打。

错误 2:以为 pipeline 只能加载一个固定模型。

pipeline("text-classification") 不指定 model 时,会使用该任务的默认模型;指定 model= 后使用你选择的模型。两者都是合法用法,差别只在「用谁」。

章末练习 ​

基础

  1. 用 pipeline("fill-mask", model="google-bert/bert-base-uncased") 让模型完成 3 个英文填空句,打印每个句子的 top-1 候选。
  2. 用 HfApi().list_models(search="gpt2", limit=5) 列出 5 个模型 id。

提高

  1. 把 1.4 的循环改成接收一个列表一次调用:classifier([text1, text2, ...]),对比返回结构与单句调用。
  2. 查看 pipeline 的 device 参数文档,说明在 CPU 与 GPU 上分别如何设置。

挑战

  1. 自己写一个小函数 run_task(task, text),内部调用 pipeline(task),对 text-classification 与 fill-mask 两个任务都能工作;无法处理的输入给出明确报错。

章末自测 ​

  1. transformers 生态中负责「数据集加载与变换」的库是?
  2. pipeline("text-classification") 中的 text-classification 是模型名还是任务名?
  3. 第一次运行 pipeline 时,模型文件会被下载到哪里?
  4. fill-mask 任务中 [MASK] 的作用是什么?
  5. 判断:所有 pipeline 任务都必须手动指定模型。
  6. HfApi().list_models(search="bert") 中的 search 参数作用是什么?