第 2 章 pipeline 与核心对象
学习目标
- 理解 pipeline 的「任务抽象」与它背后做了什么
- 掌握 transformers 三大核心对象:model(模型)、tokenizer(分词器)、config(配置)
- 会用
AutoModel、AutoTokenizer、AutoConfig手动组装一个 pipeline - 能列出 v5 支持的任务并选择合适任务
2.1 pipeline:任务抽象
上一章的 pipeline 看起来像魔法,其实它只做一件事:把任务名翻译成一串标准流程。以情感分析为例,pipeline("text-classification") 内部等价于:
- 根据任务名挑选默认模型;
AutoTokenizer.from_pretrained(model_id)加载分词器;AutoModelForSequenceClassification.from_pretrained(model_id)加载模型;- 对输入做分词、前向、softmax、按 label 映射输出。
这四步就是第 2~4 章要逐个拆开讲的东西。拆开之后,pipeline 就不再是魔法,而是一个「默认帮你选好了前三步」的便捷函数。
transformers v5 支持 27 个任务,查看方式:
from transformers.pipelines import get_supported_tasks
tasks = get_supported_tasks()
print("supported task count:", len(tasks))
print(tasks[:12])输出:
supported task count: 27
['any-to-any', 'audio-classification', 'automatic-speech-recognition', 'depth-estimation', 'document-question-answering', 'feature-extraction', 'fill-mask', 'image-classification', 'image-feature-extraction', 'image-segmentation', 'image-text-to-text', 'keypoint-matching']本书聚焦文本任务(text-classification、fill-mask、text-generation 等),但接口对图像、音频任务一致。
2.2 三大核心对象
pipeline 背后是三个对象,它们组成了 transformers 的一切:
model(模型):包含权重的神经网络,如 BERT、GPT-2。
tokenizer(分词器):把文本变成编号序列(并能在需要时还原成文本)。每种模型架构有配套的分词方式。
config(配置):描述模型的超参数——层数、隐藏层维度、词表大小等。模型结构由 config 决定,权重从 checkpoint 加载。
三者的典型生命周期:
config 描述结构 → tokenizer 负责文本 → model 负责计算
↓
from_pretrained 三者都从 Hub/本地加载2.3 用 AutoXxx 手动组装
AutoModel 系列类会根据模型 id 自动推断架构,不需要你手写 BertModel 还是 GPT2LMHeadModel:
from transformers import AutoConfig, AutoTokenizer, AutoModel
model_id = "google-bert/bert-base-uncased"
config = AutoConfig.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)
print("config 结构:", config.model_type, config.hidden_size, config.num_hidden_layers)
print("tokenizer:", type(tokenizer).__name__, "词表大小:", tokenizer.vocab_size)
print("model:", type(model).__name__)输出(真实运行):
config 结构: bert 768 12
tokenizer: BertTokenizerFast 词表大小: 30522
model: BertModel对应关系:
| Auto 类 | 负责对象 | 推断依据 |
|---|---|---|
AutoConfig | 结构配置 | 模型目录里的 config.json |
AutoTokenizer | 分词器 | 模型目录里的 tokenizer.json/vocab.txt |
AutoModel | 模型权重 | 模型目录里的 model.safetensors |
2.4 手动把对象组装成 pipeline
pipeline 接受显式传入的 model 与 tokenizer,这与只传任务名等价:
from transformers import pipeline, AutoTokenizer, AutoModelForMaskedLM
model_id = "google-bert/bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForMaskedLM.from_pretrained(model_id)
filler = pipeline("fill-mask", model=model, tokenizer=tokenizer, device=-1)
print(filler("Paris is the [MASK] of France.")[0]["token_str"])输出:
capital说明:
device=-1强制使用 CPU;device=0使用第一个 GPU;不传则按默认策略。AutoModelForMaskedLM是「带掩码语言建模头」的模型类,和裸AutoModel(只有主干)不同。选错头,pipeline 会报错或行为异常。
2.5 任务与模型头的对应
不同任务需要不同的输出头(head):分类任务要接分类头,生成任务要接语言建模头。Auto 类按任务命名:
| 任务 | Auto 类 |
|---|---|
| 掩码语言建模 | AutoModelForMaskedLM |
| 文本分类 | AutoModelForSequenceClassification |
| 因果语言建模(生成) | AutoModelForCausalLM |
| 问答 | AutoModelForQuestionAnswering |
用错类的典型症状:加载时打印 LOAD REPORT,出现大量 UNEXPECTED 权重(见第 3 章)。
动手实践
- 用
AutoTokenizer加载第 1 章你找到的中文情感模型,打印它的vocab_size与is_fast。 - 手动组装
pipeline("text-classification", model=..., tokenizer=...),输入三句中文,观察输出。 - 打印
AutoConfig.from_pretrained("google-bert/bert-base-uncased")的完整 dict,找出 5 个字段并解释含义。
常见错误
错误 1:用 AutoModel 加载分类模型。
AutoModel 只有主干、没有输出头,推理拿不到分类 logits。分类任务要用 AutoModelForSequenceClassification。
错误 2:忘记任务名与模型不匹配。
pipeline("text-classification", model="google-bert/bert-base-uncased") 会报错或行为异常,因为 bert-base-uncased 是掩码语言模型,没有分类头。任务、模型、Auto 类三者必须匹配。
错误 3:混淆 pipeline 的 device 与 model.to(device)。
pipeline 内部会自己管理设备;手动 model.to() 之后再传给 pipeline,设备可能被 pipeline 覆盖。用 device= 参数统一控制。
章末练习
基础
- 打印
get_supported_tasks()的完整列表,标出至少 5 个与文本相关的任务。 - 用
AutoConfig.from_pretrained("openai-community/gpt2")打印n_layer、n_embd、vocab_size,写出对应中文含义。
提高
- 分别用「只传任务名」和「显式传 model+tokenizer」两种方式创建
fill-maskpipeline,验证输出一致。 - 解释为什么加载同一个模型,
AutoModel与AutoModelForMaskedLM打印的 LOAD REPORT 不同。
挑战
- 自己实现一个极简版 pipeline:给定任务名
text-classification,用AutoTokenizer+AutoModelForSequenceClassification完成「分词 → 前向 → argmax → 映射到 label 名」,输出与官方 pipeline 相同结构的结果。
章末自测
- transformers 三大核心对象是哪三个?
AutoModel与AutoModelForSequenceClassification的区别是什么?pipeline(..., device=-1)中的-1表示什么?- 判断:
AutoConfig负责加载模型权重。 - 文本分类任务应使用哪个 Auto 类?
- transformers v5 共支持多少个 pipeline 任务?
