Skip to content

第 2 章 pipeline 与核心对象 ​

学习目标 ​

  • 理解 pipeline 的「任务抽象」与它背后做了什么
  • 掌握 transformers 三大核心对象:model(模型)、tokenizer(分词器)、config(配置)
  • 会用 AutoModel、AutoTokenizer、AutoConfig 手动组装一个 pipeline
  • 能列出 v5 支持的任务并选择合适任务

2.1 pipeline:任务抽象 ​

上一章的 pipeline 看起来像魔法,其实它只做一件事:把任务名翻译成一串标准流程。以情感分析为例,pipeline("text-classification") 内部等价于:

  1. 根据任务名挑选默认模型;
  2. AutoTokenizer.from_pretrained(model_id) 加载分词器;
  3. AutoModelForSequenceClassification.from_pretrained(model_id) 加载模型;
  4. 对输入做分词、前向、softmax、按 label 映射输出。

这四步就是第 2~4 章要逐个拆开讲的东西。拆开之后,pipeline 就不再是魔法,而是一个「默认帮你选好了前三步」的便捷函数。

transformers v5 支持 27 个任务,查看方式:

python
from transformers.pipelines import get_supported_tasks

tasks = get_supported_tasks()
print("supported task count:", len(tasks))
print(tasks[:12])

输出:

supported task count: 27
['any-to-any', 'audio-classification', 'automatic-speech-recognition', 'depth-estimation', 'document-question-answering', 'feature-extraction', 'fill-mask', 'image-classification', 'image-feature-extraction', 'image-segmentation', 'image-text-to-text', 'keypoint-matching']

本书聚焦文本任务(text-classification、fill-mask、text-generation 等),但接口对图像、音频任务一致。

2.2 三大核心对象 ​

pipeline 背后是三个对象,它们组成了 transformers 的一切:

model(模型):包含权重的神经网络,如 BERT、GPT-2。

tokenizer(分词器):把文本变成编号序列(并能在需要时还原成文本)。每种模型架构有配套的分词方式。

config(配置):描述模型的超参数——层数、隐藏层维度、词表大小等。模型结构由 config 决定,权重从 checkpoint 加载。

三者的典型生命周期:

text
config 描述结构 → tokenizer 负责文本 → model 负责计算
                ↓
        from_pretrained 三者都从 Hub/本地加载

2.3 用 AutoXxx 手动组装 ​

AutoModel 系列类会根据模型 id 自动推断架构,不需要你手写 BertModel 还是 GPT2LMHeadModel:

python
from transformers import AutoConfig, AutoTokenizer, AutoModel

model_id = "google-bert/bert-base-uncased"

config = AutoConfig.from_pretrained(model_id)
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)

print("config 结构:", config.model_type, config.hidden_size, config.num_hidden_layers)
print("tokenizer:", type(tokenizer).__name__, "词表大小:", tokenizer.vocab_size)
print("model:", type(model).__name__)

输出(真实运行):

config 结构: bert 768 12
tokenizer: BertTokenizerFast 词表大小: 30522
model: BertModel

对应关系:

Auto 类负责对象推断依据
AutoConfig结构配置模型目录里的 config.json
AutoTokenizer分词器模型目录里的 tokenizer.json/vocab.txt
AutoModel模型权重模型目录里的 model.safetensors

2.4 手动把对象组装成 pipeline ​

pipeline 接受显式传入的 model 与 tokenizer,这与只传任务名等价:

python
from transformers import pipeline, AutoTokenizer, AutoModelForMaskedLM

model_id = "google-bert/bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForMaskedLM.from_pretrained(model_id)

filler = pipeline("fill-mask", model=model, tokenizer=tokenizer, device=-1)
print(filler("Paris is the [MASK] of France.")[0]["token_str"])

输出:

capital

说明:

  • device=-1 强制使用 CPU;device=0 使用第一个 GPU;不传则按默认策略。
  • AutoModelForMaskedLM 是「带掩码语言建模头」的模型类,和裸 AutoModel(只有主干)不同。选错头,pipeline 会报错或行为异常。

2.5 任务与模型头的对应 ​

不同任务需要不同的输出头(head):分类任务要接分类头,生成任务要接语言建模头。Auto 类按任务命名:

任务Auto 类
掩码语言建模AutoModelForMaskedLM
文本分类AutoModelForSequenceClassification
因果语言建模(生成)AutoModelForCausalLM
问答AutoModelForQuestionAnswering

用错类的典型症状:加载时打印 LOAD REPORT,出现大量 UNEXPECTED 权重(见第 3 章)。

动手实践 ​

  1. 用 AutoTokenizer 加载第 1 章你找到的中文情感模型,打印它的 vocab_size 与 is_fast。
  2. 手动组装 pipeline("text-classification", model=..., tokenizer=...),输入三句中文,观察输出。
  3. 打印 AutoConfig.from_pretrained("google-bert/bert-base-uncased") 的完整 dict,找出 5 个字段并解释含义。

常见错误 ​

错误 1:用 AutoModel 加载分类模型。

AutoModel 只有主干、没有输出头,推理拿不到分类 logits。分类任务要用 AutoModelForSequenceClassification。

错误 2:忘记任务名与模型不匹配。

pipeline("text-classification", model="google-bert/bert-base-uncased") 会报错或行为异常,因为 bert-base-uncased 是掩码语言模型,没有分类头。任务、模型、Auto 类三者必须匹配。

错误 3:混淆 pipeline 的 device 与 model.to(device)。

pipeline 内部会自己管理设备;手动 model.to() 之后再传给 pipeline,设备可能被 pipeline 覆盖。用 device= 参数统一控制。

章末练习 ​

基础

  1. 打印 get_supported_tasks() 的完整列表,标出至少 5 个与文本相关的任务。
  2. 用 AutoConfig.from_pretrained("openai-community/gpt2") 打印 n_layer、n_embd、vocab_size,写出对应中文含义。

提高

  1. 分别用「只传任务名」和「显式传 model+tokenizer」两种方式创建 fill-mask pipeline,验证输出一致。
  2. 解释为什么加载同一个模型,AutoModel 与 AutoModelForMaskedLM 打印的 LOAD REPORT 不同。

挑战

  1. 自己实现一个极简版 pipeline:给定任务名 text-classification,用 AutoTokenizer + AutoModelForSequenceClassification 完成「分词 → 前向 → argmax → 映射到 label 名」,输出与官方 pipeline 相同结构的结果。

章末自测 ​

  1. transformers 三大核心对象是哪三个?
  2. AutoModel 与 AutoModelForSequenceClassification 的区别是什么?
  3. pipeline(..., device=-1) 中的 -1 表示什么?
  4. 判断:AutoConfig 负责加载模型权重。
  5. 文本分类任务应使用哪个 Auto 类?
  6. transformers v5 共支持多少个 pipeline 任务?