Skip to content

第 12 章 评估、分享与部署 ​

学习目标 ​

  • 用 evaluate 库计算 accuracy、f1 等指标
  • 理解评估在训练流程中的正确位置
  • 掌握 huggingface_hub 登录与 push_to_hub
  • 会编写模型卡并用 pipeline 加载本地/远端模型

12.1 evaluate:指标即模块 ​

evaluate 库把指标做成可加载的模块,统一接口:

python
from evaluate import load as load_metric

accuracy = load_metric("accuracy")
print(accuracy.compute(predictions=[0, 1, 1, 0],
                       references=[0, 1, 0, 0]))

f1 = load_metric("f1")
print(f1.compute(predictions=[0, 1, 1, 0],
                 references=[0, 1, 0, 0],
                 average="binary"))

输出:

{'accuracy': 0.75}
{'f1': 0.6666666666666666}

用法统一:compute(predictions=..., references=...)。需要 scikit-learn 支持的部分指标(如 accuracy/f1)已随本书环境安装。

12.2 评估要放在哪里 ​

评估不是训练后的「补做」,而是训练流程的一部分:

  1. 训练前:划分验证集(第 5 章),确认类别分布。
  2. 训练中:eval_strategy="epoch" 或 "steps",让 Trainer 定期评估(第 7 章)。
  3. 训练后:trainer.evaluate() 出最终指标;对错误样本做人工检查(第 8 章练习)。

多卡手动循环里,统计指标前要把张量收集回来:

python
all_preds = accelerator.gather(preds)
all_labels = accelerator.gather(batch["labels"])

否则每张卡只统计自己的 batch,数字是错的。

12.3 Hub 分享:登录 ​

上传模型前必须先登录。命令行:

$ hf auth login

或 Python:

python
from huggingface_hub import login

login()  # 交互式输入 token

token 在 https://huggingface.co/settings/tokens 创建。验证登录:

python
from huggingface_hub import whoami

print(whoami())

未登录时调用会报:

huggingface_hub.errors.LocalTokenNotFoundError: Token is required to call the /whoami-v2 endpoint, but no token found. You must provide a token or be logged in to Hugging Face with `hf auth login` or `huggingface-cli login`.

12.4 分享模型:push_to_hub ​

模型与分词器都有 push_to_hub 方法,把本地目录内容推到你的仓库:

python
model.push_to_hub("sst2-bert-tiny")      # 推权重与 config
tokenizer.push_to_hub("sst2-bert-tiny")  # 推分词器

仓库会自动创建在你的用户名下(<用户名>/sst2-bert-tiny)。如果命名空间不可写,会得到 403:

huggingface_hub.errors.HfHubHTTPError: 403 Forbidden: You don't have the rights to create a model under the namespace "xxx".

要点:

  • 上传前先确认 config.json 里有 id2label/label2id(第 8 章),否则别人推理得到 LABEL_0。
  • 用 Trainer 时,在 TrainingArguments 里设 push_to_hub=True 可在训练结束自动上传。

12.5 模型卡:README 即文档 ​

Hub 仓库的 README.md 就是模型卡,会被渲染成漂亮的介绍页。最小模板:

markdown
---
license: mit
language:
  - en
tags:
  - text-classification
  - sentiment-analysis
---

# sst2-bert-tiny

基于 `prajjwal1/bert-tiny` 在 SST-2 上微调的情感分类模型。

## 使用方法

```python
from transformers import pipeline
classifier = pipeline("text-classification",
                      model="你的用户名/sst2-bert-tiny")
print(classifier("This movie is great!"))
```

## 训练细节

- 数据:GLUE SST-2 前 2000 条
- 超参:lr=1e-4, batch=16, 2 epochs
- 验证集准确率:0.695

模型卡记录「数据、超参、指标、已知局限」,是专业模型的通行证。

12.6 部署:从本地目录或 Hub 加载 ​

第 8 章保存的本地目录可以直接被 pipeline 加载:

python
from transformers import pipeline

classifier = pipeline("text-classification",
                      model="./sst2-bert-tiny")
print(classifier("This movie was absolutely wonderful and touching."))

输出(真实运行):

[{'label': 'positive', 'score': 0.7128987312316895}]

推到 Hub 后,把本地路径换成模型 id 即可,别人也能用:

python
classifier = pipeline("text-classification",
                      model="你的用户名/sst2-bert-tiny")

生产环境进一步优化方向(超出本书范围,不再展开):text-generation-inference 服务化、vLLM 批推理、ONNX/torch.compile 加速、4/8 位量化。

动手实践 ​

  1. 用 evaluate 计算你自己微调模型的 accuracy 与 f1(第 8 章模型)。
  2. 为模型写一份模型卡,包含数据、超参、指标与一个已知失败案例。
  3. 如果已有 token:把模型推到 Hub,用模型 id 从远端重新加载并预测;没有 token 就练习到 12.4 为止,记录报错信息。

常见错误 ​

错误 1:上传前没检查 id2label。

别人加载你的模型得到 LABEL_0——推理「能跑」但不可用。上传前先跑一次 pipeline 验证输出。

错误 2:把 references 与 predictions 传反。

两个参数顺序固定:predictions 是模型输出,references 是真实标签。传反时指标照样算出来,但含义完全错。

错误 3:模型卡不写局限。

只写准确率、不写失败案例的模型卡是营销页;写清楚「在 XX 数据上会误判」才是工程文档。

章末练习 ​

基础

  1. 用 evaluate 计算一组 10 条预测的 accuracy 与 f1。
  2. 写出 whoami() 未登录时的报错类型名。

提高

  1. 为第 8 章模型写完整模型卡,并用 push_to_hub 上传(或模拟:先写 README.md 到本地模型目录)。
  2. 解释 accelerator.gather 在多卡评估中的作用。

挑战

  1. 构建一个「评估 + 报告」脚本:输入模型目录与验证集,输出 accuracy/f1、混淆矩阵、10 条错误样本及原因分析,结果写成一个 Markdown 报告。

章末自测 ​

  1. evaluate.load("accuracy") 返回什么对象?如何计算指标?
  2. 上传模型前必须完成什么?
  3. push_to_hub 会把哪些文件上传?
  4. 模型卡是什么文件的 Markdown?
  5. 判断:模型部署时本地路径与 Hub 模型 id 不能互换使用。
  6. 多卡评估时为什么要 accelerator.gather?