第 5 章 2012—2014:深度学习元年——AlexNet、VGG 与 GoogLeNet
学习目标
- 理解 AlexNet(2012)为什么是「深度学习元年」的引爆点
- 理解 VGG(2014)「小卷积核深堆叠」的思想
- 理解 GoogLeNet / Inception(2014)的多尺度并行思想
- 分别用 PyTorch 实现并训练三种经典 CNN 架构
5.1 时代背景:ImageNet 与 GPU
2012 年 10 月,Alex Krizhevsky、Ilya Sutskever 和 Geoffrey Hinton 提交了 AlexNet 论文《ImageNet Classification with Deep Convolutional Neural Networks》。它在 ImageNet 大赛上把 top-5 错误率从 26% 压到 15.3%,碾压第二名——而它靠的不只是网络结构,还有两件「新武器」:
- GPU:两块 GTX 580 并行训练,卷积可以在 GPU 上批量计算;
- 数据:ImageNet 有 120 万张标注图片,足够训练深层网络。
再加上 ReLU(解决梯度消失)、Dropout(缓解过拟合)、数据增强(扩充训练集),深度学习从边缘进入主流。2014 年,VGG 和 GoogLeNet 把错误率进一步降到 7% 以下。
5.2 AlexNet:大核、深堆叠、ReLU
AlexNet 的配方(输入 227×227):
| 层 | 操作 | 输出 |
|---|---|---|
| Conv1 | 96 个 11×11 卷积,stride 4 | 55×55×96 |
| Pool1 | 3×3 最大池化,stride 2 | 27×27×96 |
| Conv2 | 256 个 5×5 卷积 | 27×27×256 |
| Pool2 | 3×3 池化 | 13×13×256 |
| Conv3-5 | 3×3 / 3×3 / 3×3 | 13×13×256 |
| Pool5 | 3×3 池化 | 6×6×256 |
| FC | 4096 → 4096 → 1000 | 1000 |
三个关键点:全部用 ReLU(训练比 tanh 快数倍且不饱和)、Dropout(随机丢弃一半神经元,防止全连接层过拟合)、重叠池化(池化窗口 3×3、步长 2,略降错误率)。
examples/ch05_alexnet.py 把 AlexNet 适配到 28×28 的 Fashion-MNIST(卷积核等比例缩小):
AlexNet 风格参数量:11,385,546
epoch 1:训练准确率 = 87.21%
epoch 2:训练准确率 = 89.11%
epoch 3:训练准确率 = 91.73%
测试准确率 = 90.22%注意参数量高达 1138 万,绝大部分在最后三个全连接层。这是 2012 年的常态——全连接层是大头,也是后来被批评浪费的地方。
5.3 VGG:小卷积核的堆叠艺术
2014 年,牛津大学的 Karen Simonyan 和 Andrew Zisserman 提出 VGG,核心观察:两个 3×3 卷积等价于一个 5×5 卷积的感受野,但参数更少、非线性更多。
一个 5×5 卷积有 25 个参数,两个 3×3 有 18 个;同时多了一次 ReLU,表达力更强。于是 VGG 把 AlexNet 的大核全部换成 3×3,靠深度(16~19 层)取胜:VGG16 在 ImageNet 上达到 7.3% top-5 错误率。
examples/ch05_vgg.py 用「卷积块(2~3 个 3×3 卷积 + 池化)」搭建 TinyVGG:
def conv_block(in_c, out_c, reps=2):
layers = []
for _ in range(reps):
layers += [nn.Conv2d(in_c, out_c, 3, padding=1), nn.BatchNorm2d(out_c), nn.ReLU()]
in_c = out_c
layers.append(nn.MaxPool2d(2))
return nn.Sequential(*layers)TinyVGG 参数量:2,921,930
epoch 1:训练准确率 = 89.77%
epoch 2:训练准确率 = 91.42%
epoch 3:训练准确率 = 93.72%
测试准确率 = 91.83%注意 TinyVGG 只有 AlexNet 风格网络 1/4 的参数,准确率却更高(91.8% vs 90.2%)——小卷积核 + 批量归一化(BN 见第 6 章)的威力。
5.4 GoogLeNet / Inception:一层干多种活
2014 年的另一个冠军是 Google 的 GoogLeNet(向 LeNet 致敬的名字)。它的核心模块 Inception 在同一层并行跑四种操作:1×1 卷积、3×3 卷积、5×5 卷积、3×3 池化,然后把结果拼接起来。这样一层就能同时捕捉不同尺度的特征,让网络自己「挑」。
1×1 卷积是 Inception 的秘密武器:先把通道数压缩(降维),再做大核卷积,计算量大幅下降。GoogLeNet 只有 AlexNet 的 1/12 参数,却拿了 2014 年 ImageNet 冠军。
examples/ch05_inception.py 实现 Inception 模块:
class Inception(nn.Module):
def __init__(self, in_c):
super().__init__()
self.b1 = nn.Sequential(nn.Conv2d(in_c, 32, 1), nn.ReLU())
self.b2 = nn.Sequential(
nn.Conv2d(in_c, 32, 1), nn.ReLU(),
nn.Conv2d(32, 48, 3, padding=1), nn.ReLU(),
)
self.b3 = nn.Sequential(
nn.Conv2d(in_c, 16, 1), nn.ReLU(),
nn.Conv2d(16, 32, 5, padding=2), nn.ReLU(),
)
self.b4 = nn.Sequential(
nn.MaxPool2d(3, stride=1, padding=1),
nn.Conv2d(in_c, 16, 1), nn.ReLU(),
)
def forward(self, x):
return torch.cat([self.b1(x), self.b2(x), self.b3(x), self.b4(x)], dim=1)Inception 风格参数量:109,658
epoch 1:训练准确率 = 76.39%
epoch 2:训练准确率 = 79.64%
epoch 3:训练准确率 = 83.12%
测试准确率 = 82.26%只有 11 万参数就能达到 82%,Inception 的效率优势明显。准确率比 VGG 低是因为 Fashion-MNIST 太小、3 个 epoch 太短——在 ImageNet 这种大规模数据集上,它的精度才是第一梯队。
5.5 三条路线的启示
2012—2014 年给出了 CNN 设计的三个方向,今天仍在使用:
- AlexNet 路线:更大的模型 + 更强的正则化;
- VGG 路线:更小的核、更深的堆叠;
- GoogLeNet 路线:多尺度并行 + 1×1 降维。
它们共同奠定了「卷积骨干 + 分类头」的标准范式。下一步的瓶颈是深度:当网络深到 20 层以上,训练开始退化——这正是第 6 章 ResNet 要解决的问题。
完整代码
本章用到的完整示例代码:
examples/ch05_alexnet.py
"""第 5 章示例(2012):AlexNet 风格网络在 Fashion-MNIST 上训练。
AlexNet 的配方:5 层卷积 + 3 层全连接、ReLU、Dropout、重叠池化。
这里把 11×11 大卷积核按 28×28 输入做了等比例适配。
运行方式:
uv run python examples/ch05_alexnet.py
"""
import torch
from torch import nn
from torch.utils.data import DataLoader
from torchvision import datasets, transforms
torch.manual_seed(0)
device = "cuda" if torch.cuda.is_available() else "cpu"
transform = transforms.Compose([
transforms.ToTensor(),
transforms.Normalize((0.2860,), (0.3530,)),
])
train = datasets.FashionMNIST("data", train=True, download=True, transform=transform)
test = datasets.FashionMNIST("data", train=False, download=True, transform=transform)
train_loader = DataLoader(train, batch_size=256, shuffle=True)
test_loader = DataLoader(test, batch_size=1024)
class AlexNetStyle(nn.Module):
def __init__(self):
super().__init__()
self.features = nn.Sequential(
nn.Conv2d(1, 64, 5, padding=2), nn.ReLU(), nn.MaxPool2d(2), # 28→14
nn.Conv2d(64, 192, 5, padding=2), nn.ReLU(), nn.MaxPool2d(2), # 14→7
nn.Conv2d(192, 384, 3, padding=1), nn.ReLU(),
nn.Conv2d(384, 256, 3, padding=1), nn.ReLU(),
nn.Conv2d(256, 256, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2), # 7→3
)
self.classifier = nn.Sequential(
nn.Flatten(),
nn.Linear(256 * 3 * 3, 2048), nn.ReLU(), nn.Dropout(0.5),
nn.Linear(2048, 2048), nn.ReLU(), nn.Dropout(0.5),
nn.Linear(2048, 10),
)
def forward(self, x):
return self.classifier(self.features(x))
model = AlexNetStyle().to(device)
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
criterion = nn.CrossEntropyLoss()
print(f"AlexNet 风格参数量:{sum(p.numel() for p in model.parameters()):,}")
def evaluate(loader):
model.eval()
correct = total = 0
with torch.no_grad():
for xb, yb in loader:
pred = model(xb.to(device)).argmax(1)
correct += (pred == yb.to(device)).sum().item()
total += len(yb)
return correct / total
for epoch in range(3):
model.train()
for xb, yb in train_loader:
xb, yb = xb.to(device), yb.to(device)
opt.zero_grad()
loss = criterion(model(xb), yb)
loss.backward()
opt.step()
print(f"epoch {epoch + 1}:训练准确率 = {evaluate(train_loader) * 100:.2f}%")
print(f"测试准确率 = {evaluate(test_loader) * 100:.2f}%")examples/ch05_vgg.py
"""第 5 章示例(2012—2014):小型 VGG 风格网络在 Fashion-MNIST 上训练。
VGG 的核心思想是"小卷积核、深堆叠":用多个 3×3 卷积代替大卷积核。
运行方式:
uv run python examples/ch05_vgg.py
"""
import torch
from torch import nn
from torch.utils.data import DataLoader
from torchvision import datasets, transforms
torch.manual_seed(0)
device = "cuda" if torch.cuda.is_available() else "cpu"
transform = transforms.Compose([
transforms.ToTensor(),
transforms.Normalize((0.2860,), (0.3530,)),
])
train = datasets.FashionMNIST("data", train=True, download=True, transform=transform)
test = datasets.FashionMNIST("data", train=False, download=True, transform=transform)
train_loader = DataLoader(train, batch_size=256, shuffle=True)
test_loader = DataLoader(test, batch_size=1024)
def conv_block(in_c, out_c, reps=2):
layers = []
for _ in range(reps):
layers += [nn.Conv2d(in_c, out_c, 3, padding=1), nn.BatchNorm2d(out_c), nn.ReLU()]
in_c = out_c
layers.append(nn.MaxPool2d(2))
return nn.Sequential(*layers)
class TinyVGG(nn.Module):
def __init__(self):
super().__init__()
self.features = nn.Sequential(
conv_block(1, 64), # 28→14
conv_block(64, 128), # 14→7
conv_block(128, 256, reps=3), # 7→3
)
self.classifier = nn.Sequential(
nn.Flatten(),
nn.Linear(256 * 3 * 3, 512),
nn.ReLU(),
nn.Dropout(0.5),
nn.Linear(512, 10),
)
def forward(self, x):
return self.classifier(self.features(x))
model = TinyVGG().to(device)
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
criterion = nn.CrossEntropyLoss()
print(f"TinyVGG 参数量:{sum(p.numel() for p in model.parameters()):,}")
def evaluate(loader):
model.eval()
correct = total = 0
with torch.no_grad():
for xb, yb in loader:
pred = model(xb.to(device)).argmax(1)
correct += (pred == yb.to(device)).sum().item()
total += len(yb)
return correct / total
for epoch in range(3):
model.train()
for xb, yb in train_loader:
xb, yb = xb.to(device), yb.to(device)
opt.zero_grad()
loss = criterion(model(xb), yb)
loss.backward()
opt.step()
print(f"epoch {epoch + 1}:训练准确率 = {evaluate(train_loader) * 100:.2f}%")
print(f"测试准确率 = {evaluate(test_loader) * 100:.2f}%")examples/ch05_inception.py
"""第 5 章示例(2014):GoogLeNet 的 Inception 模块。
Inception 在同一层并行使用 1×1、3×3、5×5 卷积和池化,再把结果拼接,
让网络自己"挑"合适的感受野;1×1 卷积先降维,控制计算量。
运行方式:
uv run python examples/ch05_inception.py
"""
import torch
from torch import nn
from torch.utils.data import DataLoader
from torchvision import datasets, transforms
torch.manual_seed(0)
device = "cuda" if torch.cuda.is_available() else "cpu"
transform = transforms.Compose([
transforms.ToTensor(),
transforms.Normalize((0.2860,), (0.3530,)),
])
train = datasets.FashionMNIST("data", train=True, download=True, transform=transform)
test = datasets.FashionMNIST("data", train=False, download=True, transform=transform)
train_loader = DataLoader(train, batch_size=256, shuffle=True)
test_loader = DataLoader(test, batch_size=1024)
class Inception(nn.Module):
"""四条分支并行,输出通道拼接。"""
def __init__(self, in_c):
super().__init__()
self.b1 = nn.Sequential(nn.Conv2d(in_c, 32, 1), nn.ReLU())
self.b2 = nn.Sequential(
nn.Conv2d(in_c, 32, 1), nn.ReLU(),
nn.Conv2d(32, 48, 3, padding=1), nn.ReLU(),
)
self.b3 = nn.Sequential(
nn.Conv2d(in_c, 16, 1), nn.ReLU(),
nn.Conv2d(16, 32, 5, padding=2), nn.ReLU(),
)
self.b4 = nn.Sequential(
nn.MaxPool2d(3, stride=1, padding=1),
nn.Conv2d(in_c, 16, 1), nn.ReLU(),
)
def forward(self, x):
return torch.cat([self.b1(x), self.b2(x), self.b3(x), self.b4(x)], dim=1)
class TinyGoogLeNet(nn.Module):
def __init__(self):
super().__init__()
self.stem = nn.Sequential(nn.Conv2d(1, 32, 3, padding=1), nn.ReLU())
self.blocks = nn.Sequential(
Inception(32), # 32→128
nn.MaxPool2d(2),
Inception(128),
nn.MaxPool2d(2),
Inception(128),
nn.AdaptiveAvgPool2d(1),
)
self.fc = nn.Linear(128, 10)
def forward(self, x):
return self.fc(self.blocks(self.stem(x)).flatten(1))
model = TinyGoogLeNet().to(device)
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
criterion = nn.CrossEntropyLoss()
print(f"Inception 风格参数量:{sum(p.numel() for p in model.parameters()):,}")
def evaluate(loader):
model.eval()
correct = total = 0
with torch.no_grad():
for xb, yb in loader:
pred = model(xb.to(device)).argmax(1)
correct += (pred == yb.to(device)).sum().item()
total += len(yb)
return correct / total
for epoch in range(3):
model.train()
for xb, yb in train_loader:
xb, yb = xb.to(device), yb.to(device)
opt.zero_grad()
loss = criterion(model(xb), yb)
loss.backward()
opt.step()
print(f"epoch {epoch + 1}:训练准确率 = {evaluate(train_loader) * 100:.2f}%")
print(f"测试准确率 = {evaluate(test_loader) * 100:.2f}%")动手实践
- 分别运行三个示例脚本,比较参数量与测试准确率,用一句话总结「效率」与「精度」的权衡。
- 在
examples/ch05_alexnet.py中把 Dropout 概率改为 0(禁用),训练 3 轮,观察训练 / 测试准确率差距(过拟合程度)。 - 在 Inception 模块中去掉 5×5 分支,观察参数量和准确率变化。
- 把 TinyVGG 的
reps=2改为reps=3,记录参数量与准确率,验证「更深更好」是否在小数据集上成立。
常见错误
| 错误写法 | 现象 | 原因 |
|---|---|---|
| 输入尺寸与论文不一致直接照抄 | 前向报 shape 错误 | AlexNet 输入是 227×227;换数据集要重算各层尺寸 |
| Inception 各分支输出通道没算清 | 拼接后维度不对 | torch.cat(..., dim=1) 要求除通道外尺寸一致 |
| 全连接层维度用论文原值 | mat shapes cannot be multiplied | 特征图尺寸变了,flatten 后维度就变了 |
| 只训练 1 个 epoch 就对比模型 | 结论不稳定 | 对比架构要用相同 epoch 数并多次取平均 |
| Dropout 在 eval 时仍生效 | 测试准确率异常低 | 记得 model.eval(),它会关闭 dropout |
章末练习
基础
- AlexNet 的三个「新武器」分别是什么?
- 为什么两个 3×3 卷积等价于一个 5×5 卷积(感受野),但参数更少?
- Inception 模块并行了几条分支?1×1 卷积的作用是什么?
提高
- 计算 VGG 一个 3×3 卷积块(2 个卷积,通道 C→C)的参数,与一个 5×5 卷积(C→C)比较。
- 把 AlexNet 风格网络的三个全连接层改成一个,观察参数量与准确率。
- 用 3 个随机种子分别训练 TinyVGG,报告准确率均值与标准差。
挑战
- 实现一个「深度可分离」变体:把 Inception 的 3×3 分支换成 3×3 深度卷积 + 1×1 逐点卷积,对比参数与准确率(为第 10 章 EfficientNet 铺垫)。
- 实现 11 层 VGG(conv 块通道 64/128/256/256 + 两个全连接),在 Fashion-MNIST 上训练并报告完整准确率曲线。
章末自测
- AlexNet 是在什么比赛中夺冠的?
- A. Kaggle
- B. ImageNet ILSVRC
- C. CIFAR
- D. COCO
- AlexNet 使用哪种激活函数?
- A. Sigmoid
- B. tanh
- C. ReLU
- D. ELU
- VGG 的核心思想是?
- A. 大卷积核
- B. 小卷积核深堆叠
- C. 无池化
- D. 全连接分类
- 两个 3×3 卷积的感受野等价于?
- A. 3×3
- B. 5×5
- C. 6×6
- D. 9×9
- Inception 模块中 1×1 卷积的首要作用是?
- A. 增加深度
- B. 降维省计算
- C. 替代池化
- D. 加速 GPU
- Dropout 的作用是?
- A. 加速训练
- B. 缓解过拟合
- C. 增加参数量
- D. 数据增强
- VGG 提出于哪一年?
- A. 2012
- B. 2014
- C. 2015
- D. 2016
- GoogLeNet 的 Inception 模块把各分支结果?
- A. 相加
- B. 拼接
- C. 平均
- D. 相乘
- AlexNet 参数量的大头在?
- A. 卷积层
- B. 全连接层
- C. 池化层
- D. 归一化层
- 2012 年 AlexNet 夺冠的 top-5 错误率约为?
- A. 26%
- B. 15.3%
- C. 7%
- D. 3%
