Skip to content

第 5 章 2012—2014:深度学习元年——AlexNet、VGG 与 GoogLeNet ​

学习目标 ​

  • 理解 AlexNet(2012)为什么是「深度学习元年」的引爆点
  • 理解 VGG(2014)「小卷积核深堆叠」的思想
  • 理解 GoogLeNet / Inception(2014)的多尺度并行思想
  • 分别用 PyTorch 实现并训练三种经典 CNN 架构

5.1 时代背景:ImageNet 与 GPU ​

2012 年 10 月,Alex Krizhevsky、Ilya Sutskever 和 Geoffrey Hinton 提交了 AlexNet 论文《ImageNet Classification with Deep Convolutional Neural Networks》。它在 ImageNet 大赛上把 top-5 错误率从 26% 压到 15.3%,碾压第二名——而它靠的不只是网络结构,还有两件「新武器」:

  • GPU:两块 GTX 580 并行训练,卷积可以在 GPU 上批量计算;
  • 数据:ImageNet 有 120 万张标注图片,足够训练深层网络。

再加上 ReLU(解决梯度消失)、Dropout(缓解过拟合)、数据增强(扩充训练集),深度学习从边缘进入主流。2014 年,VGG 和 GoogLeNet 把错误率进一步降到 7% 以下。

5.2 AlexNet:大核、深堆叠、ReLU ​

AlexNet 的配方(输入 227×227):

层操作输出
Conv196 个 11×11 卷积,stride 455×55×96
Pool13×3 最大池化,stride 227×27×96
Conv2256 个 5×5 卷积27×27×256
Pool23×3 池化13×13×256
Conv3-53×3 / 3×3 / 3×313×13×256
Pool53×3 池化6×6×256
FC4096 → 4096 → 10001000

三个关键点:全部用 ReLU(训练比 tanh 快数倍且不饱和)、Dropout(随机丢弃一半神经元,防止全连接层过拟合)、重叠池化(池化窗口 3×3、步长 2,略降错误率)。

examples/ch05_alexnet.py 把 AlexNet 适配到 28×28 的 Fashion-MNIST(卷积核等比例缩小):

AlexNet 风格参数量:11,385,546
epoch 1:训练准确率 = 87.21%
epoch 2:训练准确率 = 89.11%
epoch 3:训练准确率 = 91.73%
测试准确率 = 90.22%

注意参数量高达 1138 万,绝大部分在最后三个全连接层。这是 2012 年的常态——全连接层是大头,也是后来被批评浪费的地方。

5.3 VGG:小卷积核的堆叠艺术 ​

2014 年,牛津大学的 Karen Simonyan 和 Andrew Zisserman 提出 VGG,核心观察:两个 3×3 卷积等价于一个 5×5 卷积的感受野,但参数更少、非线性更多。

一个 5×5 卷积有 25 个参数,两个 3×3 有 18 个;同时多了一次 ReLU,表达力更强。于是 VGG 把 AlexNet 的大核全部换成 3×3,靠深度(16~19 层)取胜:VGG16 在 ImageNet 上达到 7.3% top-5 错误率。

examples/ch05_vgg.py 用「卷积块(2~3 个 3×3 卷积 + 池化)」搭建 TinyVGG:

python
def conv_block(in_c, out_c, reps=2):
    layers = []
    for _ in range(reps):
        layers += [nn.Conv2d(in_c, out_c, 3, padding=1), nn.BatchNorm2d(out_c), nn.ReLU()]
        in_c = out_c
    layers.append(nn.MaxPool2d(2))
    return nn.Sequential(*layers)
TinyVGG 参数量:2,921,930
epoch 1:训练准确率 = 89.77%
epoch 2:训练准确率 = 91.42%
epoch 3:训练准确率 = 93.72%
测试准确率 = 91.83%

注意 TinyVGG 只有 AlexNet 风格网络 1/4 的参数,准确率却更高(91.8% vs 90.2%)——小卷积核 + 批量归一化(BN 见第 6 章)的威力。

5.4 GoogLeNet / Inception:一层干多种活 ​

2014 年的另一个冠军是 Google 的 GoogLeNet(向 LeNet 致敬的名字)。它的核心模块 Inception 在同一层并行跑四种操作:1×1 卷积、3×3 卷积、5×5 卷积、3×3 池化,然后把结果拼接起来。这样一层就能同时捕捉不同尺度的特征,让网络自己「挑」。

1×1 卷积是 Inception 的秘密武器:先把通道数压缩(降维),再做大核卷积,计算量大幅下降。GoogLeNet 只有 AlexNet 的 1/12 参数,却拿了 2014 年 ImageNet 冠军。

examples/ch05_inception.py 实现 Inception 模块:

python
class Inception(nn.Module):
    def __init__(self, in_c):
        super().__init__()
        self.b1 = nn.Sequential(nn.Conv2d(in_c, 32, 1), nn.ReLU())
        self.b2 = nn.Sequential(
            nn.Conv2d(in_c, 32, 1), nn.ReLU(),
            nn.Conv2d(32, 48, 3, padding=1), nn.ReLU(),
        )
        self.b3 = nn.Sequential(
            nn.Conv2d(in_c, 16, 1), nn.ReLU(),
            nn.Conv2d(16, 32, 5, padding=2), nn.ReLU(),
        )
        self.b4 = nn.Sequential(
            nn.MaxPool2d(3, stride=1, padding=1),
            nn.Conv2d(in_c, 16, 1), nn.ReLU(),
        )

    def forward(self, x):
        return torch.cat([self.b1(x), self.b2(x), self.b3(x), self.b4(x)], dim=1)
Inception 风格参数量:109,658
epoch 1:训练准确率 = 76.39%
epoch 2:训练准确率 = 79.64%
epoch 3:训练准确率 = 83.12%
测试准确率 = 82.26%

只有 11 万参数就能达到 82%,Inception 的效率优势明显。准确率比 VGG 低是因为 Fashion-MNIST 太小、3 个 epoch 太短——在 ImageNet 这种大规模数据集上,它的精度才是第一梯队。

5.5 三条路线的启示 ​

2012—2014 年给出了 CNN 设计的三个方向,今天仍在使用:

  • AlexNet 路线:更大的模型 + 更强的正则化;
  • VGG 路线:更小的核、更深的堆叠;
  • GoogLeNet 路线:多尺度并行 + 1×1 降维。

它们共同奠定了「卷积骨干 + 分类头」的标准范式。下一步的瓶颈是深度:当网络深到 20 层以上,训练开始退化——这正是第 6 章 ResNet 要解决的问题。

完整代码 ​

本章用到的完整示例代码:

examples/ch05_alexnet.py ​

py
"""第 5 章示例(2012):AlexNet 风格网络在 Fashion-MNIST 上训练。

AlexNet 的配方:5 层卷积 + 3 层全连接、ReLU、Dropout、重叠池化。
这里把 11×11 大卷积核按 28×28 输入做了等比例适配。

运行方式:
    uv run python examples/ch05_alexnet.py
"""
import torch
from torch import nn
from torch.utils.data import DataLoader
from torchvision import datasets, transforms

torch.manual_seed(0)
device = "cuda" if torch.cuda.is_available() else "cpu"

transform = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.2860,), (0.3530,)),
])
train = datasets.FashionMNIST("data", train=True, download=True, transform=transform)
test = datasets.FashionMNIST("data", train=False, download=True, transform=transform)
train_loader = DataLoader(train, batch_size=256, shuffle=True)
test_loader = DataLoader(test, batch_size=1024)


class AlexNetStyle(nn.Module):
    def __init__(self):
        super().__init__()
        self.features = nn.Sequential(
            nn.Conv2d(1, 64, 5, padding=2), nn.ReLU(), nn.MaxPool2d(2),      # 28→14
            nn.Conv2d(64, 192, 5, padding=2), nn.ReLU(), nn.MaxPool2d(2),    # 14→7
            nn.Conv2d(192, 384, 3, padding=1), nn.ReLU(),
            nn.Conv2d(384, 256, 3, padding=1), nn.ReLU(),
            nn.Conv2d(256, 256, 3, padding=1), nn.ReLU(), nn.MaxPool2d(2),   # 7→3
        )
        self.classifier = nn.Sequential(
            nn.Flatten(),
            nn.Linear(256 * 3 * 3, 2048), nn.ReLU(), nn.Dropout(0.5),
            nn.Linear(2048, 2048), nn.ReLU(), nn.Dropout(0.5),
            nn.Linear(2048, 10),
        )

    def forward(self, x):
        return self.classifier(self.features(x))


model = AlexNetStyle().to(device)
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
criterion = nn.CrossEntropyLoss()
print(f"AlexNet 风格参数量:{sum(p.numel() for p in model.parameters()):,}")


def evaluate(loader):
    model.eval()
    correct = total = 0
    with torch.no_grad():
        for xb, yb in loader:
            pred = model(xb.to(device)).argmax(1)
            correct += (pred == yb.to(device)).sum().item()
            total += len(yb)
    return correct / total


for epoch in range(3):
    model.train()
    for xb, yb in train_loader:
        xb, yb = xb.to(device), yb.to(device)
        opt.zero_grad()
        loss = criterion(model(xb), yb)
        loss.backward()
        opt.step()
    print(f"epoch {epoch + 1}:训练准确率 = {evaluate(train_loader) * 100:.2f}%")

print(f"测试准确率 = {evaluate(test_loader) * 100:.2f}%")

examples/ch05_vgg.py ​

py
"""第 5 章示例(2012—2014):小型 VGG 风格网络在 Fashion-MNIST 上训练。

VGG 的核心思想是"小卷积核、深堆叠":用多个 3×3 卷积代替大卷积核。

运行方式:
    uv run python examples/ch05_vgg.py
"""
import torch
from torch import nn
from torch.utils.data import DataLoader
from torchvision import datasets, transforms

torch.manual_seed(0)
device = "cuda" if torch.cuda.is_available() else "cpu"

transform = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.2860,), (0.3530,)),
])
train = datasets.FashionMNIST("data", train=True, download=True, transform=transform)
test = datasets.FashionMNIST("data", train=False, download=True, transform=transform)
train_loader = DataLoader(train, batch_size=256, shuffle=True)
test_loader = DataLoader(test, batch_size=1024)


def conv_block(in_c, out_c, reps=2):
    layers = []
    for _ in range(reps):
        layers += [nn.Conv2d(in_c, out_c, 3, padding=1), nn.BatchNorm2d(out_c), nn.ReLU()]
        in_c = out_c
    layers.append(nn.MaxPool2d(2))
    return nn.Sequential(*layers)


class TinyVGG(nn.Module):
    def __init__(self):
        super().__init__()
        self.features = nn.Sequential(
            conv_block(1, 64),      # 28→14
            conv_block(64, 128),    # 14→7
            conv_block(128, 256, reps=3),  # 7→3
        )
        self.classifier = nn.Sequential(
            nn.Flatten(),
            nn.Linear(256 * 3 * 3, 512),
            nn.ReLU(),
            nn.Dropout(0.5),
            nn.Linear(512, 10),
        )

    def forward(self, x):
        return self.classifier(self.features(x))


model = TinyVGG().to(device)
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
criterion = nn.CrossEntropyLoss()
print(f"TinyVGG 参数量:{sum(p.numel() for p in model.parameters()):,}")


def evaluate(loader):
    model.eval()
    correct = total = 0
    with torch.no_grad():
        for xb, yb in loader:
            pred = model(xb.to(device)).argmax(1)
            correct += (pred == yb.to(device)).sum().item()
            total += len(yb)
    return correct / total


for epoch in range(3):
    model.train()
    for xb, yb in train_loader:
        xb, yb = xb.to(device), yb.to(device)
        opt.zero_grad()
        loss = criterion(model(xb), yb)
        loss.backward()
        opt.step()
    print(f"epoch {epoch + 1}:训练准确率 = {evaluate(train_loader) * 100:.2f}%")

print(f"测试准确率 = {evaluate(test_loader) * 100:.2f}%")

examples/ch05_inception.py ​

py
"""第 5 章示例(2014):GoogLeNet 的 Inception 模块。

Inception 在同一层并行使用 1×1、3×3、5×5 卷积和池化,再把结果拼接,
让网络自己"挑"合适的感受野;1×1 卷积先降维,控制计算量。

运行方式:
    uv run python examples/ch05_inception.py
"""
import torch
from torch import nn
from torch.utils.data import DataLoader
from torchvision import datasets, transforms

torch.manual_seed(0)
device = "cuda" if torch.cuda.is_available() else "cpu"

transform = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize((0.2860,), (0.3530,)),
])
train = datasets.FashionMNIST("data", train=True, download=True, transform=transform)
test = datasets.FashionMNIST("data", train=False, download=True, transform=transform)
train_loader = DataLoader(train, batch_size=256, shuffle=True)
test_loader = DataLoader(test, batch_size=1024)


class Inception(nn.Module):
    """四条分支并行,输出通道拼接。"""
    def __init__(self, in_c):
        super().__init__()
        self.b1 = nn.Sequential(nn.Conv2d(in_c, 32, 1), nn.ReLU())
        self.b2 = nn.Sequential(
            nn.Conv2d(in_c, 32, 1), nn.ReLU(),
            nn.Conv2d(32, 48, 3, padding=1), nn.ReLU(),
        )
        self.b3 = nn.Sequential(
            nn.Conv2d(in_c, 16, 1), nn.ReLU(),
            nn.Conv2d(16, 32, 5, padding=2), nn.ReLU(),
        )
        self.b4 = nn.Sequential(
            nn.MaxPool2d(3, stride=1, padding=1),
            nn.Conv2d(in_c, 16, 1), nn.ReLU(),
        )

    def forward(self, x):
        return torch.cat([self.b1(x), self.b2(x), self.b3(x), self.b4(x)], dim=1)


class TinyGoogLeNet(nn.Module):
    def __init__(self):
        super().__init__()
        self.stem = nn.Sequential(nn.Conv2d(1, 32, 3, padding=1), nn.ReLU())
        self.blocks = nn.Sequential(
            Inception(32),        # 32→128
            nn.MaxPool2d(2),
            Inception(128),
            nn.MaxPool2d(2),
            Inception(128),
            nn.AdaptiveAvgPool2d(1),
        )
        self.fc = nn.Linear(128, 10)

    def forward(self, x):
        return self.fc(self.blocks(self.stem(x)).flatten(1))


model = TinyGoogLeNet().to(device)
opt = torch.optim.Adam(model.parameters(), lr=1e-3)
criterion = nn.CrossEntropyLoss()
print(f"Inception 风格参数量:{sum(p.numel() for p in model.parameters()):,}")


def evaluate(loader):
    model.eval()
    correct = total = 0
    with torch.no_grad():
        for xb, yb in loader:
            pred = model(xb.to(device)).argmax(1)
            correct += (pred == yb.to(device)).sum().item()
            total += len(yb)
    return correct / total


for epoch in range(3):
    model.train()
    for xb, yb in train_loader:
        xb, yb = xb.to(device), yb.to(device)
        opt.zero_grad()
        loss = criterion(model(xb), yb)
        loss.backward()
        opt.step()
    print(f"epoch {epoch + 1}:训练准确率 = {evaluate(train_loader) * 100:.2f}%")

print(f"测试准确率 = {evaluate(test_loader) * 100:.2f}%")

动手实践 ​

  1. 分别运行三个示例脚本,比较参数量与测试准确率,用一句话总结「效率」与「精度」的权衡。
  2. 在 examples/ch05_alexnet.py 中把 Dropout 概率改为 0(禁用),训练 3 轮,观察训练 / 测试准确率差距(过拟合程度)。
  3. 在 Inception 模块中去掉 5×5 分支,观察参数量和准确率变化。
  4. 把 TinyVGG 的 reps=2 改为 reps=3,记录参数量与准确率,验证「更深更好」是否在小数据集上成立。

常见错误 ​

错误写法现象原因
输入尺寸与论文不一致直接照抄前向报 shape 错误AlexNet 输入是 227×227;换数据集要重算各层尺寸
Inception 各分支输出通道没算清拼接后维度不对torch.cat(..., dim=1) 要求除通道外尺寸一致
全连接层维度用论文原值mat shapes cannot be multiplied特征图尺寸变了,flatten 后维度就变了
只训练 1 个 epoch 就对比模型结论不稳定对比架构要用相同 epoch 数并多次取平均
Dropout 在 eval 时仍生效测试准确率异常低记得 model.eval(),它会关闭 dropout

章末练习 ​

基础

  1. AlexNet 的三个「新武器」分别是什么?
  2. 为什么两个 3×3 卷积等价于一个 5×5 卷积(感受野),但参数更少?
  3. Inception 模块并行了几条分支?1×1 卷积的作用是什么?

提高

  1. 计算 VGG 一个 3×3 卷积块(2 个卷积,通道 C→C)的参数,与一个 5×5 卷积(C→C)比较。
  2. 把 AlexNet 风格网络的三个全连接层改成一个,观察参数量与准确率。
  3. 用 3 个随机种子分别训练 TinyVGG,报告准确率均值与标准差。

挑战

  1. 实现一个「深度可分离」变体:把 Inception 的 3×3 分支换成 3×3 深度卷积 + 1×1 逐点卷积,对比参数与准确率(为第 10 章 EfficientNet 铺垫)。
  2. 实现 11 层 VGG(conv 块通道 64/128/256/256 + 两个全连接),在 Fashion-MNIST 上训练并报告完整准确率曲线。

章末自测 ​

  1. AlexNet 是在什么比赛中夺冠的?
    • A. Kaggle
    • B. ImageNet ILSVRC
    • C. CIFAR
    • D. COCO
  2. AlexNet 使用哪种激活函数?
    • A. Sigmoid
    • B. tanh
    • C. ReLU
    • D. ELU
  3. VGG 的核心思想是?
    • A. 大卷积核
    • B. 小卷积核深堆叠
    • C. 无池化
    • D. 全连接分类
  4. 两个 3×3 卷积的感受野等价于?
    • A. 3×3
    • B. 5×5
    • C. 6×6
    • D. 9×9
  5. Inception 模块中 1×1 卷积的首要作用是?
    • A. 增加深度
    • B. 降维省计算
    • C. 替代池化
    • D. 加速 GPU
  6. Dropout 的作用是?
    • A. 加速训练
    • B. 缓解过拟合
    • C. 增加参数量
    • D. 数据增强
  7. VGG 提出于哪一年?
    • A. 2012
    • B. 2014
    • C. 2015
    • D. 2016
  8. GoogLeNet 的 Inception 模块把各分支结果?
    • A. 相加
    • B. 拼接
    • C. 平均
    • D. 相乘
  9. AlexNet 参数量的大头在?
    • A. 卷积层
    • B. 全连接层
    • C. 池化层
    • D. 归一化层
  10. 2012 年 AlexNet 夺冠的 top-5 错误率约为?
    • A. 26%
    • B. 15.3%
    • C. 7%
    • D. 3%