把 GPT 这类大语言模型(LLM)的所有魔法剥掉,核心任务只有一个:给一串词,算出下一个词是词表里每个词的概率。看到「The cat sat on the」,它算出 mat 的概率比 moon 高得多。
Vizuara 原文 1.1 节用真实 GPT-2 做了个实验:输入「After years of hard work, your effort will take you」,模型对下一个词给出前 10 名候选——「to」以 90.7% 遥遥领先,其余候选概率依次递减。注意关键点:模型给的不是一个答案,而是一整张概率表。所以 LLM 常被叫做概率引擎——它不"知道"正确答案,它只是极其擅长估计"什么词接在这里最像话"。
# ============================================================
# 【演示目标】亲眼看到"下一词预测模型"输出的东西 = 一张概率表。
# 这是 LLM 最核心的一件事:正文说"模型给的不是答案,是概率表",
# 这块代码用 1950 年代水平的玩具模型把它变成可跑的现实。
# 【思路】分四步,每步对应真实 LLM 的一个环节:
# 1. 训练 → 数语料里每对相邻词的出现次数,归一化成概率表
# (真实 LLM 用神经网络"学"这张表,我们直接数——但产物同款)
# 2. 预测 → 给一句话,查表报出下一个词的 Top-5 候选和百分比
# 3. 采样 → 按概率抽一个词(所以每次运行结果可能不同!)
# 4. 自回归 → 把抽到的词拼回句尾,回到第 2 步循环
# ============================================================
import numpy as np
# ---- 第 1 步·准备:语料与词表 ----
corpus = ( # 8 个玩具短句;故意让 the cat /
"the cat sat on the mat . the cat ate the fish . " # the mat 这类搭配高频,
"the dog sat on the floor . the dog ate the bone . " # 看概率表能否"学"到
"a cat and a dog sat together . the fish swam in the water . "
"the mat was under the cat . the bone was on the floor ."
).split() # 按空格切开 → 词的列表
vocab = sorted(set(corpus)) # 词表 = 语料中出现过的所有词(去重)
w2i = {w: i for i, w in enumerate(vocab)} # 词 → 行号:用矩阵位置存"谁后面跟谁"
n = len(vocab) # 玩具词表只有 19 个词
# ---- 第 1 步·训练:数相邻词对 → 概率表 ----
counts = np.zeros((n, n)) # counts[i][j] = 词 i 后面是词 j 的次数
for a, b in zip(corpus, corpus[1:]): # zip 错开一位配对:(第1,第2)(第2,第3)…
counts[w2i[a], w2i[b]] += 1
rows = counts.sum(axis=1, keepdims=True) # 每行的总次数(词 i 后面接任意词)
probs = np.divide(counts, rows, out=np.zeros_like(counts), where=rows > 0)
# ↑ 每行除以行和 → 行里每个格子变成概率,整行加起来 = 100%
# ---- 第 2 步·预测:给定句子,报 Top-5 候选 ----
prompt = "the cat sat on the" # 改我!用小写英文词,别超纲(词表里没有就不行)
last = prompt.lower().split()[-1] # 玩具模型只看最后一个词(真实 LLM 看全部上文)
if last not in w2i:
print("「%s」不在玩具语料里。试试 the / cat / sat / fish ..." % last)
else:
i = w2i[last] # 查到"最后一个词"对应的行
print("输入: %s" % prompt)
print("「%s」后面的 Top-5 候选词:" % last)
for j in np.argsort(probs[i])[::-1][:5]: # 该行按概率从大到小排序,取前 5
if probs[i, j] > 0:
print(" %-8s %.0f%%" % (vocab[j], probs[i, j] * 100))
# ---- 第 3+4 步·采样 + 自回归生成 ----
# np.random.choice 按 probs 这一行概率抽词:21% 抽中 cat,14% 抽中 mat……
# 抽到什么拼什么,拼完再查表——这就是正文说的"预测→拼接→循环"
out = prompt.lower().split()
for _ in range(12): # 最多再生成 12 个词(碰到句号提前停)
nxt = vocab[np.random.choice(n, p=probs[w2i[out[-1]]])]
out.append(nxt)
if nxt == ".":
break
print("自回归生成:", " ".join(out)) # 多跑几次看不同——这就是"采样"