Hand-computed Cross-Entropy Loss
Construct a prediction sequence from "random guessing" to "very confident", and observe how the cross-entropy value decreases accordingly.
After completing this case study, you will understand:Why classification tasks use cross-entropy instead of MSE — it punishes confident but wrong predictions more "severely" than MSE.
Real-Life Introduction
Exam Grading — Wrong Answers Are Worse Than No Answer
Two students' performance on a single-choice question (answer is A): Xiaohong predicts probability 0.5 for A (uncertain but correct direction), Xiaoming predicts probability 0.95 for C (very confidently wrong). Xiaohong is uncertain but directionally correct; Xiaoming is not only wrong but also very confident — such a mistake is more unforgivable.
Cross-entropy loss is like this grading standard:Reward confidence when predictions are correct, punish confidence when predictions are wrong.
Intuitive Understanding
The true label is "cat" (one-hot = [1, 0, 0]). Simulate the model's output from "random guessing" to "getting more accurate with training":
- Completely random guessing [0.33, 0.33, 0.34] → loss 1.10
- Very confident and correct [0.98, 0.01, 0.01] → loss 0.02
- Very confident but wrong [0.02, 0.9, 0.08] → loss 3.91 (nearly 4 times higher than random guessing!)
The last value deserves attention: because -log(0.02) ≈ 3.91, while -log(0.33) ≈ 1.10.
Mathematical Definition
\[ H(p, q) = -\sum_{i=1}^{C} p_i \cdot \log(q_i) \]\(p\) is the true distribution (one-hot), \(q\) is the predicted distribution. For a "confident but wrong" prediction, cross-entropy = -log(0.02) ≈ 3.91, MSE = (1-0.02)^2 ≈ 0.96. The penalty of cross-entropy is 12 times that of MSE.
Python Hands-On Practice
Example
def cross_entropy(p_true, q_pred, eps=1e-12):
q_pred = np.clip(q_pred, eps, 1 - eps)
return -np.sum(p_true * np.log(q_pred))
# True label: cat (class 0)
p_true = np.array([1, 0, 0]) # [cat, dog, bird]
predictions = {
"Completely random guessing (uniform)": np.array([0.33, 0.33, 0.34]),
"Slight feeling": np.array([0.5, 0.3, 0.2]),
"Gradually becoming reliable": np.array([0.7, 0.2, 0.1]),
"Better trained": np.array([0.9, 0.07, 0.03]),
"Very confident and correct": np.array([0.98, 0.01, 0.01]),
"Very confident but wrong!": np.array([0.02, 0.9, 0.08]),
}
print(EXAMPLE Cross-Entropy Loss Comparison (True: cat=[1,0,0])\n")
print(f"{'Prediction Description':<20} {'Loss':<10} {'Analysis'}")
print("-" * 52)
for name, q in predictions.items():
loss = cross_entropy(p_true, q)
if loss > 2: analysis = "Critical error!"
elif loss > 0.5: analysis = "Directionally correct but uncertain"
else: analysis = "Very good"
print(f"{name:<20} {loss:<10.4f} {analysis}")
# MSE vs Cross-Entropy Comparison
q_bad = predictions["Very confident but wrong!"]
mse_bad = np.mean((p_true - q_bad) ** 2)
ce_bad = cross_entropy(p_true, q_bad)
print(f"\nEXAMPLE Confident but wrong: MSE={mse_bad:.4f}, Cross-Entropy={ce_bad:.4f} (harsh {ce_bad/mse_bad:.0f}x))
EXAMPLE 交叉熵损失对比 (真实: 猫=[1,0,0]) 预测描述 损失 分析 ---------------------------------------------------- 完全瞎猜 (均匀) 1.0986 方向对但不确定 有点感觉 0.6931 方向对但不确定 逐渐靠谱 0.3567 很好 训练较好 0.1054 很好 非常自信且正确 0.0202 很好 非常自信但错误! 3.9120 严重错误! EXAMPLE 自信但错误时: MSE=0.3260, 交叉熵=3.9120 (严厉 12x)
Application Scenarios in AI
| Scenario | Loss function used |
|---|---|
| Image classification | Multi-class cross-entropy (CategoricalCrossEntropy) |
| Language model | Sum of cross-entropy at each token position — the objective function for GPT training |
| Knowledge distillation | Use the teacher model's soft labels for cross-entropy |