Hand-computed Cross-Entropy Loss

Construct a prediction sequence from "random guessing" to "very confident", and observe how the cross-entropy value decreases accordingly.

After completing this case study, you will understand:Why classification tasks use cross-entropy instead of MSE — it punishes confident but wrong predictions more "severely" than MSE.


Real-Life Introduction

Exam Grading — Wrong Answers Are Worse Than No Answer

Two students' performance on a single-choice question (answer is A): Xiaohong predicts probability 0.5 for A (uncertain but correct direction), Xiaoming predicts probability 0.95 for C (very confidently wrong). Xiaohong is uncertain but directionally correct; Xiaoming is not only wrong but also very confident — such a mistake is more unforgivable.

Cross-entropy loss is like this grading standard:Reward confidence when predictions are correct, punish confidence when predictions are wrong.


Intuitive Understanding

The true label is "cat" (one-hot = [1, 0, 0]). Simulate the model's output from "random guessing" to "getting more accurate with training":

  • Completely random guessing [0.33, 0.33, 0.34] → loss 1.10
  • Very confident and correct [0.98, 0.01, 0.01] → loss 0.02
  • Very confident but wrong [0.02, 0.9, 0.08] → loss 3.91 (nearly 4 times higher than random guessing!)

The last value deserves attention: because -log(0.02) ≈ 3.91, while -log(0.33) ≈ 1.10.


Mathematical Definition

\[ H(p, q) = -\sum_{i=1}^{C} p_i \cdot \log(q_i) \]

\(p\) is the true distribution (one-hot), \(q\) is the predicted distribution. For a "confident but wrong" prediction, cross-entropy = -log(0.02) ≈ 3.91, MSE = (1-0.02)^2 ≈ 0.96. The penalty of cross-entropy is 12 times that of MSE.


Python Hands-On Practice

Example

import numpy as np

def cross_entropy(p_true, q_pred, eps=1e-12):
    q_pred = np.clip(q_pred, eps, 1 - eps)
    return -np.sum(p_true * np.log(q_pred))

# True label: cat (class 0)
p_true = np.array([1, 0, 0])  # [cat, dog, bird]

predictions = {
    "Completely random guessing (uniform)":    np.array([0.33, 0.33, 0.34]),
    "Slight feeling":           np.array([0.5, 0.3, 0.2]),
    "Gradually becoming reliable":           np.array([0.7, 0.2, 0.1]),
    "Better trained":           np.array([0.9, 0.07, 0.03]),
    "Very confident and correct":     np.array([0.98, 0.01, 0.01]),
    "Very confident but wrong!":   np.array([0.02, 0.9, 0.08]),
}

print(EXAMPLE Cross-Entropy Loss Comparison (True: cat=[1,0,0])\n")
print(f"{'Prediction Description':<20} {'Loss':<10} {'Analysis'}")
print("-" * 52)
for name, q in predictions.items():
    loss = cross_entropy(p_true, q)
    if loss > 2: analysis = "Critical error!"
    elif loss > 0.5: analysis = "Directionally correct but uncertain"
    else: analysis = "Very good"
    print(f"{name:<20} {loss:<10.4f} {analysis}")

# MSE vs Cross-Entropy Comparison
q_bad = predictions["Very confident but wrong!"]
mse_bad = np.mean((p_true - q_bad) ** 2)
ce_bad = cross_entropy(p_true, q_bad)
print(f"\nEXAMPLE Confident but wrong: MSE={mse_bad:.4f}, Cross-Entropy={ce_bad:.4f} (harsh {ce_bad/mse_bad:.0f}x))
EXAMPLE 交叉熵损失对比 (真实: 猫=[1,0,0])

预测描述               损失        分析
----------------------------------------------------
完全瞎猜 (均匀)        1.0986     方向对但不确定
有点感觉               0.6931     方向对但不确定
逐渐靠谱               0.3567     很好
训练较好               0.1054     很好
非常自信且正确         0.0202     很好
非常自信但错误!       3.9120     严重错误!

EXAMPLE 自信但错误时: MSE=0.3260, 交叉熵=3.9120 (严厉 12x)

Application Scenarios in AI

ScenarioLoss function used
Image classificationMulti-class cross-entropy (CategoricalCrossEntropy)
Language modelSum of cross-entropy at each token position — the objective function for GPT training
Knowledge distillationUse the teacher model's soft labels for cross-entropy
Other extensions