Probability Basics -- Sample Space, Events and Conditional Probability

Probability theory is a core tool for AI to handle uncertainty.From classifier outputs to generative model sampling, probability is everywhere.


Concept Analysis

Sample Space

The set of all possible outcomes

Rolling a die: {1,2,3,4,5,6}

Event

A subset of the sample space

Even = {2,4,6}

Probability P(A)

The likelihood of A occurring, [0,1]

0 = impossible, 1 = certain

Conditional Probability: Updating Judgment Based on Partial Information

\[ P(A|B) = \frac{P(A \cap B)}{P(B)} \]

Intuition: narrow the sample space to the set where B occurs, and see what proportion of it is occupied by A.

Law of Total Probability: Summarizing Across Cases

\[ P(A) = \sum_i P(A|B_i) P(B_i) \]

Everyday Examples

Law of Large Numbers

Flip a coin 10 times, you might get 4 heads and 6 tails. Flip it 10,000 times, the proportion of heads will definitely be very close to 0.5.

The more trials, the more stable the frequency—this is the foundation of probability theory.


Python Hands-On Practice

Example

import numpy as np

# Law of Large Numbers Verification
for n in [10, 100, 1000, 10000]:
    tosses = np.random.choice(['H','T'], size=n)
    freq = np.mean(tosses == 'H')
    print(f"EXAMPLE flipped {n:5d} times, head frequency={freq:.4f}")

# Monty Hall problem: switching wins 2/3
def monty_hall(switch=True, trials=10000):
    wins = 0
    for _ in range(trials):
        car = np.random.randint(0, 3)
        choice = np.random.randint(0, 3)
        revealed = np.random.choice([d for d in range(3) if d != car and d != choice])
        if switch:
            choice = [d for d in range(3) if d != choice and d != revealed][0]
        wins += (choice == car)
    return wins/trials

print(f"\n"Monty Hall: no switch={monty_hall(False):.3f}, switch={monty_hall(True):.3f} (theoretical: 1/3, 2/3)")

Output:

EXAMPLE 掷   10次, 正面频率=0.8000
EXAMPLE 掷  100次, 正面频率=0.5500
EXAMPLE 掷 1000次, 正面频率=0.4870
EXAMPLE 掷10000次, 正面频率=0.4987

蒙提霍尔: 不换=0.329, 换=0.663 (理论: 1/3, 2/3)

Application Scenarios in AI

Classifier output = probability distribution

Softmax outputs the probability for each class, and the sum of probabilities across all classes is 1. The model doesn't just tell you "this is a cat," it also tells you "80% probability it's a cat, 15% a dog, 5% something else." This kind of probabilistic output is crucial for risk assessment and confidence calibration.

Generative model = sampling from a probability distribution

VAE samples from a learned latent distribution to generate new images, GAN samples from a random noise distribution to generate realistic images, and diffusion models start from pure noise and gradually denoise to generate high-definition images. The core operation of all these generation processes is "sampling from some probability distribution."

Stochastic Policy in Reinforcement Learning

In policy gradient methods, the agent's actions are not deterministic, but are sampled from a probability distribution output by the policy network (e.g., 70% left, 30% right). This randomness ensures exploration—if you always choose the highest-probability action, you may never discover a better policy.


Other extensions