Functions and Graphs
In this chapter, we will visually display the graphs of common functions; understanding the shape of a graph is more important than memorizing formulas.
Loss function surfaces and activation functions are essentially variants of these basic shapes. Each function comes with a switchable interactive graph; switching and dragging by hand is far more effective than looking at static illustrations.
Why specifically learn to "read graphs"
Mathematical formulas describe "rules," while graphs describe "what the rules look like."
Many core problems in deep learning are essentially "curve or surface shape" problems:
| Deep learning problem | Underlying shape problem |
|---|---|
| Why some loss functions are easier to optimize | See whether its surface is "bowl-shaped" (convex function) |
| Why sigmoid causes slower training | See how "flat" its graph is at both ends |
| Why ReLU became the default choice | See how "simple" its graph is |
Looking at each function with questions like "What shape is this? How are the ends? How is the middle?" will help you remember it better than rote memorizing formulas.
Five basic function graphs
Linear function: y = kx + b
Its shape is a straight line.k is the slope, determining the degree and direction of tilt (k>0 slopes up to the right, k<0 slopes down to the right);b is the intercept, determining the point where the line intersects the y-axis.
Both the domain and range are all real numbers, and the rate of change (slope) is the same everywhere; this is the meaning of the word "linear."
Role in AI: the most basic operation of a neuronIt is precisely a linear function in high-dimensional space. Weights control the direction and degree of tilt, and the bias controls the overall translation.
Quadratic function: y = ax² + bx + c
Its shape is a parabola.Opening upward (bowl-shaped), has a minimum value;Opening downward (inverted bowl), has a maximum value.
The vertex is the turning point of the curve and can be found by completing the square:The graph is symmetric about the vertical line through the vertex.
Role in AI: the simplest convex function looks like this. Mean squared error (MSE) approximates this "bowl-shaped" surface in many simplified scenarios—gradient descent easily finds the lowest point on a bowl-shaped surface because no matter which direction you descend, you head toward the same valley floor.
Exponential function: y = aˣ (a>0 and a≠1)
The shape is a curve that rises with accelerating speed (when a>1) or continuously decays (when 0<a<1); it never touches the x-axis, which is its horizontal asymptote.
The domain is all real numbers, and the range isWhen the input grows linearly, the output grows "explosively"—this is the core property of the exponential function.
Role in AI: in softmax,and the exponential decay in the momentum term of the Adam optimizer both exploit the properties of "amplifying differences" or "gradually forgetting".
Logarithmic function: y = logₐx (a>0 and a≠1)
It is exactly the "mirror image" of the exponential function (with the line y=x as the axis of symmetry). The curve rises rapidly from x=0, then growth becomes slower and slower.
The domain is, and the range is all real numbers. The larger the input, the slower the growth—completely the opposite of the exponential function.
Role in AI: in cross-entropy loss,it precisely leverages the logarithm's shape—"extremely steep near 0, flattening near 1"—the more confidently wrong the model is, the heavier the penalty.
Trigonometric functions: y = sin(x), y = cos(x)
The shape is a wave that repeats continuously (periodicity) and always oscillates between -1 and 1 (boundedness). sin(x) starts from the origin and goes upward; cos(x) reaches its maximum value 1 at x=0.
The period is, and the range is。
Role in AI: Transformer positional encoding directly overlays sin/cos waveforms of different frequencies to generate a unique yet regular "coordinate" for each position in the sequence (expanded in Chapter 3).
Three key AI activation functions
Activation functions are the "source of nonlinearity" for each layer of a neural network, and their shapes directly determine how easy or difficult training is.
sigmoid:σ(x) = 1 / (1 + e⁻ˣ)
The shape is an S-shaped curve that slowly climbs from 0 to 1, rising fastest in the middle (near x=0), and flattening out at both ends.
The value range is an open interval, often interpreted as 'probability'.
The biggest side effect: the slope at both ends approaches 0.When the absolute value of the input is large (e.g., x=10 or x=-10), the sigmoid curve is nearly flat, and the gradient is nearly zero.
In deep networks, if many layers all use sigmoid, the gradient gets weakened layer by layer by these "flat regions" during backpropagation, and ultimately cannot reach the earlier layers—this is...The vanishing gradient problemThe most intuitive geometric origin.
ReLU:f(x) = max(0, x)
The shape is a piecewise linear line: when x<0 it hugs the x-axis (output is always 0), and when x≥0 it is a straight line with slope 1.
The derivative is always 0 on the negative half-axis and always 1 on the positive half-axis. There's no "gradually flattening" process like sigmoid, and the computation is extremely simple (only one comparison needed).
The gradient on the positive half-axis is always 1 and does not decay as the input grows larger, greatly alleviating the vanishing gradient problem. This is the intuitive reason why ReLU became the default activation function.
italsoYesownProblem:negative半轴梯degreeHengis 0,possible导致some神经元"死掉"No再Update--这fromimage形stateTopalsoabilitystraightJie看出come。 -> It also has its own problems: the gradient on the negative half-axis is always 0, which may cause some neurons to "die" and stop updating—this can be directly seen from the shape of the image.
tanh:f(x) = (eˣ - e⁻ˣ) / (eˣ + e⁻ˣ)
An S-shaped curve very similar to sigmoid, but centered at the origin, with a range of...。
Compared to sigmoid, tanh's output is centered at 0. If the next layer's input mean is close to 0, training tends to be more stable—this is why tanh is more popular in certain scenarios.
But it doesn't solve the fundamental problem of "flattening at both ends," which is also the background for why ReLU-family functions later became popular.
Why cross-entropy uses -log(p)
In binary classification problems, the probability that the model predicts the "correct class" is..., the cross-entropy loss is。
Intuitively, you might think: useWouldn't the loss be simpler? When the two curves are placed together, the difference is immediately visible.
Examples
import math
# ---- Part 1: The sigmoid derivative decays as |x| increases ----
def sigmoid(x):
return 1 / (1 + math.exp(-x))
def sigmoid_grad(x):
s = sigmoid(x)
return s * (1 - s) # Derivative formula of sigmoid
print(sigmoid_grad(0)) # At the peak: 0.25, already the maximum value of the sigmoid derivative
print(sigmoid_grad(5)) # When |x|=5: less than 0.007 remains
print(sigmoid_grad(10)) # When |x|=10: about 0.0000456, the gradient almost vanishes
# ---- Part 2: The growth of the -log(p) penalty ----
for p in [1.0, 0.5, 0.1, 0.01]:
print(p, -math.log(p)) # The smaller p, the steeper the penalty; at p=0.01, the penalty has already reached 4.6
Running the above code produces the following output:
0.25 0.006648056670790155 4.5395807735957664e-05 1.0 0.0 0.5 0.6931471805599453 0.1 2.302585092994046 0.01 4.605170185988091
Two numbers are worth remembering: the peak of the sigmoid derivative is only0.25; when p drops from 0.5 to 0.01,the penalty rises from 0.69 to 4.6.
Three scenarios for applying "shape intuition"
| Scenario | What shape to look for | Criterion |
|---|---|---|
| Judging whether the loss function is easy to optimize | Whether the surface is close to 'bowl-shaped' (convex) | The closer to a bowl shape, the easier it is for gradient descent to find the global minimum; a bumpy (non-convex) surface may get stuck in a local optimum |
| Judging whether the activation function slows down training | Whether the two ends of the graph 'flatten out' | The flatter, the more obvious the vanishing gradient problem |
| Judging whether a function is suitable for probability output | Whether the range is bounded | Only when the range is compressed into an interval like (0,1) can it be interpreted as a probability; this is the direct reason why sigmoid and softmax are chosen |
Exercise: Without looking at formulas, judge by intuition alone—between ReLU and sigmoid, which one is more suitable for the final layer of the network for 'binary classification probability output'? Why?
Click to view the answer
sigmoid is more suitable because its range is, which can naturally be interpreted as 'the probability of belonging to a certain class'.
ReLU's range is, has no upper bound, so it cannot be directly used as a probability.
Chapter summary
| Function | Shape keywords | Corresponding AI concept |
|---|---|---|
| Linear function | Straight line, same slope everywhere | Neuron's weighted sum wx+b |
| Quadratic function | Bowl-shaped / inverted bowl-shaped | Loss function surface (convex function intuition) |
| Exponential function | Accelerated rise or decay | softmax, momentum term |
| Logarithmic function | Growth is getting slower and slower. | Cross-entropy loss |
| Trigonometric functions | Periodic waves | Positional encoding |
| sigmoid | S-shaped, flattened at both ends. | Probability output, the source of vanishing gradients. |
| ReLU | Polyline, negative half-axis returns to zero. | Default activation function |
| tanh | S-shaped, centered at 0. | Scenarios requiring "zero-mean output" |