Derivative Introduction -- Instantaneous Rate of Change and Tangent Line
Derivative = the instantaneous rate of change of a function at a point = the slope of the tangent line at that point.
Understanding the derivative helps you understand why gradient descent moves in the opposite direction.
From Average Rate of Change to Instantaneous Rate of Change
Average rate of change (secant slope) → Δx gets smaller and smaller → secant becomes tangent →Derivative = slope of the tangent line。
The Sign of the Derivative Reveals Function Behavior
f'(x) > 0
The function is rising here
The tangent line slopes upward to the right
f'(x) < 0
The function is falling here
The tangent line slopes downward to the right
f'(x) = 0
Horizontal — possibly an extremum
The tangent line is horizontal
The intuition behind gradient descent lies in the sign of the derivative:
f'(x) > 0 → the function is rising → move in the opposite direction. f'(x) < 0 → the function is falling → keep moving in this direction.
This is the root of the 'negative gradient direction'.
Real-life Example
Speedometer = derivative of the position function
The whole trip was 120 km and took 2 hours → average speed 60 km/h.
But the instantaneous speed on the dashboard keeps changing. This 'speed at this moment' is the derivative of the distance function with respect to time.
Mathematical Definition
\[ f'(x) = \lim_{\Delta x \to 0} \frac{f(x + \Delta x) - f(x)}{\Delta x} \]If this limit exists, f is differentiable at x.
Python Hands-on Practice
Example
def numerical_derivative(f, x, h=1e-5):
return (f(x + h) - f(x - h)) / (2 * h)
f = lambda x: x**2
# f'(2) = 2*2 = 4
print(f"f'(2) value={numerical_derivative(f, 2):.6f}, theoretical=4")
# Derivative at each point
for x in [-2, -1, 0, 1, 2]:
d = numerical_derivative(f, x)
dir = falling if d < 0 else (rising if d > 0 else horizontal)
print(f"x={x:2d}, f'(x)={d:5.1f}, {dir}")
f'(2) 数值=4.000000, 理论=4 x=-2, f'(x)= -4.0, 下降 x=-1, f'(x)= -2.0, 下降 x= 0, f'(x)= 0.0, 水平 x= 1, f'(x)= 2.0, 上升 x= 2, f'(x)= 4.0, 上升
Application Scenarios in AI
One-dimensional Prototype of Gradient Descent
For a univariate function, gradient descent reduces to \( x_{new} = x_{old} - \eta \cdot f'(x_{old}) \). The derivative tells you which direction the function value decreases — if the derivative is positive, move left; if negative, move right.
The derivative of the activation function determines the efficiency of backpropagation
The derivative of ReLU is 1 when x > 0 and 0 when x ≤ 0. It is extremely fast to compute, and the gradient does not decay in the positive region—this is the key reason ReLU is more suitable for deep networks than Sigmoid. The maximum derivative of Sigmoid is only 0.25, and after being multiplied across multiple layers, the gradient decays exponentially (vanishing gradient).
The derivative of the loss function is zero at the optimal solution
When the model converges to a local optimum, the partial derivatives of the loss function with respect to all parameters are close to 0—gradient descent naturally stops. This is a signal of training convergence: the gradient norm ||\nabla J|| \approx 0.
Other extensions