Partial Derivatives -- Differentiating Multivariable Functions One by One
The previous derivative was for a single-variable function f(x). A neural network has millions of parameters, and we need to take the derivative with respect to each parameter.
Partial derivative = In a multivariable function, take the derivative with respect to only one variable, treating the rest as constants.
Concept Analysis
Previously, we only handled the derivative of a single-variable function f(x). But in the real world, the outcome of things is usually determined bymultiple factorsjointly.
For example: House prices depend on multiple factors such as area, location, floor, and age. If we want to know "how much does the house price increase when the area increases by 1 square meter," we need to hold other factors fixed.
This is the core idea of partial derivatives:Change only one variable at a time, observe its effect on the result, and temporarily freeze the remaining variables.。
For \( f(x, y) = x^2 + xy + y^2 \):
- Take the partial derivative with respect to x (treat y as a constant; since y² is a constant, its derivative is 0; in xy, y is treated as a constant, and the derivative of x is 1): \( \frac{\partial f}{\partial x} = 2x + y + 0 \)
- Take the partial derivative with respect to y (treat x as a constant, similarly): \( \frac{\partial f}{\partial y} = 0 + x + 2y \)
The notation \( \partial \) (read as "round d" or "partial") denotes a partial derivative, as opposed to d in an ordinary derivative.
Second-Order Partial Derivatives
Taking the partial derivative of a partial derivative gives the second-order partial derivative:
- \( \frac{\partial^2 f}{\partial x^2} \): Take the partial derivative with respect to x, then take the partial derivative with respect to x again (measures curvature in the x-direction).
- \( \frac{\partial^2 f}{\partial x \partial y} \): First take the partial derivative with respect to x, then with respect to y (mixed partial derivative, measures the interaction effect between the two variables).
The essence of neural network training: the partial derivative of the loss function with respect toEvery parameterCompute partial derivatives, then update along the negative gradient direction.
A model with a million parameters has a gradient that is a vector of a million partial derivatives.
Everyday Example
The taste of a pot of soup
The taste of soup depends on salt, soy sauce, vinegar, and other seasonings. You want to know "how much does the taste change if you add one more spoon of salt?"
Fix the amounts of other seasonings and look only at the effect of salt—this is the idea of partial derivatives:Change only one variable at a time。
Mathematical Definition
\[ \frac{\partial f}{\partial x} = \lim_{h \to 0} \frac{f(x+h, y) - f(x, y)}{h} \]Keep y unchanged, only let x vary.
Python Hands-On Practice
Example
x, y = sp.symbols('x y')
f = x**2 + x*y + y**2
print(f"f(x,y) = {f}")
print(f"∂f/∂x = {sp.diff(f, x)}") # 2x + y
print(f"∂f/∂y = {sp.diff(f, y)}") # x + 2y
# Second-order partial derivative
print(f"∂²f/∂x² = {sp.diff(f, x, 2)}") # 2
print(f"∂²f/∂x∂y = {sp.diff(f, x, y)}") # 1
# Partial derivative of linear model loss with respect to parameters
w, b, xi, yi = sp.symbols('w b x_i y_i')
loss = (w*xi + b - yi)**2
print(f"\nloss=(w·xi+b-yi)^2")
print(f"∂L/∂w = {sp.diff(loss, w)}")
print(f"∂L/∂b = {sp.diff(loss, b)}")
The output is:
f(x,y) = x**2 + x*y + y**2 ∂f/∂x = 2*x + y ∂f/∂y = x + 2*y ∂²f/∂x² = 2 ∂²f/∂x∂y = 1 loss=(w·xi+b-yi)^2 ∂L/∂w = 2*x_i*(b + w*x_i - y_i) ∂L/∂b = 2*b + 2*w*x_i - 2*y_i
Application Scenarios in AI
Neural Network Training = Millions of Partial Derivative Computations
A model with a million parameters has a loss that is a multivariate function of a million variables. In each training round, the automatic differentiation engine computes partial derivatives with respect toevery parametercomputes partial derivatives, obtains a million partial derivative values assembled into a gradient vector, and updates all parameters along the negative gradient direction.
Local Partial Derivatives in Backpropagation
Every node in the computation graph needs to compute local partial derivatives: for an addition node \( \partial/\partial x = 1 \) (the gradient is passed through unchanged), and for a multiplication node \( \partial(xy)/\partial x = y \) (the gradient is multiplied by the other operand's value and passed back). These local partial derivatives are multiplied along the computation graph via the chain rule to obtain the final parameter gradients.
Transfer Learning with Frozen Parameters
In transfer learning, pretrained layers are often "frozen"—the parameters of these layers do not get partial derivatives computed and are not updated. In PyTorch, set requires_grad=False, and the partial derivative computation for the corresponding parameters is skipped, greatly saving memory and computation.
Other extensions