Common Differentiation Rules — Power, Exponential, and Logarithm

In the previous chapter, we understood the meaning of derivatives. In this chapter, we master actual computation. The three most commonly used types of differentiation rules in AI — memorize them and you can handle most scenarios.


Quick Reference Table of Derivatives of Basic Functions

Function f(x)Derivative f'(x)Scenarios in AI
\( c \) (constant)\( 0 \)The constant disappears when differentiating the bias term
\( x^n \)\( nx^{n-1} \)\( x^2 \) in the MSE loss
\( e^x \)\( e^x \)Sigmoid、Softmax
\( \ln x \)\( \frac{1}{x} \)Cross-entropy loss

Differentiation Rules for the Four Arithmetic Operations

RuleFormula
Sum/difference\( (f \pm g)' = f' \pm g' \)
Product\( (fg)' = f'g + fg' \)
Scalar multiplication\( (cf)' = cf' \)

The most important derivative:\( (e^x)' = e^x \)—The derivative of the exponential function equals itself.

This makes the exponential function ubiquitous in differential equations and deep learning.


Everyday Examples

Learning differentiation rules is like learning knife skills in cooking — each rule is a basic skill, and combining them lets you handle any recipe.

\( (e^x)' = e^x \) is like an unchanging ingredient: no matter how you process it, it always remains the same.


Mathematical Definition

Power function differentiation: reduce the exponent by one and multiply by the original exponent

\[ \frac{d}{dx} x^n = n x^{n-1} \]

This is the most basic differentiation formula. For example, \( (x^2)' = 2x \), \( (x^3)' = 3x^2 \).

In the MSE loss, \( (y - \hat{y})^2 \) is a power function; differentiating it yields \( 2(y - \hat{y}) \).

Exponential function differentiation: the derivative is itself

\[ \frac{d}{dx} e^x = e^x \]

This is one of the most unique functions in mathematics.The derivative of e^x is still e^x—no matter how many times you differentiate it, it remains the same.

This property makes e^x appear repeatedly in the Sigmoid function \( \sigma(x) = 1/(1+e^{-x}) \) and Softmax.

Logarithmic function differentiation

\[ \frac{d}{dx} \ln x = \frac{1}{x} \]

The derivative of the log term in cross-entropy loss uses this. The derivatives of logarithm and exponential are reciprocals—they are inverse functions of each other.

Sum and difference rule

\[ (f \pm g)' = f' \pm g' \]

The derivative of a sum = the sum of derivatives. The differentiation operation can 'pass through' plus and minus signs, differentiating each term separately.

Product rule

\[ (f \cdot g)' = f' \cdot g + f \cdot g' \]

The derivative of a product ≠ the product of derivatives. Instead, it is 'derivative of the first times the second + the first times the derivative of the second'.

This rule frequently appears in backpropagation—when computing the partial derivative of \( w \cdot x \) with respect to w, the product rule is used.

Combining these three types of functions and three rules allows you to find the derivatives of the vast majority of functions in AI.

For example, the derivative of Sigmoid \( \sigma'(x) = \sigma(x)(1-\sigma(x)) \) is obtained by comprehensively applying the quotient rule, exponential function differentiation, and the chain rule.


Hands-on Python

Example

import sympy as sp

x = sp.Symbol('x')
funcs = {'x^3': x**3, 'e^x': sp.exp(x), 'ln(x)': sp.log(x),
         'x^2 + 3x + 5': x**2 + 3*x + 5}

print("=== Symbolic differentiation verification ===")
for name, f in funcs.items():
    print(f"f(x)={name:15s} → f'(x)={sp.diff(f, x)}")

# Product rule
f = x**2 * sp.exp(x)
print(f"\n(x^2 · e^x)' = {sp.diff(f, x)}")
# Chain rule: derivative of (2x+1)^3
g = (2*x + 1)**3
print(f"((2x+1)^3)' = {sp.diff(g, x)}")

Application scenarios in AI

Gradient computation in backpropagation

Every layer of a neural network involves these basic differentiation rules. The linear layer y = Wx + b uses the product rule for the derivative with respect to parameter W, and for b uses the fact that the derivative of a constant is 0.

Derivation of the derivative of the Sigmoid activation function

\( \sigma(x) = 1/(1+e^{-x}) \), using the quotient rule gives \( \sigma'(x) = \sigma(x)(1-\sigma(x)) \). This compact form is extremely fast to compute in backpropagation—it only needs the \sigma(x) value already computed during forward propagation, no need to recompute the exponential.

Combined gradient of Softmax + cross-entropy

Softmax involves exponentials, summation, and division, and differentiating it alone is complex. But when combined with the cross-entropy loss, the gradient miraculously simplifies to \( \hat{y} - y \) (prediction - true label)—this is one of the most elegant derivative simplifications in deep learning, and the fundamental reason why Softmax+CE is standard for classification tasks.


Other extensions