Algebra Basics

This chapter introduces the transformation of equations and inequalities, the core concepts of functions, exponent and logarithm operations, and the summation symbol Σ and series convergence.

The content of this chapter is the common mathematical foundation for sigmoid, softmax, cross-entropy, and information entropy. Each knowledge point is accompanied by interactive graphical demonstrations. Dragging a slider builds intuition better than reading formulas ten times.


Solving and transforming equations and inequalities

Algebraic transformation relies on a single sentence.

Performing the same operation on both sides of an equation keeps the equation true.

As long as you treat both sides of the equation equally—adding, subtracting, multiplying, dividing (with a nonzero divisor), or taking roots (mind the sign)—the solutions of the equation will not change.

Linear equations in one variable

The most basic form is。

solveThe complete steps:

3x + 5 = 20
3x = 20 - 5        ← 两边同减 5(移项)
3x = 15
x = 15 / 3         ← 两边同除以 3
x = 5

There are only two core actions:Transposition (moving terms)(Essentially subtracting the same amount from both sides) andDividing both sides。

Quadratic equations in one variable and the discriminant

The form is, two common solution methods.

Method 1: Factoring。

x² - 5x + 6 = 0
(x - 2)(x - 3) = 0
→ x = 2 或 x = 3

Method 2: The quadratic formula, applicable in all cases but with more computation:

WherecallDiscriminant, denoted as, it directly determines the number of roots:

DiscriminantRoot situationGeometric meaning (parabola and the x-axis)
Two distinct real rootsThe parabola intersects the x-axis at two points.
One repeated rootThe parabola just touches the x-axis (tangent)
No real rootsThe parabola hangs entirely above or below the x-axis

Below, drag the sliders to change, and observe how the parabola's shape and the discriminant are linked.

Interactive demo: Discriminant and the parabola
Drag the three sliders and observe how the number of intersection points between the parabola and the x-axis changes as the sign of Δ changes.

solveThe substitution process into the quadratic formula:

a=1, b=-2, c=-3
Δ = (-2)² - 4×1×(-3) = 4 + 12 = 16
x = (2 ± √16) / 2 = (2 ± 4) / 2
→ x = 3 或 x = -1

Fractional equations: must check at the end

First find a common denominator or clear denominators, then transform it into an ordinary equation and solve.

Finally, you must substitute the solution back into the original equation to confirm it does not make any denominator zero.

Solutions that make the denominator zero must be discarded; they are calledextraneous roots。

Inequalities: the rule that is easiest to get wrong

The transformation rules for inequalities are similar to those for equations, but there is one key difference.

When both sides of an inequality are multiplied (or divided) by a negative number, the inequality sign must be reversed.

solve:

-2x > -6       ← 两边同减 6
x < 3          ← 两边同除以 -2(负数,不等号从 > 变成 <)

This rule is very important in machine learning: when deriving the convergence direction of optimization algorithms or the conditions for loss function decrease, if the sign is handled in the wrong direction, the conclusion will be completely reversed.

System of inequalitiesThe solution set of a system is to take theintersection of the solution sets of multiple inequalities.:

x > 1
x < 5
→ 解集为 1 < x < 5

Common transformation techniques

TechniqueMethodTypical use
TranspositionMove a term to the other side of the equals sign and change its sign.The basic operation for solving all equations.
Completing the squareFinding the vertex of a quadratic function, simplifying derivations of loss functions.
Substitution methodReplace a complex subexpression that appears repeatedly with a new variable.Simplify expressions with complex structure.

Examples

# Solve a quadratic equation in one variable using the discriminant and the quadratic formula
import math

def solve_quadratic(a, b, c):
    """Solve ax^2 + bx + c = 0, return all real roots"""
    delta = b * b - 4 * a * c   # Discriminant Δ = b² - 4ac
    if delta > 0:               # Δ > 0: two distinct real roots
        r = math.sqrt(delta)
        return [(-b + r) / (2 * a), (-b - r) / (2 * a)]
    elif delta == 0:            # Δ = 0: one repeated root
        return [-b / (2 * a)]
    else:                       # Δ < 0: no real roots
        return []

# Verify x² - 2x - 3 = 0; the earlier manual calculation gave x = 3 or x = -1
print(solve_quadratic(1, -2, -3))

# Case Δ = 0: x² - 4x + 4 = 0, i.e., (x-2)² = 0
print(solve_quadratic(1, -4, 4))

# Case Δ < 0: x² + x + 1 = 0
print(solve_quadratic(1, 1, 1))

Running the above code gives the following output:

[3.0, -1.0]
[2.0]
[]

Exercise: Solve the inequality。

Click to view answer

Subtract 9 from both sides:; divide both sides by -3 (negative number, inequality sign reverses):。


The concept of functions: domain, range, composite functions

Functiondescribes a mapping rule from "input to output":For each input x, there is exactly one output y corresponding to it。

Written as, read as "y is a function of x."

Domain: the set of all "valid inputs"

Determining the domain essentially means excluding values that make the expression "meaningless":

Expression featureConstraintExample
Denominatorcannot be 0
Even-indexed radicalThe value under the radical cannot be negative
LogarithmThe argument must be greater than 0

This logarithmic rule is especially important in AI: wherever log is used (such as cross-entropy loss), you must ensure the input is positive. That is why many code implementations write it as—adding a very small constant ε to preventerrors or infinite values.

Range: the set of all "possible outputs"

The range of is, because a square is always non-negative.

The sigmoid functionhas range the open interval, never equal to 0 or 1, only approaching them infinitely.

This is exactly the mathematical reason it is used to represent "probability": probabilities can theoretically approach 0% or 100% infinitely, but never truly equal them.

Composite function: function inside a function

Taking the output of one function as the input of another, written as。

The order of computation isFirst compute the inner g(x), then substitute the result into the outer f。

design,, to find:

f(g(x)) = f(x+1) = (x+1)²

is usually not equal to, the order cannot be changed arbitrarily.

Why composite functions matter for AI: A neural network is essentially a composition of many layers of functions.

A two-layer network can be written as, whereis the first layer (linear transformation + activation function),is the second layer.

The backpropagation algorithm computes gradients "layer by layer" precisely because of thechain rulefor differentiating composite functions (to be covered in detail in later calculus courses).

Examples

# Verification of composite function computation and "non-commutativity of order"
def f(x):
    return x ** 2          # Outer function f(x) = x²

def g(x):
    return x + 1           # Inner function g(x) = x + 1

# f(g(x)): first compute g(3) = 4, then substitute into f to get 16
print(f(g(3)))

# g(f(x)): first compute f(3) = 9, then substitute into g to get 10
# The two are not equal, meaning the composition order cannot be swapped
print(g(f(3)))

The output of executing the above code is:

16
10

Exercise: Let,, findand, and compare whether the two are the same.

Click to view answer

The two differ, indicating that the order of function composition cannot be swapped.


Exponentiation and logarithms

One sentence to understand the division of labor between these two functions:The exponential function is responsible for "amplifying/compressing numerical ranges"; the logarithmic function is responsible for "turning multiplication/division into addition/subtraction and compressing large-range values into a controllable range"。

They appear repeatedly in the design of loss functions and activation functions in AI.

Tends to 0 as x increases, compressing any real number into (0,1)
AI ScenariosFormulaProperty used
sigmoid function
softmax functionExponentials amplify differences, then normalize into a probability distribution
Cross-entropy lossThe logarithm turns "multiplication of probabilities" into "addition of values", avoiding underflow
Information entropylog gives entropy additivity

Rules you must be fluent in

Exponential operation rules:

Logarithmic operation rules, corresponding one-to-one with the exponential rules:

Change-of-base formula, converting logarithms of any base to natural logarithms:

Exponents and logarithms are mirrors of each other

Exponential functionand logarithmic functionare inverse functions of each other, and their graphs are symmetric about the linesymmetric.

Drag the slider to change the base a, and observe how the two curves move together.

Interactive demo: exponentials and logarithms are mirror images of each other
Drag the slider to change the base a; the two curves are always symmetric about the gray dashed line y = x.
Blue is the exponential function y = a^x, green is the logarithmic function y = log_a(x), and the gray dashed line is the axis of symmetry y = x.

Derivation exercise: why cross-entropy uses log

Assume the true label is "cat" (probability 1), and the model predicts the probability of "is cat" as。

If not using log, directly usingas the loss, when p drops from 0.5 to 0.1, the loss only grows "linearly", and the penalty is insufficient.

use: when(the model is extremely confident yet misjudges),, impose extremely severe punishment; when(the judgment is correct),, almost no punishment.

This is a direct application of the logarithmic function's shape—"steep near 0, flat near 1"—in loss function design. Chapter 2 will use an interactive image to specifically demonstrate this curve.

Examples

# Use code to verify the logarithm rule: log(mn) = log(m) + log(n)
import math

m, n = 8, 4

# Left: directly compute log₂(8×4)
left = math.log(m * n, 2)

# Right: log₂(8) + log₂(4)
right = math.log(m, 2) + math.log(n, 2)

print(left)   # log₂32 = 5
print(right)  # 3 + 2 = 5, they are equal, the rule holds
print(abs(left - right) < 1e-12)

Running the above code outputs:

5.0
5.0
True

Exercise: Calculate, and use logarithm rules to simplify into a single logarithm.

Click to see the answer


Sequences and series: understanding Σ and convergence

A sequence is a string of numbers arranged in order, denoted as, abbreviated as。

Two most basic types of sequences:

TypeCharacteristicExampleGeneral term formula
Arithmetic sequenceThe difference between consecutive terms is constant (common difference d)2, 5, 8, 11, ...(d=3)
Geometric sequenceThe ratio between consecutive terms is constant (common ratio r)2, 4, 8, 16, ...(r=2)

What is the summation symbol Σ actually calculating?

It means: let the index i go from 1 to n, and for eachcompute it, then add them all together.

The three-step method for reading a Σ formula: first look atthe summation variablewhich is the index, then look atthe summation range(from where to where), and finally look atwhat each term looks like。

In AI formulas, you can almost always see Σ, such as the mean squared error loss, and the weighted sum of a neuron。

Two summation formulas: arithmetic sequence, geometric sequence ()。

Series and convergence

A series is the sum of the "infinitely many terms" of a sequence, i.e.,。

When infinitely many numbers are added, will the result be infinite? This is the question of convergence and divergence:

If, as the number of terms increases, the accumulated sum approaches adefinite finite value, the series is said to converge; if the sum grows without bound or has no fixed trend, the series is said to diverge.

The most classic example is the geometric series:

Whenconverges when , and the sum is; whenit diverges when .

Click the "Play" button to observe how the partial sums approach the limit step by step.

Interactive Demo: Convergence Process of a Geometric Series
Drag the slider to change the common ratio r, then click "Play" to observe how the partial sums S₁, S₂, ... approach (or move away from) the limit 1/(1-r).
The red dashed line is the theoretical limit 1/(1-r). When |r| is close to 0.9, convergence becomes noticeably slower; if |r| ≥ 1 (the slider in this demo is limited to within 0.9), the series will diverge.

Why convergence matters for AI

When training neural networks, each update step of gradient descent can be viewed as a sequence; whether the weight updates can "converge" to a stable value shares the same intuition as series convergence.

Certain optimization algorithms (such as the momentum term in Adam) essentially perform aweighted exponentially decaying summation, whose mathematical form is a geometric series with a common ratio less than 1.

In recurrent neural networks (RNNs), if a multiplicative factor is multiplied repeatedly and its absolute value is less than 1, it leads to "vanishing gradients" (corresponding to the series converging to 0); if greater than 1, it leads to "exploding gradients" (corresponding to the series diverging). This is a classic problem in deep learning, rooted precisely in the intuition of geometric sequences.

Examples

# Observe the process of partial sums of a geometric series approaching the limit: r = 0.5
# Theoretical limit: 1 / (1 - 0.5) = 2

r = 0.5
limit = 1 / (1 - r)      # Theoretical limit value 2
s = 0                    # Partial sum, accumulating from 0

for i in range(6):       # Accumulate term by term: 1 + 0.5 + 0.25 + ...
    s += r ** i
    print(i + 1, s)      # Print the term count and current partial sum

print(limit)             # Theoretical limit, the partial sums get closer and closer to it

Running the above code produces the output:

1 1.0
2 1.5
3 1.75
4 1.875
5 1.9375
6 1.96875
2.0

You can see that with each added term, the gap between the partial sum and the limit is halved — this is exactly the convergence rhythm of a geometric series with a common ratio of 0.5.

Exercise: Determine whether the seriesconverges, and if it does, find its sum.

Click to view the answer

Common ratio, satisfying, it converges.


Chapter summary

TopicCore idea in one sentence
Equations and inequalitiesPerform the same operation on both sides; reverse the inequality sign when multiplying/dividing by a negative number
Concept of functionsThe domain is "what can be input", the range is "what can be output", and a composite function is "a function within a function"
Exponents and logarithmsExponents amplify/compress values, logarithms turn multiplication and division into addition and subtraction — core tools for AI loss functions
Sequences and seriesΣ means "sum according to a rule"; whether a series converges determines whether model training is stable
Other extensions