AI Introduction

Imagine this scenario:

  • In the morning, you open your phone, and the photo album has automatically sorted last night's photos by people and scenes.

  • On your way to work, you say a sentence to your phone, and the navigation app plans a route that avoids traffic jams for you.

  • When you get to your desk, you type a line into the dialog box, and AI helps you generate a draft of the entire weekly report.

These scenarios seemed like science fiction five years ago, but now they are part of everyday life. AI is not the future; it has already been embedded in the products you use every day.

The appearance of ChatGPT at the end of 2022 was a watershed moment:

  • Before this, AI was a matter for programmers and researchers.

  • After this, AI became a tool that everyone can use directly.


Definition of Artificial Intelligence

In plain language:Artificial Intelligence (AI for short) is technology that gives computers intelligent behavior similar to humans.。

The intelligent behavior mentioned here includes: understanding text, recognizing speech, identifying images, making decisions, and learning from experience.

The English abbreviation AI is pronounced asA-I(read the two letters separately), the full name isArtificial Intelligence。

What AI Can Do

The following capabilities are already quite mature in today's AI:

CapabilityTypical ApplicationsHave you used it?
Language Understanding and GenerationTranslation, writing, summarizing articles, answering questionsChatGPT、Claude
Image RecognitionFace unlock, document scanning, medical image analysisPhone album categorization, station turnstiles
Voice InteractionSpeech-to-text, voice assistantsSiri, Xiao Ai, input method voice
Content GenerationAI painting, video generation, music creationMidjourney、Suno
Code AssistanceCode completion, bug fixing, automatic code generationGitHub Copilot、Cursor

What AI Cannot Do

Knowing AI's boundaries is more important than knowing its capabilities.

AI doesn't truly understand anything—it's only doing probability calculations. When you say "the weather is really nice today," it doesn't know what "nice" feels like; it only knows what words usually follow that sentence.

AI has no consciousness, no emotions, and no goals of its own; it doesn't want to do anything.

AI can hallucinate, that is, it will confidently fabricate facts, names, papers, and data that don't exist, because it is essentially predicting the next word rather than verifying facts.

AI lacks judgment in new situations it hasn't seen; if there are no similar scenarios in the training data, its performance may be very poor.

Remember:AI is a powerful assistant, not a reliable authority。

For high-risk decisions involving medical care, law, investments, etc., AI's output can only be used for reference, not as the final basis.


A Brief History of AI Development

AI didn't appear out of nowhere; its history spans more than 70 years, with three major boom-and-bust cycles.

First Wave: Symbolic AI (1950s–1980s)

In 1956, a group of scientists met at Dartmouth College,Artificial IntelligenceThis term was formally proposed.

The mainstream idea at the time was simple: write human expert knowledge into a set of rules, store them in a computer, and reason according to the rules when a problem arises. This approach was called expert systems.

For example, a medical diagnosis expert system might store thousands of rules: if a patient has a fever and cough, it could be a cold; if the fever exceeds 39 degrees and lasts three days, a blood test is recommended.

The problem is that the rules of the real world are endless. You write five thousand rules, and then a patient comes in with a six-thousandth symptom combination, and the system has no idea what to do.

Moreover, modifying rules is extremely painful—adding a new rule can conflict with hundreds of existing ones.

In the late 1980s, people realized this path would not work, and AI entered its first winter.

Second Wave: Rise of Machine Learning (1990s–2010s)

Researchers changed their approach:Instead of having humans write rules, let machines find patterns from data on their own.。

For example, to identify spam, instead of writing a rule like "if it containsprizethen it is spam," show the machine ten thousand labeled emails (five thousand normal, five thousand spam) and let the algorithm figure out what spam usually looks like.

This period gave birth to classic algorithms such as support vector machines (SVM), random forests, and logistic regression, which worked well in scenarios like spam filtering, credit card fraud detection, and product recommendation.

But there was a bottleneck: features had to be designed by humans. For example, in image recognition, you first had to manually extract features such as "edges," "color distribution," and "texture" before feeding them to the algorithm. The quality of human feature extraction determined the upper limit of the model.

Third Wave: Deep Learning and Large Models (2010s–present)

In 2012, a deep neural network called AlexNet overwhelmingly outperformed traditional methods in the ImageNet image recognition competition, officially marking the beginning of the deep learning era.

The core breakthrough of deep learning was:Even features no longer needed manual design—the model learned layer by layer from raw data by itself.。

In 2017, Google published the paper "Attention Is All You Need," proposing the Transformer architecture.

In November 2022, OpenAI released ChatGPT, which surpassed one hundred million users in two months, truly bringing AI to the general public.

From 2023 to the present, large models such as GPT-4, Claude, and Gemini have continued to iterate, with new directions such as multimodal AI, AI agents, and reasoning models constantly emerging.

The core thread of the three waves, summed up in one sentence:

WaveCore ideaWho does the workRepresentative event
First (1950s–1980s)Humans write rules, machines execute themProgrammers write rules1956 Dartmouth Conference
Second (1990s–2010s)Machines learn patterns from dataHumans design features, algorithms learn patterns1997 Deep Blue defeats chess champion
Third (2010s–present)Deep networks + massive data + massive computing powerModels even learn features themselves2022 ChatGPT release

The evolution path of AI: rule-driven → data-driven → deep learning → large models → agents. More specifically, we can divide it into five stages:

  • 1950s: People raised a question—can machines think like humans?
  • 1960s–1980s: Attempts were made to encode knowledge and rules into machines (symbolicism, expert systems), but results were limited.
  • 1990s–2000s: Instead of hand-writing rules, machines were allowed to learn from data (machine learning).
  • 2010s: Deep learning rose to prominence, and neural networks achieved breakthroughs by relying on massive data and computing power.
  • 2020s: The era of large models arrived, and a single model began to possess general-purpose capabilities (chatting, writing, programming, reasoning).
  • Future: AI is moving from answering questions to autonomously completing tasks (Agent, AGI).


Three Easily Confused Concepts: AI, ML, DL

AI (Artificial Intelligence), ML (Machine Learning), and DL (Deep Learning) are often conflated by the media, but their definitions are not the same—they are not the same thing.

First, look at a relationship diagram:

AI、机器学习、深度学习三者包含关系图

Containment Relationship: AI ⊃ ML ⊃ DL

  • Artificial intelligence (AI) is the largest circle; any technology that uses machines to simulate intelligent behavior counts as AI.

  • Machine Learning (ML) is a subset of AI, specifically referring to methods that automatically learn from data.

  • Deep Learning (DL) is a subset of ML, specifically referring to methods based on multi-layer neural networks.

Not all AI is ML, and not all ML is DL.

Use an Example to Clarify the Differences Among the Three

Suppose we want to build a program to determine whether an email is spam. The three approaches correspond to three levels:

Example

# ============================================
# Approach 1: Rule-based (AI, but not ML)
# Programmer writes rules by hand; there is no "learning" process.
# ============================================

def is_spam_by_rules(email_text: str) -> bool:
    """Use keyword rules to judge spam — purely manual rules, no learning."""
    # If any of the following keywords is hit, it is judged as spam.
    spam_keywords = [
        "Congratulations on winning",
        "Claim for free",
        "Click to claim",
        "Limited-time offer",
        "example gift pack",   # For testing; can be replaced in real scenarios.
    ]
    for keyword in spam_keywords:
        if keyword in email_text:
            return True
    return False


# Test: an email with the following content
test_email = "Congratulations! You have obtained a example gift pack, click to claim!"
result = is_spam_by_rules(test_email)
print(f"Rule-based result: {'spam' if result else 'normal email'}")
# Output: Rule-based judgment result: spam

The advantage of the rule-based method is that it is simple and direct; the disadvantage is that rules can never be fully written out or cover everything. If a spam email uses a different phrasing, such as "Congratulations, Your Excellency, on winning the grand prize," the rule will miss it.

Example

# ============================================
# Approach 2: Classic Machine Learning (ML, but not DL)
# Automatically learn patterns from historical data; no need to write rules by hand.
# ============================================

def train_simple_classifier(emails: list, labels: list):
    """Train the simplest machine learning classifier
Idea: Count the frequency of each word in spam and normal emails,
Use the frequency difference to judge new emails."""

    # Count how many times each word appears in the two types of emails.
    spam_word_count = {}   # Total number of times the word appears in spam.
    ham_word_count = {}    # Total number of times the word appears in normal emails.
    spam_total = 0         # Total word count of spam.
    ham_total = 0          # Total word count of normal emails.

    for email_text, label in zip(emails, labels):
        words = email_text.split()
        for word in words:
            if label == "spam":
                spam_word_count[word] = spam_word_count.get(word, 0) + 1
                spam_total += 1
            else:
                ham_word_count[word] = ham_word_count.get(word, 0) + 1
                ham_total += 1

    # Return the learned statistical information (this is the "model").
    return {
        "spam_word_count": spam_word_count,
        "ham_word_count": ham_word_count,
        "spam_total": spam_total,
        "ham_total": ham_total,
    }


def predict(model: dict, email_text: str) -> str:
    """Use the learned model to predict new emails."""
    words = email_text.split()
    spam_score = 0.0   # Spam score
    ham_score = 0.0    # Normal email score

    for word in words:
        # Calculate the probability of this word appearing in spam (add smoothing to avoid division by zero).
        spam_prob = (model["spam_word_count"].get(word, 0) + 1) / (model["spam_total"] + 1)
        ham_prob = (model["ham_word_count"].get(word, 0) + 1) / (model["ham_total"] + 1)
        spam_score += spam_prob
        ham_score += ham_prob

    return "spam" if spam_score > ham_score else "ham"


# Training data: 4 labeled emails
emails = [
    "Congratulations on winning Claim for free Gift pack",           # spam
    "Limited-time offer Click to claim example gift pack",       # spam
    "Meeting tomorrow remember to bring report",                # normal
    "Weekend together eat Are you free",              # normal
]
labels = ["spam", "spam", "ham", "ham"]

model = train_simple_classifier(emails, labels)

# Predict a new email
new_email = "Congratulations on winning the grand prize Hurry to claim"
result = predict(model, new_email)
print(f"ML method judgment result: {'spam' if result == 'spam' else 'normal email'}")
# Output: ML method judgment result: spam

The ML method no longer requires hand-written rules; it learns automatically from data.

But it still requires humans to design "features" — in this example, the feature is "using word frequency to make judgments." If features are poorly designed, the model's performance cannot improve.

Example

# ============================================
# Approach 3: Deep Learning (DL, a subset of ML)
# Use a multi-layer neural network; let the model learn features itself.
# Here, pseudocode + comments are used to explain the principle, without relying on any framework.
# ============================================

# Idea for deep learning-based spam classification:
# 1. Convert each word into a vector (a sequence of numbers); this step is called "word embedding".
# 2. Feed these vectors into a multi-layer neural network.
# 3. Each network layer automatically extracts increasingly abstract features.
# 4. The last layer outputs the probability of "spam" or "normal".

# Below is a conceptual three-layer neural network structure:

class SimpleNeuralNetwork:
    """Concept demonstration: the structure of a three-layer neural network (does not include training logic)."""

    def __init__(self, input_size: int, hidden_size: int, output_size: int):
        """Initialize the weights of each network layer."""
        import random
        # Layer 1: input → hidden layer (apply a nonlinear transformation to the input features).
        self.w1 = [[random.random() for _ in range(hidden_size)]
                   for _ in range(input_size)]
        # Layer 2: Hidden layer → Output layer (maps abstract features to final classification)
        self.w2 = [[random.random() for _ in range(output_size)]
                   for _ in range(hidden_size)]

    def forward(self, x: list) -> list:
        """Forward propagation: input data passes through the network to obtain output"""
        # First layer transformation
        hidden = [sum(x[i] * self.w1[i][j] for i in range(len(x)))
                  for j in range(len(self.w1[0]))]
        # ReLU activation function (negative values become 0, positive values remain unchanged)
        hidden = [max(0, h) for h in hidden]
        # Second layer transformation yields the final output
        output = [sum(hidden[i] * self.w2[i][j] for i in range(len(hidden)))
                  for j in range(len(self.w2[0]))]
        return output


# Assume each word has been converted into a 100-dimensional vector (done automatically by the word embedding layer)
# Input 100-dim vector → hidden layer 64 neurons → output 2 values (spam probability, normal probability)
model = SimpleNeuralNetwork(input_size=100, hidden_size=64, output_size=2)
print("Three-layer neural network structure created: 100 → 64 → 2")
# Output: Three-layer neural network structure created: 100 → 64 → 2

Comparison of three approaches:

ApproachWho does the workCategoryPros and cons
Rule-based methodProgrammers manually write rulesAI (non-ML)Simple and direct, but rules are endless and hard to maintain
Classic MLAlgorithms learn patterns from data; features are designed by humansML (non-DL)Works well with small amounts of data, strong interpretability
Deep learningEven features are learned by the model itselfDLWorks well with massive data, but requires a lot of computing power

When someone says "our company is working on AI," you can ask: "Are you using rules, traditional machine learning, or deep learning?" — This helps you quickly determine their technical approach.


Weak AI, Strong AI, Super AI

Besides classifying AI by technical approach, AI can also be divided into three levels by "intelligence level."

Weak AI (Narrow AI) — Already Realized

Can only perform well on specific tasks; switch tasks and it's at a loss.

AlphaGo can beat world champions at Go, but ask it to write a poem, and it can't.

ChatGPT is great at chatting, but give it a medical image to diagnose, and it can't.

All AI products we can access today are weak AI.

"Weak" doesn't mean weak capability; it means narrow scope — it excels only in one or a few specific domains.

Strong AI (AGI, Artificial General Intelligence) — Not Yet Realized

Able to learn and work in any field like a human. No need for separate training on each new task; a few examples are enough to get started.

AGI is the ultimate goal of companies such as OpenAI, Anthropic, and Google DeepMind.

No widely recognized AGI has emerged yet. When will it be realized? Optimistic estimates say 5-10 years, pessimistic estimates say over 50 years; no one can say for sure.

Super AI (ASI, Superintelligence) — Purely Theoretical

Comprehensively surpasses human intelligence in all domains. Can make scientific discoveries that humans cannot, and solve problems that humans cannot understand.

This is a common setting in science fiction, but it is still very far from reality.

Relationship among the three levels:

LevelCapability scopeCurrent statusExamples
Weak AISingle domainAlready widespreadChatGPT, facial recognition, AlphaGo
Strong AI (AGI)General domain, comparable to humansNot yet achieved—
Super AI (ASI)Comprehensively surpasses humansPurely theoretical concept—

Correcting Common Misconceptions

There are two most widespread misconceptions about AI, and they deserve to be addressed explicitly.

Misconception 1: AI is conscious

AI has no consciousness.

Don't be deceived by AI's tone. It can write poems that bring you to tears, but it has no feelings about that poem, and doesn't even know what crying means.

When you chat with ChatGPT and it says "I feel" or "I think," no "feeling" or "thinking" is actually happening.

Essentially, AI is doing only one thing:Predicting the most likely next character based on all the preceding text.。

For example: you input "今天天气true" (the weather today is really), it predicts the next character is "OK" (good) with 85% probability, "热" (hot) with 8%, "冷" (cold) with 5%, and then chooses "OK". Next, based on "今天天气trueOK" (the weather today is really good), it predicts the next character is "啊" (ah)... generating character by character in this way.

The entire process has no subjective experience, no self-awareness, no intention; it merely mimics the probability distribution of human language.

Misconception 2: AI will replace humans

A more accurate statement:AI will not replace humans, but humans who use AI will replace humans who don't use AI。

History has repeatedly proven this pattern:

  • After the calculator appeared, abacus operators disappeared, but the work efficiency of mathematicians and engineers increased tenfold.

  • After search engines appeared, librarians became fewer, but everyone's ability to acquire knowledge jumped by several orders of magnitude.

  • AI is the same: repetitive, rule-based work will be replaced, but work requiring judgment, creativity, and interpersonal collaboration will be enhanced.

For individuals, the most important thing is not to fear AI, but tolearn to collaborate with AI。

Other Extensions