Frontier Research Trends

The pace of progress in the AI field is often astonishing. A new idea seen in an academic paper today may become a product feature everyone uses six months later.

Scaling Law, MoE mixture of experts, long context techniques, inference-time compute scaling, multimodal fusion, AI Agent, few-shot learning, AI for Science—these directions being explored in the lab will define the shape of AI products for the next decade.

This module does not aim to explain every technical detail, but rather to help you build a panoramic view of frontier research: know what important directions exist, what problems each direction is solving, how far it has progressed, and where it might go in the future.

The value of understanding frontier research lies not in following trends, but inSee the thread of technological evolution, and find the unchanging laws amid change.。


Scaling Law (Scale Law)

Scaling Law has been the core guiding principle behind the success of large models over the past few years—larger models, more data, and stronger compute bring better performance.

What is Scaling Law

Simply put:Within a certain range, model performance predictably improves as computation, data, and parameters increase.。

This is like farming: within a reasonable range, the more fertilizer, water, and fertile soil, the better the harvest.

Scaling Law was first systematically articulated by OpenAI in the 2020 paper "Scaling Laws for Neural Language Models".

They found that when you scale model parameters, data, and computation by N times, the model's loss decreases as log(N)—that is, although marginal returns diminish, as long as you keep investing, performance will continue to improve.

Chinchilla Optimal Training Law

In 2022, DeepMind's paper "Training Compute-Optimal Large Language Models" (the Chinchilla paper) brought an important correction.

The prevailing approach before was to make the model as large as possible and then train it with relatively little data.

The Chinchilla paper pointed out:Previous models were too large and had too little data.。

They proposed that, given a certain compute budget, model parameters and training data should be scaled up together at a certain ratio.

Specifically: when computation doubles, model parameters should be scaled up by about 1.4 times, and training data should also be scaled up by about 1.4 times.

ModelParametersTraining data sizeRelease dateCharacteristics
GPT-3175B300B tokens2020Large model, relatively little data
Chinchilla70B1.4T tokens2022Slightly smaller model, much more data
LLaMA 270B2T tokens2023Following the Chinchilla approach

Chinchilla's impact has been far-reaching: subsequent mainstream models, from LLaMA to GPT-4, have placed more emphasis on the data ratio, rather than blindly pursuing ultra-large parameter counts.

Limitations and Controversies of Scaling Law

Scaling Law is not a panacea. Its limitations are reflected in several aspects:

First isdiminishing marginal returns—to double performance might require tenfold or even hundredfold computational investment.

Second isthe unpredictability of emergent abilities—Some abilities (such as complex reasoning) do not appear at all when the model is small, but suddenly emerge after reaching a certain scale. However, no one can accurately predict when the next ability will emerge or what it will be.

The third isdata bottleneck—The total amount of high-quality text data is limited. At the current rate of consumption, we may soon hit the data ceiling.

Core insight of Scaling Law:Scale is not everything, but without scale, nothing is possible. Today's models still benefit from larger scale, but researchers are also exploring new paths that go beyond simply scaling up.


Mixture of Experts (MoE)

The Mixture of Experts (MoE) is an architectural design that makes the model larger without significantly increasing inference cost.

Dense Model vs Sparse Model

Traditional large models are "dense" — every time a token is input, all parameters of the model are used.

MoE is "sparse" — each token only uses a small portion of the model's parameters (i.e., a few "experts"), while the other parameters remain dormant.

This is like going to a hospital:

Dense model: When you see a doctor, doctors from all departments come for a consultation — comprehensive, but too costly.

MoE model: When you see a doctor, the triage desk assigns you to a few relevant departments (e.g., internal medicine + cardiology), and only doctors from those departments diagnose you — ensuring professionalism while controlling cost.

CharacteristicDense modelMoE sparse model
Total parametersUsually smallerCan be very large
Parameters used per tokenAll parametersA small portion of experts
Inference costProportional to parameter countRelatively controllable
Training difficultyRelatively simpleRequires solving load balancing and other issues
Representative modelsGPT-3、LLaMASwitch Transformer、Mixtral、GPT-4

Gating Mechanism (Gating)

The core of MoE isgating network—it decides which experts each token should be sent to for processing.

The input to the gating network is the feature of the current token, and the output is the weight of each expert.

The usual approach is: select the Top-K experts with the highest weights (e.g., K=2 or K=8), send the token only to those experts, and then weight-sum their outputs.

The gating network itself is learnable — during training, it gradually learns "what type of content is suitable for which experts".

Expert Routing Algorithm

MoE has a unique challenge:Load balancing。

If the gating network always assigns most tokens to a few experts, the other experts won't get sufficient training, and model capacity is wasted.

Researchers have proposed various routing algorithms to solve this problem:

One approach is to add a "load balancing loss" to the gating network's loss function to encourage each expert to be used uniformly.

Another is to adopt more complex routing strategies, such as "capacity limits"—each expert has a maximum throughput, and when it's full, tokens are routed to the next most suitable expert.

Representative Model: Mixtral

At the end of 2023, Mixtral 8x7B released by Mistral AI truly brought MoE into the mainstream.

Mixtral has 8 experts with 7B parameters each, and selects Top-2 experts for each token, so each token actually uses about 14B parameters, but the total parameter count is 47B.

The result: Mixtral's inference speed and cost are comparable to a 14B dense model, but its performance approaches that of a 70B model.

This characteristic of "doing big things with small money" has made MoE a focus of industry attention.

Training and Inference Challenges

Although MoE is attractive, it also introduces additional complexity:

Training: it requires handling issues such as expert load balancing, communication overhead (in multi-machine distributed training), and expert dropout.

Inference: although only a few experts are activated per token, the entire model still needs to be loaded into GPU memory—this places higher demands on memory capacity.

However, researchers are using techniques such as model parallelism and dynamic expert offloading to alleviate these issues.

Examples

# ============================================
# A simple conceptual demonstration of the MoE gating mechanism
# Show how to select appropriate experts based on input
# ============================================

import random
from typing import List, Tuple


class Expert:
    """A simple expert model (conceptual demonstration)"""

    def __init__(self, expert_id: int, specialty: str):
        self.expert_id = expert_id
        self.specialty = specialty  # Expertise domains (e.g., "math", "code", "literature")
        # Simulated expert parameters (in practice, these are neural network weights)
        self.weights = [random.random() for _ in range(10)]

    def forward(self, x: List[float]) -> float:
        """Expert processes input and produces output"""
        # Simple weighted sum simulation (in practice, it's neural network computation)
        return sum(x[i] * self.weights[i % 10] for i in range(len(x)))


class MoEGate:
    """MoE gating network: decides which experts each input is assigned to"""

    def __init__(self, num_experts: int, top_k: int = 2):
        self.num_experts = num_experts
        self.top_k = top_k
        # Parameters of the gating network itself
        self.gate_weights = [[random.random() for _ in range(10)]
                             for _ in range(num_experts)]

    def compute_scores(self, x: List[float]) -> List[float]:
        """Compute each expert's score for the current input"""
        scores = []
        for expert_weights in self.gate_weights:
            score = sum(x[i] * expert_weights[i] for i in range(len(x)))
            scores.append(score)
        return scores

    def select_experts(self, x: List[float]) -> List[Tuple[int, float]]:
        """Select Top-K experts and return a list of (expert index, weight)"""
        scores = self.compute_scores(x)

        # Sort by score in descending order
        scored_experts = [(i, score) for i, score in enumerate(scores)]
        scored_experts.sort(key=lambda x: x[1], reverse=True)

        # Select Top-K
        top_experts = scored_experts[:self.top_k]

        # Perform softmax normalization on the Top-K weights (simplified version)
        total = max(0.0001, sum(score for _, score in top_experts))
        return [(idx, score / total) for idx, score in top_experts]


class SimpleMoE:
    """A simplified mixture-of-experts model"""

    def __init__(self, num_experts: int = 8, top_k: int = 2):
        # Create multiple experts
        specialties = ["Mathematics", "Code", "Literature", "History",
                      "Science", "Art", "Sports", "Philosophy"]
        self.experts = [Expert(i, specialties[i % len(specialties)])
                       for i in range(num_experts)]
        # Create gating network
        self.gate = MoEGate(num_experts, top_k)

    def forward(self, x: List[float]) -> float:
        """MoE forward propagation"""
        # 1. Gating network selects experts
        selected_experts = self.gate.select_experts(x)

        print(f"Input features: {[f'{v:.2f}' for v in x[:5]]}...")
        print(f"Selected experts:")
        for idx, weight in selected_experts:
            print(f" - Expert {idx} ({self.experts[idx].specialty}), weight: {weight:.3f}")

        # 2. Selected experts process the input
        expert_outputs = []
        for idx, weight in selected_experts:
            output = self.experts[idx].forward(x)
            expert_outputs.append((output, weight))

        # 3. Weighted merge of expert outputs
        final_output = sum(output * weight for output, weight in expert_outputs)

        return final_output


# ============================================
# Demonstrate the MoE workflow
# ============================================

# Create an MoE model with 8 experts, selecting Top-2
moe = SimpleMoE(num_experts=8, top_k=2)

# Simulate several different types of inputs (represented by random vectors)
print("=" * 50)
print("EXAMPLE MoE Concept Demo")
print("=" * 50)

# Input 1: Simulate "Mathematics" type input
input1 = [0.9, 0.8, 0.1, 0.1, 0.2, 0.1, 0.1, 0.1, 0.1, 0.1]
output1 = moe.forward(input1)
print(f"MoE output: {output1:.4f}")
print()

# Input 2: Simulate "Code" type input
input2 = [0.1, 0.2, 0.9, 0.8, 0.1, 0.1, 0.1, 0.1, 0.1, 0.1]
output2 = moe.forward(input2)
print(f"MoE output: {output2:.4f}")
print()

# Input 3: Simulate "Literature" type input
input3 = [0.1, 0.1, 0.1, 0.1, 0.9, 0.8, 0.1, 0.1, 0.1, 0.1]
output3 = moe.forward(input3)
print(f"MoE output: {output3:.4f}")

print()
print("=" * 50)
print("Hint: In real MoE, the gating network gradually learns through training")
print(" to learn what types of inputs are suitable for which experts.")
print("=" * 50)

The core insight of MoE:Not all parameters need to work at the same time. Let each expert focus on what they are good at, and use the gating network for scheduling, which can expand total capacity while controlling inference cost.


Long Context Technology

Early GPT-3 had a context window of only 2K tokens (about 1,500 Chinese characters); today's mainstream models have reached 8K, 32K, 128K, or even 1M tokens.

Long context is not simply "stretching the window" — it requires solving a series of technical challenges.

Core Challenges of Long Context

The most direct problem iscomputational complexity。

The self-attention mechanism of Transformer has O(n²) complexity — when the context length doubles, the computation quadruples.

Expanding the window from 2K to 128K increases computation by 4096 times — this is computationally infeasible.

The second challenge isPosition Encoding。

The sinusoidal position encoding or learnable position encoding used by the original Transformer has poor extrapolation ability beyond the training length—just like a person who has only trained on a 100-meter track suddenly being asked to run 1000 meters, they would be very uncomfortable.

The third challenge isMemory effect。

Even if the model can technically process long texts, it may "read and forget"—the content in the middle is hard to effectively utilize.

RoPE Extrapolation Methods (YaRN / LongRoPE)

Rotary Position Embedding (RoPE) is one of the most popular position encoding schemes currently.

The core idea of RoPE is:Encoding position information as rotation angles。

The advantage of this is: relative position relationships are naturally encoded—if two tokens are k positions apart, their rotation angles differ by k units.

However, the extrapolation effect of the original RoPE still degrades beyond the training length.

Researchers have proposed improvement methods:

YaRN(Yet Another RoPE Scaling)—by finely adjusting RoPE's scaling factor, the model can maintain good performance even beyond the training length.

LongRoPE—A method proposed by Meta that extends the LLaMA 2 model's context window from 4K to 128K or even 256K through progressive expansion and fine-tuning.

The common idea behind these techniques is:Without retraining the entire model, only adjusting some parameters of the position encodingcan achieve a significant expansion of the context window.

Sparse Attention (Sliding Window / Longformer)

Another approach to solving the O(n²) complexity is:Not having every token attend to all other tokens。

Sliding Window Attention—Each token only attends to tokens within a fixed window before and after it (e.g., 4096 tokens before and after).

This is like when reading a novel, you only remember the content of the most recent few pages, rather than recalling the entire book with every word you read.

Longformer—Combining sliding window and global attention, certain special tokens (such as <[BOS_never_used_51bce0c785ca2f68081bfa7d91973934]> and summary markers) can attend to all tokens, while other tokens only attend to local information.

This allows handling long texts while ensuring global information is not lost.

State Space Model: Mamba

At the end of 2023, the emergence of Mamba brought a completely new approach.

Mamba belongs toState Space Model (SSM)in the category, its complexity is O(n) — if the context length doubles, the computation only doubles as well.

Mamba's core innovation isSelective State Space— allowing the model to dynamically decide based on the input which information to remember and which to ignore.

This is like taking notes: you don't write down every word, but selectively record the key points.

Mamba's advantages are:

1. It can theoretically handle infinitely long contexts

2. Fast inference speed and low memory usage

3. Excellent performance on long-sequence tasks

However, Mamba also has limitations: it is still slightly inferior to mature Transformers on short-text tasks, and its ecosystem is relatively underdeveloped.

Technical solutionsComplexityRepresentative modelsApplicable scenarios
Full-attention TransformerO(n²)GPT-3、LLaMAShort to medium-length text
RoPE extrapolationO(n²)Claude、GPT-4Medium-length text
Sparse attentionO(n) or O(n log n)Longformer、MistralDocument-level long text
State Space ModelO(n)Mamba、HyenaUltra-long sequences (code, genomes, etc.)

The evolution direction of long-context technology:Finding the best balance among performance, speed, and memory. Different scenarios may require different technical solutions.


Inference-Time Compute Scaling

The training and inference of traditional large models are separated — training consumes a large amount of computation, while inference uses fixed forward computation.

The idea of inference-time compute scaling is:Allowing the model to "think" more steps during inference, trading extra computation for better results.

The Approach of o1 / DeepSeek-R1

In 2024, OpenAI's o1 and DeepSeek's DeepSeek-R1 represented an important breakthrough in this direction.

The core of these models is:Having the model first perform a "reasoning chain of thought" before giving the final answer。

Specifically:

1. After receiving a question, the model does not answer directly, but first generates a series of "thought processes" (visible or hidden to the user)

2. During the thinking process, the model can try different approaches, correct its own mistakes, and break down complex problems

3. After thinking sufficiently, it then gives the final answer

This is like solving a math problem: you don't write the answer directly, but first derive it step by step on scratch paper, and organize the answer after confirming it is correct.

Slow Thinking vs Fast Thinking

This direction echoes psychologist Kahneman's "dual-system theory":

System 1 (fast thinking)—Intuitive, fast, and effortless. It corresponds to a single forward pass in traditional models.

System 2 (slow thinking)—Analytical, slow, and effortful. It corresponds to multi-step thinking during inference.

Previous large models were mainly "System 1," while models like o1 have begun to introduce "System 2" capabilities.

It is important to note:Slow thinking requires more time and cost—A request that thinks for 10 seconds may cost 10 times as much as an immediate answer.

Best-of-N Sampling

Another simpler inference-time scaling method is Best-of-N (also called Self-Consistency).

The approach is straightforward:

1. Have the model generate N different answers to the same question

2. Use some method to select the best one (e.g., majority voting, model self-evaluation, external validation)

This is like answering a multiple-choice question on a test: you don't just guess once; you try multiple solutions and then choose the most consistent answer.

The advantage of Best-of-N is that it is simple, general, and requires almost no extra training—but the disadvantage is the high cost (it needs to generate N times).

Monte Carlo Tree Search in LLM

Another active direction is to bring from reinforcement learning theMonte Carlo Tree Search (MCTS)into large models.

AlphaGo used MCTS to achieve superhuman performance.

The idea of MCTS is:

1. View the problem as a "search tree"—each node is a thinking state, each edge is a thinking step.

2. Explore different thinking paths through multiple simulations.

3. Select the most promising path based on the exploration results.

Combining MCTS with LLMs is a very promising direction—it can both leverage LLMs' rich knowledge and bring out MCTS's planning capabilities.

MethodCore ideaAdvantagesDisadvantages
Chain of Thought (CoT)Have the model think step by stepSimple and effectiveIncreases token consumption
o1 / R1 styleDeep chain of thought + reinforcement learningStrong reasoning abilityComplex training
Best-of-NChoose one from manySimple and generalHigh cost
MCTS + LLMSearch for the optimal thinking pathStrong planning abilityComplex implementation

The core trade-off of inference-time compute scaling:How much time and cost are you willing to spend to get better answers?Different scenarios (casual chat vs. mathematical proof) may require different settings.


Latest Advances in Multimodal Large Models

From GPT-4V to Gemini to GPT-4o, multimodal capabilities are evolving rapidly—current models can not only read text, but also see images, understand videos, hear audio, and even generate multimedia content.

GPT-4o / Gemini 1.5 Pro

The new generation of multimodal models in 2024 has several notable features:

Smoother interaction—GPT-4o supports voice conversation, can understand images and videos in real time, and the interaction experience is closer to a real person.

Stronger video understanding—Gemini 1.5 Pro can process videos up to 1 hour long and accurately answer questions about video content.

Native multimodal generation—Models are no longer "look first, then answer," but can simultaneously process and generate content in multiple modalities.

The capabilities of these models are often astonishing: give it a hand-drawn sketch, and it can help you write code to implement it; give it a photo of a physics problem, and it can explain it step by step; give it a video, and it can analyze characters' emotions and plot development.

Native Multimodal Training

Early multimodal models were typically "stitched together"—pretrain a language model, train a vision encoder, and finally combine them.

The new generation of models is moving towardnative multimodal training—processing text, images, audio, and other modalities simultaneously from the very beginning of training.

This brings several benefits:

1. Tighter cross-modal associations—the model can truly understand "this word corresponds to this region in the image"

2. More natural multi-turn interaction—can alternately present text and images to form a coherent conversation

3. Emergent new capabilities—such as "telling a story from a picture" or "identifying objects by sound"

Breakthroughs in Video Understanding

Video understanding is a direction that has advanced rapidly in the past year.

Early approaches split videos into independent frames and encoded each frame separately—this loses temporal continuity.

New ideas include:

Temporal attention—allowing the model to not only focus on spatial content but also track changes over time.

Uniform sampling—uniformly sampling key frames from long videos, covering the whole while controlling computational cost.

Video tokenization—treating video as a kind of "language" and processing it in a way similar to text.

Current models can already: understand complex video tutorials, analyze key moments in sports matches, and even learn operational skills from videos.

ModelCapability focusRelease time
GPT-4VImage understanding2023
Gemini 1.0Native multimodal2023
Claude 3Visual reasoning2024
Gemini 1.5 ProLong-form video understanding2024
GPT-4oReal-time multimodal interaction2024

Future directions of multimodality:From "seeing and hearing" to "true understanding". The next-stage challenges are physical reasoning, causal understanding, and abstract thinking under multimodality.


AI Agent Research Frontiers

The goal of AI Agent (intelligent agent) is to enable AI to "do things" like humans — having goals, planning, executing, and reflecting.

World Model

The world model is a core component of the Agent — it is the Agent's "mental model" of the world.

Simply put, the world model allows the Agent to answer:

"If I do this, what will happen next?"

This is like playing chess: you don't move directly, but first simulate a few steps in your mind and predict your opponent's reaction.

Research directions for world models include:

Pretrained world models— pretrained on large amounts of interaction data, giving the Agent common-sense physics knowledge.

Imagination planning— letting the Agent try multiple approaches in "imagination" and choose the best one.

Causal understanding— letting the Agent not only know "what happened" but also understand "why it happened."

Embodied Intelligence (Embodied AI)

Embodied AI focuses on:AI that has a "body" in the physical or virtual world。

For example:

— robots grasping objects and navigating in the real world

— AI building complex structures in Minecraft

— digital humans interacting with the environment in virtual settings

The challenge of embodied AI is that the physical world is continuous, unpredictable, and noisy — quite different from the discrete world of text.

But it is precisely this "sense of reality" that makes embodied AI an important path for learning common sense.

The Importance of Code Execution Capability

A key capability of modern Agents is:Being able to write code and execute code。

Code is the Agent's "universal tool":

— when calculation is needed, write a piece of Python code

— when data analysis is needed, write SQL or Pandas

— when task automation is needed, write scripts to call various APIs

— when complex reasoning is needed, write code to implement algorithms

Code execution capability is essentially a form of "extensibility" — it allows the Agent to create new tools as needed, rather than being limited to preset capabilities.

This is also why models like Claude and GPT-4 put great effort into strengthening their coding abilities.

The core challenge of AI Agents:How to convert the "knowledge" of large models into "actions" in the real world. This requires planning ability, tool use ability, and the ability to learn from trial and error.


Few-Shot and Zero-Shot Learning

Traditional machine learning requires a large amount of labeled data to learn a task well.

The paradigm shift brought by large models is:Give a few examples, or even just instructions, and the model can learn new tasks。

The Principle of In-Context Learning

In-Context Learning (context learning) is one of the most amazing capabilities of large models.

Simply put: you don't need to modify model parameters; just give a few examples in the input, and the model can learn new tasks.

For example:

Input:

apple → fruit

carrot → vegetable

banana → ?

The model can automatically understand the pattern and output "fruit".

But how does In-Context Learning actually work?

Researchers have proposed several explanations:

Implicit Gradient Descent— the in-context examples are equivalent to performing a kind of "gradient update", but it only takes effect during inference.

Task Retrieval— the model retrieves similar tasks from training data, then reuses the solution for that task.

Bayesian Inference— the model performs Bayesian updates over the space of possible hypotheses based on the examples.

There is no completely consistent answer yet, but one thing is certain:Models learn "how to learn" during pretraining。

The Impact of Instruction Tuning

Instruction Tuning is a key technique for making models better at following human instructions.

The approach is simple: collect a large number of "instruction-task" pairs and use them to fine-tune the model.

For example:

Instruction: Translate the following sentence into French

Input: Hello, how are you?

Output: Bonjour, comment ça va?

After Instruction Tuning, the model becomes more "obedient"—more likely to do what the instruction asks, rather than generating arbitrarily.

More importantly, Instruction Tuning bringsgeneralization ability— the model performs well on seen tasks, and can also perform well on completely unseen new tasks.

This is like a good student: he doesn't memorize mechanically, but has learned "how to understand and execute instructions."

Learning paradigmsData requiredRepresentative methodsApplicable scenarios
Traditional supervised learningLarge amounts of labeled dataFine-tuning the full modelHigh-value tasks with sufficient data
Few-shot learningA few examplesIn-Context LearningQuick trial for new tasks
Zero-shot learningOnly instructionsInstruction TuningSimple, low-risk tasks
Chain-of-Thought (CoT)Examples with reasoning stepsStep-by-step promptingComplex reasoning tasks

The core value of few-shot learning:Lowering the barrier to applying AINow, you don't need to be a machine learning expert or collect large amounts of data—just being able to write prompts is enough to let AI help you get things done.


AI and Scientific Discovery

AI is not only changing our lives, but also changing scientific research itself—from protein structures to mathematical theorems, from drug discovery to material design, AI is becoming a powerful assistant for scientists.

The Significance of AlphaFold

In 2020, AlphaFold was a milestone: it solved the "protein folding problem" that had puzzled the biology community for 50 years.

Simply put, a protein is a chain of amino acids, but it folds into a complex 3D structure—and this structure determines its function.

Previously, resolving a protein structure could take years and a large amount of experimental funding.

Now, AlphaFold can predict an accurate 3D structure within minutes.

The impact of AlphaFold is profound:

— helping understand disease mechanisms

— accelerating drug development

— designing entirely new proteins

More importantly, AlphaFold has proven that:AI can make substantial contributions at the boundaries of human knowledge。

AI for Mathematics

Mathematical research is considered the pinnacle of human intelligence—but AI is also beginning to make its mark in this field.

Recent advances include:

Theorem-proving assistants— Proof assistants such as Lean and Isabelle, combined with large models, can help mathematicians formalize proofs.

Conjecture discovery— AI can discover patterns in mathematical objects and propose new conjectures.

Symbolic reasoning— AI can perform symbolic computation, simplification, and derivation.

For example, DeepMind's AlphaTensor discovered more efficient matrix multiplication algorithms; Meta's AI helped mathematicians make new discoveries in the field of topology.

Current AI cannot yet conduct mathematical research independently, but it has already become a "collaborative partner" for mathematicians.

Drug Discovery

Drug development is a long and expensive process—on average taking 10 years and $1 billion.

AI is changing this situation:

Target discovery—AI analyzes biological data to find disease-related protein targets.

Molecular generation—AI generates novel candidate molecules and predicts their properties.

ADMET prediction—AI predicts drug absorption, metabolism, toxicity, etc.

Synthetic route design—AI designs optimal routes for organic synthesis.

Multiple AI drug companies have entered clinical trial stages—although no AI-discovered drug has been marketed yet, the potential in this direction is enormous.

Scientific fieldAI applicationsRepresentative achievements
Structural biologyProtein structure predictionAlphaFold、RoseTTAFold
MathematicsTheorem proving, conjecture discoveryAlphaTensor、Lean AI
ChemistryMolecular generation, reaction predictionAlphaFold for ligands
Drug developmentTarget discovery, candidate designMultiple companies entering clinical trials
Materials scienceNew material designBattery materials, catalysts

Core insights of AI-empowered scientific discovery:Scientific research is essentially "searching in a vast possibility space"—and what AI excels at is precisely finding patterns in massive data and high-dimensional spaces.

Other extensions