AI Agent Tools and External Integration

This chapter introduces how Agents interact with the external world.

Tool invocation capability is the key to connecting Agents with the external environment.

By integrating various tools, Agents can execute code, access APIs, operate file systems, and more.


Computer Use

Computer Use enables Agents to operate computer interfaces like humans.

This includes graphical interfaces such as browsers and desktop applications.

This is an important step toward achieving artificial general intelligence (AGI).

Core Capabilities

Screenshot Parsing: Understand the content on the screen and identify interactive elements.

GUI Element Recognition: Recognize interface elements such as buttons, input boxes, and menus.

Mouse/Keyboard Control: Simulate human operations, performing actions such as clicking and typing.

Browser Automation: Web navigation, form filling, searching, etc.

Working Principle

The workflow of Computer Use can be summarized as: observe, understand, decide, and execute.

First step, screenshot. Capture the content of the current screen or window.

Second step, visual analysis. Use a vision model to analyze the screenshot and identify interface elements and states.

Third step, decision-making. Determine the next action based on the task goal.

Fourth step, execution. Perform actions such as mouse clicks and keyboard input.

Repeat the loop until the task is completed.

Code Implementation

Basic Implementation of Computer Use Agent

class ComputerUseAgent:
    """
Implementation of Computer Use Agent
Able to operate computer interfaces like a human
    """

   
    def __init__(self, vision_model, action_executor, planner):
        # Vision Model: Analyze Screenshot
        self.vision_model = vision_model
        # Action Executor: Perform mouse and keyboard operations
        self.action_executor = action_executor
        # Planner: Decide next action
        self.planner = planner
        # Maximum execution steps
        self.max_steps = 100
   
    def observe(self, screenshot):
        """
Parse screenshot, identify interactive elements
:param screenshot: Screenshot image
:return: List of interface elements
        """

        # Use vision model to analyze screenshot
        analysis = self.vision_model.analyze(screenshot)
       
        # Return recognized interface elements
        # Each element includes: type, location, content, interactivity
        return analysis.elements
   
    def act(self, action):
        """
Execute action
:param action: Action description, e.g., {"type": "click", "x": 100, "y": 200}
        """

        return self.action_executor.execute(action)
   
    def run(self, task):
        """
Run task loop
:param task: Task description
:return: Task result
        """

        # Initialize task state
        self.planner.set_task(task)
       
        for step in range(self.max_steps):
            # Step 1: Capture current screen
            screenshot = self.get_screen()
           
            # Step 2: Observe - Identify interface elements
            elements = self.observe(screenshot)
           
            # Step 3: Decide - Determine next action
            action = self.planner.decide_action(elements)
           
            # Check if task is complete
            if action.is_final:
                return action.result
           
            # Step 4: Execute action
            self.act(action)
           
            # Optional: Wait for interface update
            self.wait_for_update()
       
        return "Reached maximum step limit"


class VisionModel:
    """Vision Model: Analyze Screenshot"""
   
    def analyze(self, screenshot):
        """
Analyze screenshot
Return list of interface elements
        """

        # Use multimodal model for analysis
        prompt = """
Analyze this screenshot and identify all interactive interface elements.
Including: buttons, input fields, links, menus, etc.

For each element, please provide:
1. Type (button, input, link, etc.)
2. Location (bounding box coordinates)
3. Content (button text, input placeholder, etc.)
4. Interactivity (visibility, enabled state)
"""

        result = self.vision_model.analyze_image(screenshot, prompt)
        return ScreenAnalysisResult(elements=result.elements)


class ActionExecutor:
    """Action Executor: Perform mouse and keyboard operations"""
   
    def execute(self, action):
        """
Execute action
:param action: Action object
        """

        if action.type == "click":
            self.mouse.click(action.x, action.y)
        elif action.type == "type":
            self.keyboard.type_text(action.text)
        elif action.type == "scroll":
            self.mouse.scroll(action.direction, action.amount)
        elif action.type == "press":
            self.keyboard.press_key(action.key)
       
        return ActionResult(success=True)


class Planner:
    """Planner: Decide next action"""
   
    def __init__(self, llm):
        self.llm = llm
        self.task = None
        self.history = []
   
    def set_task(self, task):
        """Set current task"""
        self.task = task
        self.history = []
   
    def decide_action(self, elements):
        """
Determine next action based on current interface state
        """

        prompt = f"""
Current task: {self.task}

History of executed actions:
{self.history}

Interactive elements on the current interface:
{elements}

Please decide the next action.
If the task is complete, return is_final=True.
Otherwise, return the action to be executed (type, coordinates, parameters, etc.).
"""

        response = self.llm.generate(prompt)
        return Action.parse(response)

Application Scenarios

Computer Use is particularly suitable for the following scenarios:

Web automation: such as automatically filling forms, scraping data, and performing web operations.

Desktop application operations: such as opening files, editing documents, and operating software.

Test automation: such as automatically executing UI tests and regression tests.

Note: Computer Use is still under development, and has issues such as slow execution and occasional errors. For scenarios with a clear API, directly calling the API is usually more efficient than Computer Use.


Detailed Explanation of MCP Protocol

MCP (Model Context Protocol) is an open standard protocol.

It enables AI models to securely connect with external tools and data sources.

MCP's design goal is to become the "USB interface" of the AI field.

Core Design Principles

Standardization: Unified protocol format, allowing different Agents and tools to interoperate.

Security: Clear permission control, Agents can only access authorized resources.

Extensibility: Easy to add new tools and data sources.

Architecture Components

MCP Host: Host environment for running AI applications, such as Claude Desktop, AI coding assistants, etc.

MCP Client: A client that maintains a 1:1 connection with the MCP Server.

MCP Server: Server program that provides tools and resources.

Communication Flow

First step, connection establishment. The Client establishes a connection with the Server and exchanges capability information.

Second step, tool discovery. The Client queries the Server for available tools and resources.

Third step, tool invocation. The Client sends a tool invocation request, and the Server executes it and returns the result.

Fourth step, resource access. The Client can read and write resources provided by the Server.

MCP Server Implementation Example

MCP Server Implementation

from mcp.server import Server
from mcp.types import Tool, Resource
import asyncio

# Create MCP Server instance
app = Server("filesystem")

# Define file system tools
@app.list_tools()
async def list_tools():
    """List all available tools"""
    return [
        Tool(
            name="read_file",
            description="Read file content",
            inputSchema={
                "type": "object",
                "properties": {
                    "path": {
                        "type": "string",
                        "description": "File path"
                    },
                    "encoding": {
                        "type": "string",
                        "description": "File encoding, defaults to utf-8",
                        "default": "utf-8"
                    }
                },
                "required": ["path"]
            }
        ),
        Tool(
            name="write_file",
            description="Write file content",
            inputSchema={
                "type": "object",
                "properties": {
                    "path": {
                        "type": "string",
                        "description": "File path"
                    },
                    "content": {
                        "type": "string",
                        "description": "File content"
                    }
                },
                "required": ["path", "content"]
            }
        ),
        Tool(
            name="list_directory",
            description="List directory contents",
            inputSchema={
                "type": "object",
                "properties": {
                    "path": {
                        "type": "string",
                        "description": "Directory path"
                    }
                }
            }
        )
    ]

@app.call_tool()
async def call_tool(name, arguments):
    """Execute tool call"""
    if name == "read_file":
        return await read_file(arguments["path"], arguments.get("encoding", "utf-8"))
    elif name == "write_file":
        return await write_file(arguments["path"], arguments["content"])
    elif name == "list_directory":
        return await list_directory(arguments["path"])
    else:
        raise ValueError(f"Unknown tool: {name}")

async def read_file(path, encoding):
    """Read file"""
    with open(path, "r", encoding=encoding) as f:
        content = f.read()
    return [{"type": "text", "text": content}]

async def write_file(path, content):
    """Write file"""
    with open(path, "w", encoding="utf-8") as f:
        f.write(content)
    return [{"type": "text", "text": "File written successfully"}]

async def list_directory(path):
    """List directory"""
    import os
    entries = os.listdir(path)
    return [{"type": "text", "text": "\n".join(entries)}]

# Run server
if __name__ == "__main__":
    import mcp.server.stdio
    asyncio.run(mcp.server.stdio.serve(app))

API Integration Patterns

There are three common ways for an Agent to integrate external APIs.

REST API is the most commonly used Web API style.

GraphQL provides more flexible data querying capabilities.

gRPC is suitable for remote procedure calls in high-performance scenarios.

REST API Integration

REST (Representational State Transfer) is a Web API design style.

It uses HTTP methods (GET, POST, PUT, DELETE) for operations.

REST API Client Implementation

import requests
from typing import Dict, Any, Optional

class RESTAPIClient:
    """
REST API Client
Encapsulates HTTP requests and provides a concise calling interface
    """

   
    def __init__(self, base_url: str, headers: Optional[Dict] = None):
        """
Initialize API client
:param base_url: API base URL
:param headers: default request headers
        """

        self.base_url = base_url.rstrip("/")
        self.session = requests.Session()
       
        # Set default request headers
        default_headers = {
            "Content-Type": "application/json",
            "Accept": "application/json"
        }
        if headers:
            default_headers.update(headers)
        self.session.headers.update(default_headers)
   
    def call(
        self,
        method: str,
        endpoint: str,
        params: Optional[Dict] = None,
        data: Optional[Dict] = None,
        **kwargs
    ) -> Dict[Any, Any]:
        """
Send API request
:param method: HTTP method (GET, POST, PUT, DELETE)
:param endpoint: API endpoint
:param params: URL query parameters
:param data: request body data
:return: response data
        """

        url = f"{self.base_url}/{endpoint.lstrip('/')}"
       
        response = self.session.request(
            method=method.upper(),
            url=url,
            params=params,
            json=data,
            **kwargs
        )
       
        # Check HTTP status code
        response.raise_for_status()
       
        # Parse response
        if response.content:
            return response.json()
        return {}
   
    def get(self, endpoint: str, params: Optional[Dict] = None):
        """GET request"""
        return self.call("GET", endpoint, params=params)
   
    def post(self, endpoint: str, data: Dict):
        """POST request"""
        return self.call("POST", endpoint, data=data)
   
    def put(self, endpoint: str, data: Dict):
        """PUT request"""
        return self.call("PUT", endpoint, data=data)
   
    def delete(self, endpoint: str, params: Optional[Dict] = None):
        """DELETE request"""
        return self.call("DELETE", endpoint, params=params)


class APIAgent:
    """
    API Agent
Calling external APIs through natural language
    """

   
    def __init__(self, api_client: RESTAPIClient):
        self.api_client = api_client
        self.tool_descriptions = self.define_tools()
   
    def define_tools(self):
        """
Define tools available to the Agent
Return tool descriptions so the LLM understands how to call them
        """

        return [
            {
                "name": "get_weather",
                "description": "Get weather information for a specified city",
                "parameters": {
                    "type": "object",
                    "properties": {
                        "city": {
                            "type": "string",
                            "description": "City name, e.g., Beijing, Shanghai"
                        }
                    },
                    "required": ["city"]
                }
            },
            {
                "name": "get_forecast",
                "description": "Get weather forecast for a specified city",
                "parameters": {
                    "type": "object",
                    "properties": {
                        "city": {
                            "type": "string",
                            "description": "City name"
                        },
                        "days": {
                            "type": "integer",
                            "description": "Forecast days, defaulting to 3",
                            "default": 3
                        }
                    },
                    "required": ["city"]
                }
            }
        ]
   
    def execute_tool(self, tool_name: str, arguments: Dict):
        """
Execute tool call
:param tool_name: tool name
:param arguments: tool arguments
:return: execution result
        """

        if tool_name == "get_weather":
            return self.api_client.get(f"/weather/{arguments['city']}")
        elif tool_name == "get_forecast":
            params = {"days": arguments.get("days", 3)}
            return self.api_client.get(f"/forecast/{arguments['city']}", params=params)
        else:
            raise ValueError(f"Unknown tool: {tool_name}")

API Integration Best Practices

Error Handling: Gracefully handle network errors, timeouts, service unavailability, etc.

Retry Mechanism: For transient errors, implement exponential backoff retry.

Rate Limiting: Respect the API's rate limits to avoid being banned.

Caching Strategy: Cache frequently requested data to reduce API calls.

Security Considerations: Do not hardcode sensitive information such as API keys; use environment variables to manage them.


Chapter Summary

This chapter introduces the core technologies for Agents to interact with the external world.

Computer UseEnables Agents to operate GUI interfaces, suitable for scenarios that require interaction with graphical applications.

MCP ProtocolProvides a standardized way to integrate tools, serving as the "USB interface" of the Agent era.

API IntegrationIs the most common way to integrate external systems, including REST, GraphQL, gRPC, etc.

Choosing the appropriate integration method depends on the specific scenario.

For graphical interface operations, Computer Use is a universal solution.

For standardized tool integration, MCP is the future development direction.

For Web API integration, REST API remains the most mainstream choice.

Other Extensions