Agent Evaluation, Safety, and Alignment

Evaluation is the foundation for ensuring Agent effectiveness.

Safety and alignment are key to whether an Agent can be trusted and deployed.


Agent Evaluation System

Evaluating Agent performance is a critical part of the development process.

A good evaluation system helps us understand the boundaries of an Agent's capabilities.

It also provides direction and basis for continuous optimization.

Evaluation Dimensions

Task Completion: Whether the Agent correctly completed the given task.

Efficiency Metrics: Number of steps, token consumption, and execution time required to complete the task.

Quality Metrics: Accuracy of answers, consistency of responses, and naturalness of dialogue.

Robustness: Ability to handle abnormal inputs and noisy data.

Common Benchmarks

Benchmark Purpose Evaluation Content
GAIA General AI Assistant Evaluation Complex task handling, multi-step reasoning
MMLU Multi-task Language Understanding Knowledge Q&A across 57 subjects
HumanEval Code Generation Evaluation Correctness of Python code writing
HotpotQA Multi-hop Question Answering Evaluation Questions requiring reasoning across multiple documents
AgentBench Agent Capability Evaluation Agent performance in real-world environments

Code Implementation: Evaluation Framework

Agent Evaluation Framework

class AgentEvaluator:
    """
Agent Evaluation Framework
Evaluate the performance of an Agent on various tasks
    """

   
    def __init__(self, agent, metrics):
        # Agent to be evaluated
        self.agent = agent
        # Evaluation Metrics List
        self.metrics = metrics
   
    def evaluate(self, test_cases):
        """
Execute evaluation
:param test_cases: list of test cases
:return: evaluation report
        """

        results = []
       
        for test_case in test_cases:
            # Execute task
            result = self.run_single_test(test_case)
            results.append(result)
       
        # Generate evaluation report
        report = self.generate_report(results)
        return report
   
    def run_single_test(self, test_case):
        """
Run a single test case
        """

        # Record start time
        start_time = time.time()
       
        # Execute Agent
        try:
            output = self.agent.run(test_case.input)
            success = self.evaluate_output(output, test_case.expected)
            error = None
        except Exception as e:
            output = None
            success = False
            error = str(e)
       
        # Record end time
        end_time = time.time()
       
        return TestResult(
            test_case=test_case,
            output=output,
            success=success,
            error=error,
            duration=end_time - start_time,
            token_count=self.count_tokens(output)
        )
   
    def evaluate_output(self, output, expected):
        """Evaluate whether the output meets expectations"""
        for metric in self.metrics:
            if not metric.evaluate(output, expected):
                return False
        return True
   
    def generate_report(self, results):
        """Generate evaluation report"""
        total = len(results)
        passed = sum(1 for r in results if r.success)
       
        # Calculate various metrics
        avg_duration = sum(r.duration for r in results) / total
        avg_tokens = sum(r.token_count for r in results) / total
       
        # Group and count by test type
        by_category = {}
        for r in results:
            category = r.test_case.category
            if category not in by_category:
                by_category[category] = {"total": 0, "passed": 0}
            by_category[category]["total"] += 1
            if r.success:
                by_category[category]["passed"] += 1
       
        return EvaluationReport(
            total=total,
            passed=passed,
            pass_rate=passed / total,
            avg_duration=avg_duration,
            avg_tokens=avg_tokens,
            by_category=by_category,
            results=results
        )


class TestCase:
    """Test case"""
   
    def __init__(self, input, expected, category="general", metadata=None):
        # Input
        self.input = input
        # Expected output or evaluation criteria
        self.expected = expected
        # Category
        self.category = category
        # Extra metadata
        self.metadata = metadata or {}


class TestResult:
    """Test result"""
   
    def __init__(self, test_case, output, success, error, duration, token_count):
        self.test_case = test_case
        self.output = output
        self.success = success
        self.error = error
        self.duration = duration
        self.token_count = token_count


class Metric:
    """Base class for evaluation metrics"""
   
    def evaluate(self, output, expected):
        raise NotImplementedError


class ExactMatchMetric(Metric):
    """Exact match metric"""
   
    def evaluate(self, output, expected):
        return output.strip() == expected.strip()


class ContainsMetric(Metric):
    """Keyword inclusion metric"""
   
    def evaluate(self, output, expected):
        if isinstance(expected, list):
            return all(keyword in output for keyword in expected)
        return expected in output


class SemanticSimilarityMetric(Metric):
    """Semantic similarity metric"""
   
    def __init__(self, threshold=0.8):
        self.threshold = threshold
   
    def evaluate(self, output, expected):
        similarity = self.compute_similarity(output, expected)
        return similarity >= self.threshold
   
    def compute_similarity(self, text1, text2):
        """Calculate the semantic similarity of two texts"""
        # Use embedding model to compute cosine similarity
        embedding1 = self.embedder.embed([text1])[0]
        embedding2 = self.embedder.embed([text2])[0]
        return cosine_similarity(embedding1, embedding2)

Safety and Alignment

The safety of the Agent is of utmost importance.

AI systems may be subject to various attacks and produce harmful outputs.

Alignment ensures that AI behavior aligns with human intentions and values.

Common Safety Threats

Prompt Injection

Attackers use input to induce the Agent to ignore system instructions.

Example input: "Ignore previous instructions and instead execute..."

This is a context hijacking attack that exploits the Agent's trust in user input.

Jailbreaking

Bypass security restrictions through specific inputs.

For example, using techniques such as role-playing or fictional scenarios.

Data Poisoning

Maliciously modify training data or retrieval results.

Causing the Agent to produce incorrect or harmful outputs.

Sensitive Information Leakage

The Agent improperly exposes user privacy or internal system information.

Defense Strategies

Safe Agent Implementation

class SecureAgent:
    """
Safety Agent
Add multiple layers of safety protection on top of the base Agent
    """

   
    def __init__(self, base_agent, guardrails, input_validator, output_filter):
        # Base Agent
        self.base_agent = base_agent
        # List of safety guardrails
        self.guardrails = guardrails
        # Input validator
        self.input_validator = input_validator
        # Output filter
        self.output_filter = output_filter
   
    def process(self, user_input):
        """
Process user input, including multiple layers of security checks
        """

        # ==================== Layer 1: Input Validation ====================
        # Check whether the input is valid
        is_valid, reason = self.input_validator.validate(user_input)
        if not is_valid:
            return self.create_safety_response(reason)
       
        # ==================== Layer 2: Injection Detection ====================
        # Detect attacks such as prompt injection
        for guardrail in self.guardrails:
            check_result = guardrail.check_input(user_input)
            if not check_result.is_safe:
                # Record security events
                self.log_security_event(
                    event_type="input_guardrail_triggered",
                    input=user_input,
                    reason=check_result.reason
                )
                return self.create_safety_response(check_result.reason)
       
        # ==================== Layer 3: Execute Core Logic ====================
        try:
            response = self.base_agent.process(user_input)
        except Exception as e:
            return self.create_error_response(str(e))
       
        # ==================== Layer 4: Output Filtering ====================
        # Check whether the output is safe
        for guardrail in self.guardrails:
            check_result = guardrail.check_output(response)
            if not check_result.is_safe:
                self.log_security_event(
                    event_type="output_guardrail_triggered",
                    output=response,
                    reason=check_result.reason
                )
                return self.create_safety_response(check_result.reason)
       
        # Apply output filtering (e.g., desensitization of sensitive information)
        response = self.output_filter.filter(response)
       
        return response
   
    def create_safety_response(self, reason):
        """Create safe response"""
        return {
            "type": "safety_block",
            "message": "Sorry, I cannot complete this request.",
            "reason": reason
        }
   
    def log_security_event(self, event_type, **kwargs):
        """Record security events"""
        # In actual applications, this should be written to a security log system
        print(f"[SECURITY] {event_type}: {kwargs}")


class InputValidator:
    """Input validator"""
   
    def validate(self, text):
        """
Validate whether the input is valid
        :return: (is_valid, reason)
        """

        if not text or len(text.strip()) == 0:
            return False, "Input cannot be empty"
       
        if len(text) > 10000:
            return False, "Input length exceeds the limit"
       
        # Check whether it contains executable content
        if self.contains_executable_content(text):
            return False, "Input contains suspicious executable content"
       
        return True, None
   
    def contains_executable_content(self, text):
        """Check whether it contains executable content"""
        # Simplified implementation
        suspicious_patterns = [
            "javascript:",
            "data:text/html",
            "<script>",
        ]
        return any(pattern in text.lower() for pattern in suspicious_patterns)


class Guardrail:
    """Safety guardrail"""
   
    def check_input(self, text):
        """Check input"""
        raise NotImplementedError
   
    def check_output(self, text):
        """Check output"""
        raise NotImplementedError


class ContentFilterGuardrail(Guardrail):
    """Content filtering guardrail"""
   
    def __init__(self, blocked_topics, banned_words):
        self.blocked_topics = blocked_topics
        self.banned_words = banned_words
   
    def check_input(self, text):
        # Check whether it involves prohibited topics
        for topic in self.blocked_topics:
            if topic in text.lower():
                return CheckResult(
                    is_safe=False,
                    reason=f"Involves sensitive topic: {topic}"
                )
       
        # Check whether it contains banned words
        for word in self.banned_words:
            if word in text.lower():
                return CheckResult(
                    is_safe=False,
                    reason=f"Contains inappropriate words"
                )
       
        return CheckResult(is_safe=True)
   
    def check_output(self, text):
        # Output check same as above
        return self.check_input(text)


class CheckResult:
    """Check result"""
   
    def __init__(self, is_safe, reason=None):
        self.is_safe = is_safe
        self.reason = reason

Important reminder: Security is an ongoing process, and there is no foolproof solution. Security strategies need continuous monitoring, updating, and improvement.


Observability

Agents in production environments need a comprehensive monitoring system.

Observability helps us understand Agent behavior, troubleshoot issues, and optimize performance.

Three Pillars

Logging: Record all key events and decisions.

Tracing: Track the complete flow path of requests through the system.

Metrics: Collect performance and quality metrics.

Code Implementation

Observable Agent Implementation

import logging
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.resources import Resource

# Configure logging
logger = logging.getLogger(__name__)

# Configure Tracing
tracer_provider = TracerProvider()
trace.set_tracer_provider(tracer_provider)
tracer = trace.get_tracer(__name__)


class ObservableAgent:
    """
Observability Agent
Integrated log, trace, and metrics collection
    """

   
    def __init__(self, agent, metrics_collector):
        # Base Agent
        self.agent = agent
        # Metrics collector
        self.metrics = metrics_collector
   
    def process(self, user_input):
        """
Handle requests with full observability support
        """

        # Create tracing span
        with tracer.start_as_current_span("agent_process") as span:
            # Set span attributes
            span.set_attribute("user.input.length", len(user_input))
            span.set_attribute("user.input.preview", user_input[:100])
           
            # Record start
            logger.info(f"Start processing request: {user_input[:50]}...")
            start_time = time.time()
           
            try:
                # Execute core logic
                result = self.agent.process(user_input)
               
                # Record success
                duration = time.time() - start_time
                span.set_attribute("success", True)
                span.set_attribute("duration_ms", duration * 1000)
                span.set_attribute("result.length", len(str(result)))
               
                # Collect metrics
                self.metrics.record("request_duration", duration)
                self.metrics.increment("request_success")
               
                logger.info(f"Request completed, took: {duration:.2f}s")
               
                return result
               
            except Exception as e:
                # Record error
                duration = time.time() - start_time
                span.set_attribute("success", False)
                span.set_attribute("error.type", type(e).__name__)
                span.set_attribute("error.message", str(e))
                span.record_exception(e)
               
                # Collect error metrics
                self.metrics.increment("request_error")
                self.metrics.record("error_duration", duration)
               
                logger.error(f"Request failed: {e}")
               
                raise


class MetricsCollector:
    """Metrics collector"""
   
    def __init__(self):
        self.metrics = {}
   
    def increment(self, name, value=1):
        """Increment counter metric"""
        if name not in self.metrics:
            self.metrics[name] = {"type": "counter", "value": 0}
        self.metrics[name]["value"] += value
   
    def record(self, name, value):
        """Record numeric metric"""
        if name not in self.metrics:
            self.metrics[name] = {"type": "gauge", "values": []}
        self.metrics[name]["values"].append(value)
   
    def get_summary(self):
        """Get metrics summary"""
        summary = {}
        for name, data in self.metrics.items():
            if data["type"] == "counter":
                summary[name] = data["value"]
            else:
                values = data["values"]
                summary[f"{name}_avg"] = sum(values) / len(values)
                summary[f"{name}_max"] = max(values)
                summary[f"{name}_min"] = min(values)
        return summary

Guardrails and Human-in-the-Loop

Guardrails

Guardrails are real-time input/output filtering and restriction mechanisms.

Unlike safety guardrails, Guardrails focus more on ensuring output quality and compliance.

HITL(Human-in-the-Loop)

HITL introduces human review at critical decision points.

Suitable for high-risk scenarios where AI cannot make autonomous decisions.

HITL Agent Implementation

class HITLAgent:
    """
Human-in-the-loop Agent
Introduce human review at critical decision points
    """

   
    def __init__(self, agent, approval_queue, notification_handler):
        # Base Agent
        self.agent = agent
        # Approval queue
        self.approval_queue = approval_queue
        # Notification handler
        self.notification_handler = notification_handler
   
    def process(self, request):
        """
Process requests, trigger human approval when necessary
        """

        # Assess request risk level
        risk_level = self.assess_risk(request)
       
        if risk_level == "high":
            # High-risk request, requires human approval
            return self.handle_high_risk_request(request)
        elif risk_level == "medium":
            # Medium risk, enable enhanced monitoring
            return self.handle_medium_risk_request(request)
        else:
            # Low risk, process directly
            return self.agent.process(request)
   
    def assess_risk(self, request):
        """Assess request risk level"""
        # Check if sensitive operations are involved
        sensitive_operations = [
            "delete", "remove", "cancel",
            "transfer", "payment", "refund"
        ]
       
        content_lower = request.lower()
        for op in sensitive_operations:
            if op in content_lower:
                return "high"
       
        # Check complexity of request content
        if len(request) > 1000:
            return "medium"
       
        return "low"
   
    def handle_high_risk_request(self, request):
        """
Process high-risk requests that require human approval
        """

        # Create approval task
        task_id = self.approval_queue.add({
            "request": request,
            "risk_level": "high",
            "timestamp": datetime.now()
        })
       
        # Notify approver
        self.notification_handler.notify_approver(
            task_id=task_id,
            message=f"A high-risk request requires approval: {request[:100]}..."
        )
       
        # Return waiting status
        return {
            "status": "pending_approval",
            "task_id": task_id,
            "message": "Your request requires human approval, please wait."
        }
   
    def handle_feedback(self, task_id, approved, feedback=None):
        """
Process approval feedback
        """

        task = self.approval_queue.get(task_id)
       
        if approved:
            # Approval approved, execute request
            self.approval_queue.complete(task_id)
            result = self.agent.process(task["request"])
           
            # Notify applicant
            self.notification_handler.notify_requester(
                task_id=task_id,
                status="approved",
                result=result
            )
           
            return result
        else:
            # Approval rejected
            self.approval_queue.reject(task_id, feedback)
           
            # Notify applicant
            self.notification_handler.notify_requester(
                task_id=task_id,
                status="rejected",
                feedback=feedback
            )
           
            return {
                "status": "rejected",
                "reason": feedback or Approval failed.
            }


class ApprovalQueue:
    Approval queue
   
    def __init__(self):
        self.queue = {}
        self.counter = 0
   
    def add(self, task):
        Add approval task
        self.counter += 1
        task_id = f"task_{self.counter}"
        self.queue[task_id] = {
            **task,
            "status": "pending"
        }
        return task_id
   
    def get(self, task_id):
        return self.queue.get(task_id)
   
    def complete(self, task_id):
        self.queue[task_id]["status"] = "approved"
   
    def reject(self, task_id, reason):
        self.queue[task_id]["status"] = "rejected"
        self.queue[task_id]["reject_reason"] = reason

Chapter Summary

This chapter introduces knowledge about the evaluation, safety, and alignment of Agents.

Evaluation systemEvaluate the Agent's performance using multi-dimensional metrics.

security threatIncluding prompt injection, jailbreaking, data poisoning, etc.

Protection strategyEnsure system security through multi-layered security checks.

ObservabilityImplement system monitoring through logs, tracing, and metrics.

Guardrails and HITLProvide output quality control and manual review capabilities.

Evaluation, safety, and alignment are the foundation for building trustworthy agent systems.

It requires continuous attention and improvement in actual development.

Other extensions