Agent Evaluation, Safety, and Alignment
Evaluation is the foundation for ensuring Agent effectiveness.
Safety and alignment are key to whether an Agent can be trusted and deployed.
Agent Evaluation System
Evaluating Agent performance is a critical part of the development process.
A good evaluation system helps us understand the boundaries of an Agent's capabilities.
It also provides direction and basis for continuous optimization.
Evaluation Dimensions
Task Completion: Whether the Agent correctly completed the given task.
Efficiency Metrics: Number of steps, token consumption, and execution time required to complete the task.
Quality Metrics: Accuracy of answers, consistency of responses, and naturalness of dialogue.
Robustness: Ability to handle abnormal inputs and noisy data.
Common Benchmarks
| Benchmark | Purpose | Evaluation Content |
|---|---|---|
| GAIA | General AI Assistant Evaluation | Complex task handling, multi-step reasoning |
| MMLU | Multi-task Language Understanding | Knowledge Q&A across 57 subjects |
| HumanEval | Code Generation Evaluation | Correctness of Python code writing |
| HotpotQA | Multi-hop Question Answering Evaluation | Questions requiring reasoning across multiple documents |
| AgentBench | Agent Capability Evaluation | Agent performance in real-world environments |
Code Implementation: Evaluation Framework
Agent Evaluation Framework
"""
Agent Evaluation Framework
Evaluate the performance of an Agent on various tasks
"""
def __init__(self, agent, metrics):
# Agent to be evaluated
self.agent = agent
# Evaluation Metrics List
self.metrics = metrics
def evaluate(self, test_cases):
"""
Execute evaluation
:param test_cases: list of test cases
:return: evaluation report
"""
results = []
for test_case in test_cases:
# Execute task
result = self.run_single_test(test_case)
results.append(result)
# Generate evaluation report
report = self.generate_report(results)
return report
def run_single_test(self, test_case):
"""
Run a single test case
"""
# Record start time
start_time = time.time()
# Execute Agent
try:
output = self.agent.run(test_case.input)
success = self.evaluate_output(output, test_case.expected)
error = None
except Exception as e:
output = None
success = False
error = str(e)
# Record end time
end_time = time.time()
return TestResult(
test_case=test_case,
output=output,
success=success,
error=error,
duration=end_time - start_time,
token_count=self.count_tokens(output)
)
def evaluate_output(self, output, expected):
"""Evaluate whether the output meets expectations"""
for metric in self.metrics:
if not metric.evaluate(output, expected):
return False
return True
def generate_report(self, results):
"""Generate evaluation report"""
total = len(results)
passed = sum(1 for r in results if r.success)
# Calculate various metrics
avg_duration = sum(r.duration for r in results) / total
avg_tokens = sum(r.token_count for r in results) / total
# Group and count by test type
by_category = {}
for r in results:
category = r.test_case.category
if category not in by_category:
by_category[category] = {"total": 0, "passed": 0}
by_category[category]["total"] += 1
if r.success:
by_category[category]["passed"] += 1
return EvaluationReport(
total=total,
passed=passed,
pass_rate=passed / total,
avg_duration=avg_duration,
avg_tokens=avg_tokens,
by_category=by_category,
results=results
)
class TestCase:
"""Test case"""
def __init__(self, input, expected, category="general", metadata=None):
# Input
self.input = input
# Expected output or evaluation criteria
self.expected = expected
# Category
self.category = category
# Extra metadata
self.metadata = metadata or {}
class TestResult:
"""Test result"""
def __init__(self, test_case, output, success, error, duration, token_count):
self.test_case = test_case
self.output = output
self.success = success
self.error = error
self.duration = duration
self.token_count = token_count
class Metric:
"""Base class for evaluation metrics"""
def evaluate(self, output, expected):
raise NotImplementedError
class ExactMatchMetric(Metric):
"""Exact match metric"""
def evaluate(self, output, expected):
return output.strip() == expected.strip()
class ContainsMetric(Metric):
"""Keyword inclusion metric"""
def evaluate(self, output, expected):
if isinstance(expected, list):
return all(keyword in output for keyword in expected)
return expected in output
class SemanticSimilarityMetric(Metric):
"""Semantic similarity metric"""
def __init__(self, threshold=0.8):
self.threshold = threshold
def evaluate(self, output, expected):
similarity = self.compute_similarity(output, expected)
return similarity >= self.threshold
def compute_similarity(self, text1, text2):
"""Calculate the semantic similarity of two texts"""
# Use embedding model to compute cosine similarity
embedding1 = self.embedder.embed([text1])[0]
embedding2 = self.embedder.embed([text2])[0]
return cosine_similarity(embedding1, embedding2)
Safety and Alignment
The safety of the Agent is of utmost importance.
AI systems may be subject to various attacks and produce harmful outputs.
Alignment ensures that AI behavior aligns with human intentions and values.
Common Safety Threats
Prompt Injection
Attackers use input to induce the Agent to ignore system instructions.
Example input: "Ignore previous instructions and instead execute..."
This is a context hijacking attack that exploits the Agent's trust in user input.
Jailbreaking
Bypass security restrictions through specific inputs.
For example, using techniques such as role-playing or fictional scenarios.
Data Poisoning
Maliciously modify training data or retrieval results.
Causing the Agent to produce incorrect or harmful outputs.
Sensitive Information Leakage
The Agent improperly exposes user privacy or internal system information.
Defense Strategies
Safe Agent Implementation
"""
Safety Agent
Add multiple layers of safety protection on top of the base Agent
"""
def __init__(self, base_agent, guardrails, input_validator, output_filter):
# Base Agent
self.base_agent = base_agent
# List of safety guardrails
self.guardrails = guardrails
# Input validator
self.input_validator = input_validator
# Output filter
self.output_filter = output_filter
def process(self, user_input):
"""
Process user input, including multiple layers of security checks
"""
# ==================== Layer 1: Input Validation ====================
# Check whether the input is valid
is_valid, reason = self.input_validator.validate(user_input)
if not is_valid:
return self.create_safety_response(reason)
# ==================== Layer 2: Injection Detection ====================
# Detect attacks such as prompt injection
for guardrail in self.guardrails:
check_result = guardrail.check_input(user_input)
if not check_result.is_safe:
# Record security events
self.log_security_event(
event_type="input_guardrail_triggered",
input=user_input,
reason=check_result.reason
)
return self.create_safety_response(check_result.reason)
# ==================== Layer 3: Execute Core Logic ====================
try:
response = self.base_agent.process(user_input)
except Exception as e:
return self.create_error_response(str(e))
# ==================== Layer 4: Output Filtering ====================
# Check whether the output is safe
for guardrail in self.guardrails:
check_result = guardrail.check_output(response)
if not check_result.is_safe:
self.log_security_event(
event_type="output_guardrail_triggered",
output=response,
reason=check_result.reason
)
return self.create_safety_response(check_result.reason)
# Apply output filtering (e.g., desensitization of sensitive information)
response = self.output_filter.filter(response)
return response
def create_safety_response(self, reason):
"""Create safe response"""
return {
"type": "safety_block",
"message": "Sorry, I cannot complete this request.",
"reason": reason
}
def log_security_event(self, event_type, **kwargs):
"""Record security events"""
# In actual applications, this should be written to a security log system
print(f"[SECURITY] {event_type}: {kwargs}")
class InputValidator:
"""Input validator"""
def validate(self, text):
"""
Validate whether the input is valid
:return: (is_valid, reason)
"""
if not text or len(text.strip()) == 0:
return False, "Input cannot be empty"
if len(text) > 10000:
return False, "Input length exceeds the limit"
# Check whether it contains executable content
if self.contains_executable_content(text):
return False, "Input contains suspicious executable content"
return True, None
def contains_executable_content(self, text):
"""Check whether it contains executable content"""
# Simplified implementation
suspicious_patterns = [
"javascript:",
"data:text/html",
"<script>",
]
return any(pattern in text.lower() for pattern in suspicious_patterns)
class Guardrail:
"""Safety guardrail"""
def check_input(self, text):
"""Check input"""
raise NotImplementedError
def check_output(self, text):
"""Check output"""
raise NotImplementedError
class ContentFilterGuardrail(Guardrail):
"""Content filtering guardrail"""
def __init__(self, blocked_topics, banned_words):
self.blocked_topics = blocked_topics
self.banned_words = banned_words
def check_input(self, text):
# Check whether it involves prohibited topics
for topic in self.blocked_topics:
if topic in text.lower():
return CheckResult(
is_safe=False,
reason=f"Involves sensitive topic: {topic}"
)
# Check whether it contains banned words
for word in self.banned_words:
if word in text.lower():
return CheckResult(
is_safe=False,
reason=f"Contains inappropriate words"
)
return CheckResult(is_safe=True)
def check_output(self, text):
# Output check same as above
return self.check_input(text)
class CheckResult:
"""Check result"""
def __init__(self, is_safe, reason=None):
self.is_safe = is_safe
self.reason = reason
Important reminder: Security is an ongoing process, and there is no foolproof solution. Security strategies need continuous monitoring, updating, and improvement.
Observability
Agents in production environments need a comprehensive monitoring system.
Observability helps us understand Agent behavior, troubleshoot issues, and optimize performance.
Three Pillars
Logging: Record all key events and decisions.
Tracing: Track the complete flow path of requests through the system.
Metrics: Collect performance and quality metrics.
Code Implementation
Observable Agent Implementation
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.resources import Resource
# Configure logging
logger = logging.getLogger(__name__)
# Configure Tracing
tracer_provider = TracerProvider()
trace.set_tracer_provider(tracer_provider)
tracer = trace.get_tracer(__name__)
class ObservableAgent:
"""
Observability Agent
Integrated log, trace, and metrics collection
"""
def __init__(self, agent, metrics_collector):
# Base Agent
self.agent = agent
# Metrics collector
self.metrics = metrics_collector
def process(self, user_input):
"""
Handle requests with full observability support
"""
# Create tracing span
with tracer.start_as_current_span("agent_process") as span:
# Set span attributes
span.set_attribute("user.input.length", len(user_input))
span.set_attribute("user.input.preview", user_input[:100])
# Record start
logger.info(f"Start processing request: {user_input[:50]}...")
start_time = time.time()
try:
# Execute core logic
result = self.agent.process(user_input)
# Record success
duration = time.time() - start_time
span.set_attribute("success", True)
span.set_attribute("duration_ms", duration * 1000)
span.set_attribute("result.length", len(str(result)))
# Collect metrics
self.metrics.record("request_duration", duration)
self.metrics.increment("request_success")
logger.info(f"Request completed, took: {duration:.2f}s")
return result
except Exception as e:
# Record error
duration = time.time() - start_time
span.set_attribute("success", False)
span.set_attribute("error.type", type(e).__name__)
span.set_attribute("error.message", str(e))
span.record_exception(e)
# Collect error metrics
self.metrics.increment("request_error")
self.metrics.record("error_duration", duration)
logger.error(f"Request failed: {e}")
raise
class MetricsCollector:
"""Metrics collector"""
def __init__(self):
self.metrics = {}
def increment(self, name, value=1):
"""Increment counter metric"""
if name not in self.metrics:
self.metrics[name] = {"type": "counter", "value": 0}
self.metrics[name]["value"] += value
def record(self, name, value):
"""Record numeric metric"""
if name not in self.metrics:
self.metrics[name] = {"type": "gauge", "values": []}
self.metrics[name]["values"].append(value)
def get_summary(self):
"""Get metrics summary"""
summary = {}
for name, data in self.metrics.items():
if data["type"] == "counter":
summary[name] = data["value"]
else:
values = data["values"]
summary[f"{name}_avg"] = sum(values) / len(values)
summary[f"{name}_max"] = max(values)
summary[f"{name}_min"] = min(values)
return summary
Guardrails and Human-in-the-Loop
Guardrails
Guardrails are real-time input/output filtering and restriction mechanisms.
Unlike safety guardrails, Guardrails focus more on ensuring output quality and compliance.
HITL(Human-in-the-Loop)
HITL introduces human review at critical decision points.
Suitable for high-risk scenarios where AI cannot make autonomous decisions.
HITL Agent Implementation
"""
Human-in-the-loop Agent
Introduce human review at critical decision points
"""
def __init__(self, agent, approval_queue, notification_handler):
# Base Agent
self.agent = agent
# Approval queue
self.approval_queue = approval_queue
# Notification handler
self.notification_handler = notification_handler
def process(self, request):
"""
Process requests, trigger human approval when necessary
"""
# Assess request risk level
risk_level = self.assess_risk(request)
if risk_level == "high":
# High-risk request, requires human approval
return self.handle_high_risk_request(request)
elif risk_level == "medium":
# Medium risk, enable enhanced monitoring
return self.handle_medium_risk_request(request)
else:
# Low risk, process directly
return self.agent.process(request)
def assess_risk(self, request):
"""Assess request risk level"""
# Check if sensitive operations are involved
sensitive_operations = [
"delete", "remove", "cancel",
"transfer", "payment", "refund"
]
content_lower = request.lower()
for op in sensitive_operations:
if op in content_lower:
return "high"
# Check complexity of request content
if len(request) > 1000:
return "medium"
return "low"
def handle_high_risk_request(self, request):
"""
Process high-risk requests that require human approval
"""
# Create approval task
task_id = self.approval_queue.add({
"request": request,
"risk_level": "high",
"timestamp": datetime.now()
})
# Notify approver
self.notification_handler.notify_approver(
task_id=task_id,
message=f"A high-risk request requires approval: {request[:100]}..."
)
# Return waiting status
return {
"status": "pending_approval",
"task_id": task_id,
"message": "Your request requires human approval, please wait."
}
def handle_feedback(self, task_id, approved, feedback=None):
"""
Process approval feedback
"""
task = self.approval_queue.get(task_id)
if approved:
# Approval approved, execute request
self.approval_queue.complete(task_id)
result = self.agent.process(task["request"])
# Notify applicant
self.notification_handler.notify_requester(
task_id=task_id,
status="approved",
result=result
)
return result
else:
# Approval rejected
self.approval_queue.reject(task_id, feedback)
# Notify applicant
self.notification_handler.notify_requester(
task_id=task_id,
status="rejected",
feedback=feedback
)
return {
"status": "rejected",
"reason": feedback or Approval failed.
}
class ApprovalQueue:
Approval queue
def __init__(self):
self.queue = {}
self.counter = 0
def add(self, task):
Add approval task
self.counter += 1
task_id = f"task_{self.counter}"
self.queue[task_id] = {
**task,
"status": "pending"
}
return task_id
def get(self, task_id):
return self.queue.get(task_id)
def complete(self, task_id):
self.queue[task_id]["status"] = "approved"
def reject(self, task_id, reason):
self.queue[task_id]["status"] = "rejected"
self.queue[task_id]["reject_reason"] = reason
Chapter Summary
This chapter introduces knowledge about the evaluation, safety, and alignment of Agents.
Evaluation systemEvaluate the Agent's performance using multi-dimensional metrics.
security threatIncluding prompt injection, jailbreaking, data poisoning, etc.
Protection strategyEnsure system security through multi-layered security checks.
ObservabilityImplement system monitoring through logs, tracing, and metrics.
Guardrails and HITLProvide output quality control and manual review capabilities.
Evaluation, safety, and alignment are the foundation for building trustworthy agent systems.
It requires continuous attention and improvement in actual development.
Other extensions