Skip to content
AgentEnv Framework

More Useful Env Extensions

Some extensions we have found useful

Forced Errors

urn:agentenv:set-errors/v1 makes a tool fail, to test how an agent copes with a flaky API. Here every_nth fails every third call, and error_rate would fail a share of calls at random instead. An exception in a tool reaches the agent as a tool error with its message.

email/server.py
import random

from agentenv_protocol import AgentEnvEnvironment, environment_card, extension, tool

ERRORS = {"rate_limit": "rate limit exceeded, retry later", "timeout": "request timed out", "unavailable": "service unavailable"}


@environment_card(name="email")
class EmailEnv(AgentEnvEnvironment):
    def __init__(self) -> None:
        self.inbox: list[dict] = []
        self.sent: list[dict] = []
        self.faults: dict[str, dict] = {}

    @extension(uri="urn:agentenv:set-errors/v1", description="Make a tool fail.")
    async def set_errors(self, tool_name: str, every_nth: int = 0, error_rate: float = 0.0, error_type: str = "rate_limit") -> dict:
        self.faults[tool_name] = {"every_nth": every_nth, "error_rate": error_rate, "error_type": error_type, "calls": 0}
        return {"tool_name": tool_name, "every_nth": every_nth, "error_rate": error_rate, "error_type": error_type}

    def _fault(self, tool_name: str) -> None:
        fault = self.faults.get(tool_name)
        if fault is None:
            return
        fault["calls"] += 1
        nth = fault["every_nth"]
        failing = fault["calls"] % nth == 0 if nth else random.random() < fault["error_rate"]
        if failing:
            raise RuntimeError(ERRORS[fault["error_type"]])

    @tool()
    async def send(self, to: str, subject: str, body: str) -> dict:
        """Send an email."""
        self._fault("send")
        self.sent.append({"to": to, "subject": subject, "body": body})
        return {"sent": len(self.sent)}

The Acting User

urn:agentenv:set-acting-user/v1 sets who the server acts as, so one env can put the agent in different people's shoes. Here it decides whose mail email_search finds and who send sends as. Raising for a user the server doesn't know answers 500, so a task stops instead of running as the wrong person.

email/server.py (the same EmailEnv)
USERS = {"alex@example.com", "priya@example.com", "jordan@example.com"}


@environment_card(name="email")
class EmailEnv(AgentEnvEnvironment):
    def __init__(self) -> None:
        ...
        self.user_email = "alex@example.com"

    @extension(uri="urn:agentenv:set-acting-user/v1", description="Act as this user.")
    async def set_acting_user(self, user_email: str) -> dict:
        if user_email not in USERS:
            raise ValueError(f"unknown user: {user_email}")
        self.user_email = user_email
        return {"user_email": user_email}

    @tool()
    async def email_search(self, query: str) -> dict:
        """Find the acting user's emails whose subject or body contains the query."""
        q = query.lower()
        mine = [e for e in self.inbox if e["to"] == self.user_email]
        return {"emails": [e for e in mine if q in (e["subject"] + " " + e["body"]).lower()]}

    @tool()
    async def send(self, to: str, subject: str, body: str) -> dict:
        """Send an email as the acting user."""
        self._fault("send")
        self.sent.append({"from": self.user_email, "to": to, "subject": subject, "body": body})
        return {"sent": len(self.sent)}

Async Waits

urn:agentenv:set-async-wait/v1 makes a tool's work take time, so an agent has to wait and poll as it would against a real system. Here each job reports_submit starts is ready after a wait drawn between min_seconds and max_seconds. A server that follows the virtual clock can count that wait in virtual seconds instead, so a fast clock shortens it.

reports/server.py
import random
import time

from agentenv_protocol import AgentEnvEnvironment, environment_card, extension, tool


@environment_card(name="reports")
class ReportsEnv(AgentEnvEnvironment):
    def __init__(self) -> None:
        self.waits: dict[str, tuple[float, float]] = {}
        self.jobs: dict[str, dict] = {}

    @extension(uri="urn:agentenv:set-async-wait/v1", description="Make a tool's jobs take this long.")
    async def set_async_wait(self, tool_name: str, min_seconds: float, max_seconds: float | None = None) -> dict:
        self.waits[tool_name] = (min_seconds, max_seconds if max_seconds is not None else min_seconds)
        return {"tool_name": tool_name, "wait_range_seconds": list(self.waits[tool_name])}

    @tool()
    async def reports_submit(self, query: str) -> dict:
        """Start a report; poll reports_status for it."""
        wait = random.uniform(*self.waits.get("reports_submit", (0, 0)))
        job_id = f"job_{len(self.jobs) + 1}"
        self.jobs[job_id] = {"query": query, "ready_at": time.monotonic() + wait}
        return {"job_id": job_id, "status": "pending"}

    @tool()
    async def reports_status(self, job_id: str) -> dict:
        """Whether a report is ready, and its rows once it is."""
        job = self.jobs[job_id]
        if time.monotonic() < job["ready_at"]:
            return {"job_id": job_id, "status": "pending"}
        return {"job_id": job_id, "status": "ready", "rows": 12}

In a Task

In a task, the apply_server_config step calls these extensions before the agent runs, as Env control describes.

Last updated on

Ask a question · Report an issue

On this page