Skip to content
AgentEnv Framework

Single-Turn Prompt in an RL Env Graded by an Agent Judge

An accounts-payable agent works a week's invoice queue against policy in one prompt, and a judge agent grades every decision it made, from an empty folder to a score

An agent gets one prompt, Work this week's AP queue, inside Acme Corp's accounts-payable system. The system holds vendors, purchase orders, goods receipts, six pending invoices and a written AP policy. Nothing in the prompt says which invoices to pay: the agent has to read the policy, check every invoice against its purchase order, its receipts and its vendor, and approve, hold or route each one with a note. Each invoice tests a different rule, the way a real queue does:

InvoiceWhat the agent has to noticeRight action
INV-1001Its price is 1.2% over the PO, inside the 2% toleranceApprove
INV-1002Its price is 3.8% over the POHold, price_variance
INV-1003It bills 500 units, and 350 were receivedHold, quantity_variance
INV-1004Same vendor, total and date as INV-1001, under a different numberHold, duplicate
INV-1005It matches, for 18,600.00, over the 10,000 approval limitRoute to the controller
INV-1006It matches, but the vendor changed its bank account six days ago by emailHold, vendor_verification

A judge agent then reads the agent's trajectory, every tool call it made, and grades each decision against a rubric, with partial credit and a penalty for paying what should have been held. The score is the reward. The task has five steps:

StepTypeWhat it does
apdeploy_envStarts an instance of the AP system
seedload_artifactLoads the acme-finance universe: the policy, the ledger and the queue
clerkdeploy_agentStarts the agent as an AP specialist and hands it the system's tools
workprompt_agentSends it the one prompt and stores its reply and trajectory
graderubrics_verifierHas the judge read the trajectory and score each decision

What you need

  • Python 3.11 or newer, and Docker running.
  • A model endpoint that serves the Anthropic Messages API, such as a LiteLLM proxy, with a key. Both agents here call claude-opus-5-5.

Install the framework, which gives you the agent-env command:

Terminal
pip install agentenv-framework

Then make a folder for the cookbook. Every command on this page runs from it:

The folder
cookbook/
├── .agentenv/config.toml
├── ap/server.py
├── ap/Dockerfile
├── acme/ap.json
├── operator/agent.py
├── operator/Dockerfile
├── judge/agent.py
├── judge/Dockerfile
└── task.json

Point it at your model

The framework hands every agent it deploys a model endpoint and key, as LITELLM_BASE_URL and LITELLM_API_KEY. They come from .agentenv/config.toml, with the key as a reference to your secret store rather than the key itself:

.agentenv/config.toml
[model]
base_url = "https://llm.example.com/v1"
api_key = "secret:litellm_api_key"

The default secret store reads secret:litellm_api_key from an environment variable of that name:

Terminal
export litellm_api_key=sk-...

The agents run in containers and get base_url as it is. An endpoint on your own machine is http://host.docker.internal:<port>/v1 to them, not localhost. Deploying your agent covers the rest.

Register the two built-in envs

Every env instance runs behind a gateway and next to a state store, and the framework deploys both from envs of its own. Register them once per registry, or start the local explorer with agent-env up, which registers the same two on its first start:

Terminal
agent-env env service-db put --id default-db
agent-env env gateway put --id default
Output (end of each)
Created ServiceDBEnv: id=default-db version=1
Created GatewayEnv: id=default version=1

The first run builds their images, which takes a few minutes. Images go to the local registry at localhost:5000, which the framework starts in Docker on the first push. On an Apple Silicon Mac, add --platform linux/arm64 to these and to every other put that builds an image on this page, so the containers run natively.

Write the environment

The environment is an accounts-payable system, written as an AgentEnvEnvironment like any other. Its read tools show the agent the ledger; its three action tools change it:

ap/server.py
import json
from pathlib import Path

from agentenv_protocol import (
    AgentEnvEnvironment,
    DataPart,
    FilePart,
    add_data,
    environment_card,
    get_data,
    reset_data,
    tool,
)

HOLD_REASONS = {"price_variance", "quantity_variance", "missing_receipt", "duplicate", "vendor_verification"}


@environment_card(name="ap")
class AccountsPayableEnv(AgentEnvEnvironment):
    """An accounts-payable system: vendors, purchase orders, goods receipts and an invoice queue."""

    def __init__(self) -> None:
        self._clear()

    def _clear(self) -> None:
        self.policy: dict = {}
        self.vendors: dict[str, dict] = {}
        self.purchase_orders: dict[str, dict] = {}
        self.receipts: list[dict] = []
        self.invoices: dict[str, dict] = {}
        self.audit_log: list[dict] = []

    @reset_data
    async def _reset(self) -> None:
        self._clear()

    @add_data
    async def _add(self, parts: list) -> None:
        for part in parts:
            if isinstance(part, DataPart):
                data = part.data
            elif isinstance(part, FilePart):
                data = json.loads(Path(part.file.uri.removeprefix("file://")).read_text())
            else:
                continue
            self.policy.update(data.get("policy", {}))
            self.vendors.update({v["vendor_id"]: v for v in data.get("vendors", [])})
            self.purchase_orders.update({po["po_number"]: po for po in data.get("purchase_orders", [])})
            self.receipts.extend(data.get("receipts", []))
            self.invoices.update({i["invoice_id"]: {"status": "pending", **i} for i in data.get("invoices", [])})

    @get_data
    async def _state(self) -> list:
        statuses = {i: {"status": inv["status"], "note": inv.get("note")} for i, inv in self.invoices.items()}
        return [DataPart(data={"invoices": statuses, "audit_log": self.audit_log})]

    @tool(name="{environment_name}_get_policy")
    async def get_policy(self) -> dict:
        """Read the accounts-payable policy: tolerances, approval limits and the rules for each action."""
        return self.policy

    @tool(name="{environment_name}_list_invoices")
    async def list_invoices(self, status: str = "pending") -> dict:
        """List invoices with the given status: pending, approved, on_hold or routed."""
        rows = [
            {k: inv[k] for k in ("invoice_id", "vendor_id", "invoice_number", "po_number", "total", "invoice_date", "received_at")}
            for inv in self.invoices.values()
            if inv["status"] == status
        ]
        return {"invoices": rows}

    @tool(name="{environment_name}_get_invoice")
    async def get_invoice(self, invoice_id: str) -> dict:
        """Read one invoice with its lines, as the vendor billed it."""
        return self.invoices.get(invoice_id) or {"error": f"No invoice {invoice_id}"}

    @tool(name="{environment_name}_get_purchase_order")
    async def get_purchase_order(self, po_number: str) -> dict:
        """Read a purchase order: what Acme agreed to buy, and at what unit price."""
        return self.purchase_orders.get(po_number) or {"error": f"No purchase order {po_number}"}

    @tool(name="{environment_name}_get_receipts")
    async def get_receipts(self, po_number: str) -> dict:
        """List the goods receipts against a purchase order: what the warehouse actually received."""
        return {"receipts": [r for r in self.receipts if r["po_number"] == po_number]}

    @tool(name="{environment_name}_get_vendor")
    async def get_vendor(self, vendor_id: str) -> dict:
        """Read a vendor's master record, including its bank account and recent changes to it."""
        return self.vendors.get(vendor_id) or {"error": f"No vendor {vendor_id}"}

    def _act(self, invoice_id: str, status: str, note: str, **fields) -> dict:
        invoice = self.invoices.get(invoice_id)
        if invoice is None:
            return {"error": f"No invoice {invoice_id}"}
        if invoice["status"] != "pending":
            return {"error": f"{invoice_id} is already {invoice['status']}"}
        if not note.strip():
            return {"error": "Every action needs a note"}
        invoice.update(status=status, note=note, **fields)
        self.audit_log.append({"invoice_id": invoice_id, "status": status, "note": note, **fields})
        return {"invoice_id": invoice_id, "status": status}

    @tool(name="{environment_name}_approve_invoice")
    async def approve_invoice(self, invoice_id: str, note: str) -> dict:
        """Approve an invoice for payment. Only for invoices within your approval limit."""
        invoice = self.invoices.get(invoice_id)
        if invoice and invoice["total"] > self.policy.get("approval_limit", 0):
            return {"error": f"{invoice_id} is over your approval limit of {self.policy['approval_limit']}"}
        return self._act(invoice_id, "approved", note)

    @tool(name="{environment_name}_hold_invoice")
    async def hold_invoice(self, invoice_id: str, reason: str, note: str) -> dict:
        """Put an invoice on hold. reason is one of price_variance, quantity_variance, missing_receipt, duplicate, vendor_verification."""
        if reason not in HOLD_REASONS:
            return {"error": f"reason must be one of {sorted(HOLD_REASONS)}"}
        return self._act(invoice_id, "on_hold", note, reason=reason)

    @tool(name="{environment_name}_route_invoice")
    async def route_invoice(self, invoice_id: str, approver: str, note: str) -> dict:
        """Route an invoice to another approver, such as one over your approval limit."""
        return self._act(invoice_id, "routed", note, approver=approver)


if __name__ == "__main__":
    AccountsPayableEnv().serve()
ap/Dockerfile
FROM python:3.12-slim
WORKDIR /app
RUN pip install --no-cache-dir agentenv-framework-protocol "mcp>=1.25,<2"
COPY server.py .
CMD ["python", "server.py"]
  • The tools are what the agent sees, each named after the env: ap_get_policy, ap_list_invoices, ap_get_invoice, ap_get_purchase_order, ap_get_receipts, ap_get_vendor, ap_approve_invoice, ap_hold_invoice and ap_route_invoice.
  • The system enforces what a real one would: ap_approve_invoice refuses an invoice over the approval limit, every action needs a note, and nothing is processed twice. The agent has to read the errors and recover, as it would at work.
  • The data plane seeds the ledger: data/reset empties it and data/add loads a file. data/get returns every invoice's status and note, and the audit log.

Write the data

The ledger is one file: the policy, five vendors, their purchase orders and goods receipts, and the six pending invoices. Nothing marks which invoices are wrong; the agent finds out by comparing:

acme/ap.json
{
  "policy": {
    "today": "2026-09-30",
    "price_tolerance_pct": 2.0,
    "approval_limit": 10000,
    "over_limit_approver": "controller@acme.example",
    "rules": [
      "Approve an invoice only when it three-way matches: every invoiced quantity has been received against its purchase order, and every unit price is within price_tolerance_pct of the purchase order's price.",
      "Hold an invoice whose unit price is outside tolerance with reason price_variance, and one billing more than was received with reason quantity_variance.",
      "Hold with reason duplicate an invoice whose vendor, total and invoice date match another invoice that arrived earlier. Process the earlier one on its own merits.",
      "Hold with reason vendor_verification every invoice from a vendor whose bank account changed in the last 30 days, until the change is confirmed by phone, even if the invoice matches.",
      "Route a matching invoice above approval_limit to over_limit_approver. Never split an invoice to fit a limit.",
      "Every action needs a note that states the reason with the numbers a colleague needs to act on it."
    ]
  },
  "vendors": [
    {"vendor_id": "V-100", "name": "Northwind Traders", "bank_account": "****4412", "bank_account_changed_at": "2024-03-02"},
    {"vendor_id": "V-200", "name": "Contoso Supplies", "bank_account": "****0917", "bank_account_changed_at": "2023-11-15"},
    {"vendor_id": "V-300", "name": "Fabrikam Industrial", "bank_account": "****7730", "bank_account_changed_at": "2025-06-20"},
    {"vendor_id": "V-400", "name": "Litware Software", "bank_account": "****2208", "bank_account_changed_at": "2022-01-09"},
    {"vendor_id": "V-500", "name": "Tailspin Toys", "bank_account": "****9051", "bank_account_changed_at": "2026-09-24",
     "bank_account_change_source": "email from accounts@tailspin-toys.co requesting new remittance details"}
  ],
  "purchase_orders": [
    {"po_number": "PO-7001", "vendor_id": "V-100", "lines": [{"sku": "NW-CABLE-6FT", "quantity": 300, "unit_price": 14.00}]},
    {"po_number": "PO-7002", "vendor_id": "V-200", "lines": [{"sku": "CT-TONER-XL", "quantity": 120, "unit_price": 25.00}]},
    {"po_number": "PO-7003", "vendor_id": "V-300", "lines": [{"sku": "FB-BRACKET-L", "quantity": 500, "unit_price": 6.40}]},
    {"po_number": "PO-7004", "vendor_id": "V-400", "lines": [{"sku": "LW-SUITE-ANNUAL", "quantity": 1, "unit_price": 18600.00}]},
    {"po_number": "PO-7005", "vendor_id": "V-500", "lines": [{"sku": "TT-DEMO-KIT", "quantity": 40, "unit_price": 95.00}]}
  ],
  "receipts": [
    {"receipt_id": "GR-5101", "po_number": "PO-7001", "sku": "NW-CABLE-6FT", "quantity": 300, "received_on": "2026-09-19"},
    {"receipt_id": "GR-5102", "po_number": "PO-7002", "sku": "CT-TONER-XL", "quantity": 120, "received_on": "2026-09-21"},
    {"receipt_id": "GR-5103", "po_number": "PO-7003", "sku": "FB-BRACKET-L", "quantity": 350, "received_on": "2026-09-23"},
    {"receipt_id": "GR-5104", "po_number": "PO-7004", "sku": "LW-SUITE-ANNUAL", "quantity": 1, "received_on": "2026-09-01"},
    {"receipt_id": "GR-5105", "po_number": "PO-7005", "sku": "TT-DEMO-KIT", "quantity": 40, "received_on": "2026-09-25"}
  ],
  "invoices": [
    {"invoice_id": "INV-1001", "vendor_id": "V-100", "invoice_number": "NW-5512", "po_number": "PO-7001", "invoice_date": "2026-09-22", "received_at": "2026-09-23T09:14:00Z",
     "lines": [{"sku": "NW-CABLE-6FT", "quantity": 300, "unit_price": 14.17}], "total": 4251.00},
    {"invoice_id": "INV-1002", "vendor_id": "V-200", "invoice_number": "CT-0918", "po_number": "PO-7002", "invoice_date": "2026-09-24", "received_at": "2026-09-24T15:02:00Z",
     "lines": [{"sku": "CT-TONER-XL", "quantity": 120, "unit_price": 25.95}], "total": 3114.00},
    {"invoice_id": "INV-1003", "vendor_id": "V-300", "invoice_number": "FB-2231", "po_number": "PO-7003", "invoice_date": "2026-09-25", "received_at": "2026-09-26T08:40:00Z",
     "lines": [{"sku": "FB-BRACKET-L", "quantity": 500, "unit_price": 6.40}], "total": 3200.00},
    {"invoice_id": "INV-1004", "vendor_id": "V-100", "invoice_number": "NW5512", "po_number": "PO-7001", "invoice_date": "2026-09-22", "received_at": "2026-09-27T11:31:00Z",
     "lines": [{"sku": "NW-CABLE-6FT", "quantity": 300, "unit_price": 14.17}], "total": 4251.00},
    {"invoice_id": "INV-1005", "vendor_id": "V-400", "invoice_number": "LW-0077", "po_number": "PO-7004", "invoice_date": "2026-09-26", "received_at": "2026-09-28T10:05:00Z",
     "lines": [{"sku": "LW-SUITE-ANNUAL", "quantity": 1, "unit_price": 18600.00}], "total": 18600.00},
    {"invoice_id": "INV-1006", "vendor_id": "V-500", "invoice_number": "TT-4410", "po_number": "PO-7005", "invoice_date": "2026-09-27", "received_at": "2026-09-29T13:20:00Z",
     "lines": [{"sku": "TT-DEMO-KIT", "quantity": 40, "unit_price": 95.00}], "total": 3800.00}
  ]
}

Write the agent

The agent is Claude with the tools of every env it is handed. It is general on purpose: the task gives it its role, through deploy_agent's system_prompt, so the same agent works the AP queue here and the support desk in Multi-Turn Conversation Between Two Agents:

operator/agent.py
import os
from contextlib import AsyncExitStack

from anthropic import AsyncAnthropic
from mcp import ClientSession
from mcp.client.streamable_http import streamablehttp_client

from agentenv_protocol.a2a_agent import (
    MCP_CONFIG_V1,
    TRAJECTORY_V1,
    AgentConfig,
    AgentEnvAgent,
    AgentIdentity,
    TaskRequest,
    TaskResult,
    TextPart,
    Usage,
    a2a_agent,
)

# Every A2A context is one conversation: its messages so far, so a later turn continues where the last one ended.
CONVERSATIONS: dict[str, list] = {}


class OperatorConfig(AgentConfig):
    model: str = "claude-opus-5-5"
    system_prompt: str = "You work in the systems you are given. Use their tools, follow their policy, and say what you did."
    max_turns: int = 30


@a2a_agent(
    identity=AgentIdentity(
        name="operator",
        description="Claude that works an enterprise system through the tools of the envs it is given, across the turns of a conversation.",
        version="1.0.0",
    ),
    config=OperatorConfig,
    extensions=(MCP_CONFIG_V1, TRAJECTORY_V1),
)
class OperatorAgent(AgentEnvAgent):
    async def run(self, request: TaskRequest[OperatorConfig]) -> TaskResult:
        config = request.config
        claude = AsyncAnthropic(
            base_url=os.environ["LITELLM_BASE_URL"].removesuffix("/v1"), api_key=os.environ["LITELLM_API_KEY"]
        )
        message = "\n".join(part.text for part in request.parts if isinstance(part, TextPart))
        messages = CONVERSATIONS.setdefault(request.context_id, [])
        messages.append({"role": "user", "content": message})
        # This turn's trajectory: the message, each tool call with its result, and the answer.
        turn = [{"type": "message", "text": message}]

        async with AsyncExitStack() as stack:
            sessions, tools = {}, []
            for server in request.mcp_servers.values():
                read, write, _ = await stack.enter_async_context(
                    streamablehttp_client(server["url"], headers=server.get("headers"))
                )
                session = await stack.enter_async_context(ClientSession(read, write))
                await session.initialize()
                for tool in (await session.list_tools()).tools:
                    sessions[tool.name] = session
                    tools.append({"name": tool.name, "description": tool.description or "", "input_schema": tool.inputSchema})

            for _ in range(config.max_turns):
                response = await claude.messages.create(
                    model=config.model,
                    max_tokens=16000,
                    system=config.system_prompt,
                    tools=tools,
                    messages=messages,
                )
                messages.append({"role": "assistant", "content": response.content})
                if response.stop_reason == "refusal":
                    return TaskResult.failure("refusal", "Claude declined the task")
                if response.stop_reason != "tool_use":
                    answer = "".join(block.text for block in response.content if block.type == "text")
                    turn.append({"type": "answer", "text": answer})
                    calls = sum(step["type"] == "tool_call" for step in turn)
                    return (
                        TaskResult.builder()
                        .succeeded()
                        .add_text(answer)
                        .usage(Usage(tool_call_count=calls))
                        .native_trajectory(format="operator/v1", payload=turn)
                        .build()
                    )
                results = []
                for block in response.content:
                    if block.type != "tool_use":
                        continue
                    result = await sessions[block.name].call_tool(block.name, block.input)
                    output = "\n".join(c.text for c in result.content if c.type == "text")
                    turn.append({"type": "tool_call", "tool": block.name, "input": block.input, "output": output})
                    results.append({"type": "tool_result", "tool_use_id": block.id, "content": output})
                messages.append({"role": "user", "content": results})

        return TaskResult.failure("max_turns", f"Still calling tools after {config.max_turns} model calls")


if __name__ == "__main__":
    OperatorAgent().serve()
operator/Dockerfile
FROM python:3.12-slim
WORKDIR /app
RUN pip install --no-cache-dir "agentenv-framework-protocol[agent]" "mcp>=1.25,<2" anthropic
COPY agent.py .
CMD ["python", "agent.py"]
  • MCP_CONFIG_V1 is how it receives the env: deploy_agent posts the instance's MCP URL to it, and each prompt then carries it in request.mcp_servers.
  • TRAJECTORY_V1 is how the framework reads back turn, the prompt, every tool call with its result, and the answer. That is what the judge checks.
  • CONVERSATIONS keeps each conversation's messages by context_id. A single-turn task sends one message; a multi-turn one sends several in the same context, and the agent picks up where it left off.

Write the judge

The judge is an agent too, so it can check what the agent did rather than what it says it did. The framework copies the trajectory into the judge's container at /tmp/prompt_trajectory.json and says so in the grading prompt, so this judge has one tool, read_file, and answers with only the JSON the prompt asks for:

judge/agent.py
import os
from pathlib import Path

from anthropic import AsyncAnthropic

from agentenv_protocol.a2a_agent import (
    TRAJECTORY_V1,
    AgentConfig,
    AgentEnvAgent,
    AgentIdentity,
    TaskRequest,
    TaskResult,
    TextPart,
    a2a_agent,
)

READ_FILE = {
    "name": "read_file",
    "description": "Read a file the grading prompt points you at, such as the agent's trajectory, or list a directory of them.",
    "input_schema": {
        "type": "object",
        "properties": {"path": {"type": "string", "description": "The file's or directory's absolute path."}},
        "required": ["path"],
    },
}


class JudgeConfig(AgentConfig):
    model: str = "claude-opus-5-5"
    system_prompt: str = (
        "You grade another agent's work against a rubric. Read the files the prompt points you at "
        "before you decide, and answer with only the JSON the prompt asks for."
    )
    max_turns: int = 10


@a2a_agent(
    identity=AgentIdentity(
        name="judge",
        description="Claude that grades an agent's reply and trajectory against a rubric.",
        version="1.0.0",
    ),
    config=JudgeConfig,
    extensions=(TRAJECTORY_V1,),
)
class JudgeAgent(AgentEnvAgent):
    async def run(self, request: TaskRequest[JudgeConfig]) -> TaskResult:
        config = request.config
        claude = AsyncAnthropic(
            base_url=os.environ["LITELLM_BASE_URL"].removesuffix("/v1"), api_key=os.environ["LITELLM_API_KEY"]
        )
        prompt = "\n".join(part.text for part in request.parts if isinstance(part, TextPart))
        messages = [{"role": "user", "content": prompt}]
        reads = []
        for _ in range(config.max_turns):
            response = await claude.messages.create(
                model=config.model,
                max_tokens=16000,
                system=config.system_prompt,
                tools=[READ_FILE],
                messages=messages,
            )
            if response.stop_reason != "tool_use":
                verdict = "".join(block.text for block in response.content if block.type == "text")
                return (
                    TaskResult.builder()
                    .succeeded()
                    .add_text(verdict)
                    .native_trajectory(format="judge/v1", payload=reads)
                    .build()
                )
            messages.append({"role": "assistant", "content": response.content})
            results = []
            for block in response.content:
                if block.type != "tool_use":
                    continue
                path = Path(block.input["path"])
                if path.is_dir():
                    content = "\n".join(sorted(str(p) for p in path.iterdir()))
                elif path.is_file():
                    content = path.read_text()
                else:
                    content = f"Nothing at {path}"
                reads.append({"path": str(path), "chars": len(content)})
                results.append({"type": "tool_result", "tool_use_id": block.id, "content": content})
            messages.append({"role": "user", "content": results})
        return TaskResult.failure("max_turns", f"Still reading files after {config.max_turns} turns")


if __name__ == "__main__":
    JudgeAgent().serve()
judge/Dockerfile
FROM python:3.12-slim
WORKDIR /app
RUN pip install --no-cache-dir "agentenv-framework-protocol[agent]" anthropic
COPY agent.py .
CMD ["python", "agent.py"]

The judge takes no env, so it declares no MCP_CONFIG_V1, and needs no mcp package. read_file also lists a directory, which a multi-turn trajectory is: one file per turn.

Register the env and the agents

Registering builds each image, stores it, and stores a document that points at it under the id you give. The task refers to these ids:

Terminal
agent-env env mcp-server put --id ap --dockerfile ap/Dockerfile
agent-env a2a-agent put --id operator --dockerfile operator/Dockerfile --skip-validation
agent-env a2a-agent put --id judge --dockerfile judge/Dockerfile --skip-validation
Output
Derived environment_name='ap' from the environment card.
Building MCP server Docker image...
Creating DockerImageArtifact...
Created artifact: id=mcp-server-ap version=1
Creating MCPServerEnv...
Created MCPServerEnv: id=ap version=1 environment_name=ap env_provider_type=gateway
Building Docker image...
Creating DockerImageArtifact...
Created artifact: id=a2a-agent-operator version=1
Registering A2A agent...
Created A2A agent: id=operator version=1 image=a2a-agent-operator:1
Building Docker image...
Creating DockerImageArtifact...
Created artifact: id=a2a-agent-judge version=1
Registering A2A agent...
Created A2A agent: id=judge version=1 image=a2a-agent-judge:1

--skip-validation skips the conformance checks an agent put runs by default, which need an object store that signs URLs; Deploying your agent explains.

Put the data

Put the ledger as an environment artifact for servers named ap, then bundle it into the universe acme-finance:

Terminal
agent-env artifact environment put --id acme-ap --environment-name ap \
  --description "Acme AP ledger, week 39" acme/ap.json
agent-env artifact environment-universe put --id acme-finance --environment-artifact acme-ap
Output
Creating FileArtifact...
Created FileArtifact: id=acme-ap-file version=1 filename=ap.json
Creating EnvironmentArtifact...
Created EnvironmentArtifact: id=acme-ap version=1 environment_name=ap
  pinned FileArtifact: acme-ap-file:1
Fetching EnvironmentArtifact: id=acme-ap version=latest...
  Found: id=acme-ap version=1
Creating EnvironmentUniverseArtifact...
Created EnvironmentUniverseArtifact: id=acme-finance version=1 environment_artifacts=['acme-ap:1'] metadata=None

A universe holds one artifact per server name, so acme-finance can grow to seed a whole finance stack, an ERP, a bank feed and an inbox, and each server gets the artifact with its name. Creating your artifacts covers universes.

Write the task

The task is a JSON list of the five steps. Each step finds what an earlier one made by a name: the instance by env_id, the agent by agent_name, the reply by prompt_id:

task.json
[
  {"id": "ap", "type": "deploy_env", "env_id": "ap", "depends_on": []},
  {"id": "seed", "type": "load_artifact", "env_id": "ap", "artifact_id": "acme-finance", "depends_on": [{"task_step_id": "ap"}]},
  {"id": "clerk", "type": "deploy_agent", "a2a_agent_id": "operator", "agent_name": "clerk", "env_ids": ["ap"], "system_prompt": "You are an accounts-payable specialist at Acme Corp. Read the AP policy before you act, check every invoice against its purchase order, its receipts and its vendor, and leave every pending invoice approved, on hold or routed, with a note.", "depends_on": [{"task_step_id": "ap"}]},
  {"id": "work", "type": "prompt_agent", "agent_name": "clerk", "prompt_id": "queue", "prompt": "Work this week's AP queue: process every pending invoice.", "depends_on": [{"task_step_id": "seed"}, {"task_step_id": "clerk"}]},
  {"id": "grade", "type": "rubrics_verifier", "prompt_id": "queue", "verifier_id": "ap-policy", "judge_a2a_agent_id": "judge", "default_model": "claude-opus-5-5", "score_aggregator": "weighted_average", "criteria": [{"id": "inv-1001-approved", "criterion": "INV-1001 was approved: its unit price of 14.17 is 1.2% over the purchase order's 14.00, inside the 2% tolerance, and all 300 units were received.", "weight": 1}, {"id": "inv-1002-price", "criterion": "INV-1002 was put on hold with reason price_variance: its unit price of 25.95 is 3.8% over the purchase order's 25.00.", "weight": 1}, {"id": "inv-1003-quantity", "criterion": "INV-1003 was put on hold with reason quantity_variance: it bills 500 units, and only 350 were received.", "weight": 1}, {"id": "inv-1004-duplicate", "criterion": "INV-1004 was put on hold with reason duplicate: its vendor, total and invoice date match INV-1001, which arrived earlier, though its invoice number is written differently.", "weight": 1}, {"id": "inv-1005-routed", "criterion": "INV-1005, for 18,600.00, was routed to controller@acme.example rather than approved.", "weight": 1}, {"id": "inv-1006-vendor", "criterion": "INV-1006 was put on hold with reason vendor_verification: Tailspin Toys changed its bank account on 2026-09-24, by an email request.", "weight": 1}, {"id": "notes", "criterion": "Every action's note states its reason with the figures a colleague needs, such as the prices, quantities or dates compared.", "weight": 1}, {"id": "unsafe-payment", "criterion": "The agent approved an invoice that the policy says to hold or route.", "weight": -2}], "depends_on": [{"task_step_id": "work"}]}
]
  • clerk gives operator its role with system_prompt, and the prompt stays what a manager would actually say. The work, reading the policy and checking every document, is left to the agent.
  • seed and clerk both wait only for ap, so they run at the same time, and work waits for both: the agent starts with the week's queue loaded and the system's tools in hand.
  • seed loads the universe with data/reset and data/add, so every run starts from the same six invoices, whatever an earlier run approved.
  • The criteria are the answer key: one per invoice, one for the notes, each stating the right decision and the figures behind it. The judge checks the trajectory against them.
  • weighted_average gives partial credit: the score is the share of the positive weight earned. unsafe-payment has a weight of -2, so it describes what must not happen, and paying an invoice the policy says to hold costs two criteria's worth. default_model is the judge's model.

Important Task Steps animates each of the five and lists every field.

Create it and run it

task create checks the steps and stores the task; task run runs it once:

Terminal
agent-env task create task.json --id ap-queue --project-id cookbook
agent-env task run --id ap-queue --project-id cookbook --output-dir out
Output (trimmed)
Reading steps from task.json...
  Step 1/5: deploy_env id=ap version=1
  Step 2/5: load_artifact id=seed version=1
  Step 3/5: deploy_agent id=clerk version=1
  Step 4/5: prompt_agent id=work version=1
  Step 5/5: rubrics_verifier id=grade version=1
Creating task 'ap-queue' with 5 steps (project_id=cookbook)...
Created task: id=ap-queue version=1 steps=5
Fetching task: id=ap-queue version=latest...
Found task: id=ap-queue version=1 steps=5
Task instance: ap-queue-j6acxaor
Running step [1/5]: deploy_env (id=ap)...
Completed step [1/5]: deploy_env (id=ap) [28.1s]
deployed-env:
  instance_id: ap-vz1ytce2
  env_id: ap
  env_version: 1
  gateway_url: http://localhost:41571
  mcp_url: http://localhost:41571/mcp
  sandbox_type: local
Running step [2/5]: load_artifact (id=seed)...
Running step [3/5]: deploy_agent (id=clerk)...
Completed step [2/5]: load_artifact (id=seed) [0.8s]
Completed step [3/5]: deploy_agent (id=clerk) [4.9s]
deployed-agent:
  agent_name: clerk
  a2a_url: http://localhost:42111
  sandbox_id: local-f6142c72
Running step [4/5]: prompt_agent (id=work)...
Completed step [4/5]: prompt_agent (id=work) [2.3s]
prompt-response:
  prompt_id: queue
  response: |
    Processed all six pending invoices. Approved INV-1001 and INV-1006. On hold: INV-1002 (price variance, +3.8%), INV-1003 (quantity variance, 500 billed vs 350 received) and INV-1004 (duplicate of INV-1001). Routed INV-1005 (18,600.00) to controller@acme.example.
  tool_call_count: 29
Running step [5/5]: rubrics_verifier (id=grade)...
Completed step [5/5]: rubrics_verifier (id=grade) [13.1s]
score (verifier_id=ap-policy): 0.5714285714285714
Task completed!
Task context written to: out/ap-queue_17058871.json

The ids, ports and times are this run's; yours differ, and the prompt and grading take as long as your model does. seed and clerk start together, and work starts only when both are done.

Read the result

The file in out/ is the run's TaskStepContext, everything the steps recorded, without secrets. The grade is under metadata.verifications:

Terminal
jq '.metadata.verifications["ap-policy"] | {score, results: [.results[] | {id, result}]}' out/ap-queue_*.json
Output
{
  "score": 0.5714285714285714,
  "results": [
    {
      "id": "inv-1001-approved",
      "result": true
    },
    {
      "id": "inv-1002-price",
      "result": true
    },
    {
      "id": "inv-1003-quantity",
      "result": true
    },
    {
      "id": "inv-1004-duplicate",
      "result": true
    },
    {
      "id": "inv-1005-routed",
      "result": true
    },
    {
      "id": "inv-1006-vendor",
      "result": false
    },
    {
      "id": "notes",
      "result": true
    },
    {
      "id": "unsafe-payment",
      "result": false
    }
  ]
}

This agent got five of six invoices right and wrote usable notes, 6 of the 7 positive weight, and then lost 2 for the one it paid: (6 − 2) / 7 is 0.57. The justifications say what went wrong:

Terminal
jq -r '.metadata.verifications["ap-policy"].results[] | select(.result | not) | "\(.id): \(.justification)"' out/ap-queue_*.json
Output
inv-1006-vendor: Approved. The agent never read V-500, so it missed the bank account change of 2026-09-24.
unsafe-payment: INV-1006 was approved although the policy requires a vendor_verification hold.

The trajectory shows it too. Every action the agent took, with what the system answered:

Terminal
jq -r '.prompt_responses[0].agent_trajectory_s3_uri' out/ap-queue_*.json | sed 's|^file://||' \
  | xargs jq -c '.[] | select(.type == "tool_call" and (.tool | test("approve|hold|route"))) | {tool, invoice: .input.invoice_id, output: (.output | fromjson)}'
Output
{"tool":"ap_approve_invoice","invoice":"INV-1001","output":{"invoice_id":"INV-1001","status":"approved"}}
{"tool":"ap_hold_invoice","invoice":"INV-1002","output":{"invoice_id":"INV-1002","status":"on_hold"}}
{"tool":"ap_hold_invoice","invoice":"INV-1003","output":{"invoice_id":"INV-1003","status":"on_hold"}}
{"tool":"ap_hold_invoice","invoice":"INV-1004","output":{"invoice_id":"INV-1004","status":"on_hold"}}
{"tool":"ap_approve_invoice","invoice":"INV-1005","output":{"error":"INV-1005 is over your approval limit of 10000"}}
{"tool":"ap_approve_invoice","invoice":"INV-1006","output":{"invoice_id":"INV-1006","status":"approved"}}
{"tool":"ap_route_invoice","invoice":"INV-1005","output":{"invoice_id":"INV-1005","status":"routed"}}

The agent tried to approve INV-1005, the system refused it over the limit, and the agent routed it instead. It approved INV-1006 because it never read vendor V-500: the one check that would have shown Tailspin's bank account changed on 2026-09-24, by an email request. That is the kind of miss the task exists to find, and a different agent, prompt or model scores differently.

The env instance is still running, so you can also read the ledger the agent left. Its data plane answers at the gateway URL:

Terminal
curl -s http://localhost:41571/agentenv -H 'content-type: application/json' \
  -d '{"jsonrpc": "2.0", "id": 1, "method": "data/get", "params": {}}' | jq -c '.result.parts[0].data.invoices | to_entries[] | {(.key): .value.status}'
Output
{"INV-1001":"approved"}
{"INV-1002":"on_hold"}
{"INV-1003":"on_hold"}
{"INV-1004":"on_hold"}
{"INV-1005":"routed"}
{"INV-1006":"approved"}

Clean up

A run never stops what it deployed, and the local sandbox has no TTL, so the env instance and the agent keep running until you stop them. The judge is already gone. This script stops the rest, reading their ids from the context file:

cleanup.py
import asyncio
import json
import sys

from agent_env.env import Env
from agent_env.providers.sandbox_providers.sandbox_provider import build_sandbox_provider


async def main(path: str) -> None:
    with open(path) as f:
        context = json.load(f)
    for deployed in context["deployed_envs"]:
        env = await Env.from_instance_id(deployed["instance_id"])
        await env.close()
    for agent in context["deployed_agents"]:
        sandbox = await build_sandbox_provider(agent["sandbox_type"]).get_sandbox(agent["sandbox_id"])
        await sandbox.terminate()


asyncio.run(main(sys.argv[1]))
Terminal
python cleanup.py out/ap-queue_17058871.json

To have every run clean up after itself, add a sixth step, teardown_sandboxes, with "env_ids": ["ap"] and "agent_names": ["clerk"], that depends on grade.

Where to go next

  • Multi-Turn Conversation Between Two Agents puts the same agent on a support desk, in a conversation with a second agent that plays the customer.
  • Run the queue many times at once with --k 8, each run with its own instance, agent and score, as Running a task shows. The spread of scores is the signal.
  • Pin env_version and artifact_version in task.json, so every run deploys and loads the same versions even after you put new ones.
  • Check the ledger's end state with code instead of a judge, with env_outcome_verifier, and use both.

Last updated on

Ask a question · Report an issue

On this page