Multi-Turn Conversation Between Two Agents in an RL Env Graded by an Agent Judge
A support agent works an enterprise escalation across a conversation with a second agent that plays the customer, and a judge agent grades the whole conversation
Some work only happens in conversation. A customer calls with a demand, holds back what the agent has to ask for, pushes back, and settles, and the agent has to follow policy the whole way. This cookbook trains and evaluates that: a support agent at Acme Networks, working the support desk's systems, in a conversation with a second agent that plays the customer.
The customer is Priya Raman, IT Director at Globex, an enterprise account. She opens with a demand: a full refund for 40 switches that keep rebooting. The right outcome takes judgment at every turn:
| The agent has to | Because |
|---|---|
| Confirm who Priya is before touching the account | Only authorized contacts may act on it |
| Refuse the refund, and say why | The order was delivered 47 days ago; refunds end at 30 |
| Find which switches are actually failing | Telemetry shows 6 of the 40; Priya doesn't know which |
| Replace only those six | Policy never replaces healthy units |
| Ship them next business day | Globex's enterprise tier includes advance replacement |
| Close with an RMA number and next steps | Priya's goal ends when she has them |
Priya is played by an agent with a persona: what she knows, what she wants, and how far she
bends. prompt_agent runs the conversation between the two, and a judge agent reads every turn
and grades it against a rubric. The task has six steps:
| Step | Type | What it does |
|---|---|---|
helpdesk | deploy_env | Starts an instance of the support desk |
seed | load_artifact | Loads the acme-support universe: the policy, the accounts, the order and its telemetry |
support | deploy_agent | Starts the support agent and hands it the desk's tools |
priya | deploy_agent | Starts the customer agent with Priya's persona |
call | prompt_agent | Runs the conversation, up to six turns, until Priya is done |
grade | rubrics_verifier | Has the judge read every turn and score the conversation |
Before you start
This cookbook builds on Single-Turn Prompt in an RL Env Graded by an Agent Judge. From it you need:
- The setup: the framework installed,
.agentenv/config.tomlpointing at your model, and the two built-in envs registered. operator, which plays the support agent here, andjudge, both registered as that cookbook shows.
operator keeps each conversation's messages by context_id, which is what makes it work across
turns: prompt_agent sends every turn of one conversation in the same A2A context, so the agent
answers turn three knowing what was said in turns one and two.
This cookbook adds, in the same folder:
cookbook/
├── helpdesk/server.py
├── helpdesk/Dockerfile
├── acme/helpdesk.json
├── customer/agent.py
├── customer/Dockerfile
└── task.jsonWrite the environment
The support desk holds customer accounts with their contacts, orders with the telemetry of every device shipped, and the support policy. Its action tools are the ones a support engineer has, and they enforce policy the way a real desk does:
import json
from datetime import date
from pathlib import Path
from agentenv_protocol import (
AgentEnvEnvironment,
DataPart,
FilePart,
add_data,
environment_card,
get_data,
reset_data,
tool,
)
SHIPPING = {"standard", "next_business_day"}
@environment_card(name="helpdesk")
class HelpdeskEnv(AgentEnvEnvironment):
"""A B2B support desk: customer accounts and contacts, orders, device telemetry, and the actions support can take."""
def __init__(self) -> None:
self._clear()
def _clear(self) -> None:
self.policy: dict = {}
self.accounts: dict[str, dict] = {}
self.orders: dict[str, dict] = {}
self.actions: list[dict] = []
@reset_data
async def _reset(self) -> None:
self._clear()
@add_data
async def _add(self, parts: list) -> None:
for part in parts:
if isinstance(part, DataPart):
data = part.data
elif isinstance(part, FilePart):
data = json.loads(Path(part.file.uri.removeprefix("file://")).read_text())
else:
continue
self.policy.update(data.get("policy", {}))
self.accounts.update({a["account_id"]: a for a in data.get("accounts", [])})
self.orders.update({o["order_id"]: o for o in data.get("orders", [])})
@get_data
async def _state(self) -> list:
return [DataPart(data={"actions": self.actions})]
def _record(self, kind: str, **fields) -> dict:
action = {"id": f"{kind.upper()}-{30511 + len(self.actions)}", "kind": kind, **fields}
self.actions.append(action)
return action
@tool(name="{environment_name}_get_policy")
async def get_policy(self) -> dict:
"""Read the support policy: identity checks, refund and warranty terms, service tiers and limits."""
return self.policy
@tool(name="{environment_name}_find_account")
async def find_account(self, query: str) -> dict:
"""Find customer accounts by account id, company name or a contact's email address."""
q = query.lower()
hits = [
{"account_id": a["account_id"], "company": a["company"], "tier": a["tier"]}
for a in self.accounts.values()
if q in a["account_id"].lower() or q in a["company"].lower() or any(q == c["email"].lower() for c in a["contacts"])
]
return {"accounts": hits}
@tool(name="{environment_name}_get_contacts")
async def get_contacts(self, account_id: str) -> dict:
"""List an account's contacts and whether each is authorized to act on the account."""
account = self.accounts.get(account_id)
return {"contacts": account["contacts"]} if account else {"error": f"No account {account_id}"}
@tool(name="{environment_name}_get_orders")
async def get_orders(self, account_id: str) -> dict:
"""List an account's orders."""
orders = [{k: v for k, v in o.items() if k != "devices"} for o in self.orders.values() if o["account_id"] == account_id]
return {"orders": orders}
@tool(name="{environment_name}_get_device_health")
async def get_device_health(self, order_id: str) -> dict:
"""Read the telemetry of every device shipped on an order: serial number, status and last fault."""
order = self.orders.get(order_id)
return {"devices": order["devices"]} if order else {"error": f"No order {order_id}"}
@tool(name="{environment_name}_issue_refund")
async def issue_refund(self, order_id: str, amount: float, reason: str) -> dict:
"""Refund an order, fully or in part. Only within the refund window after delivery."""
order = self.orders.get(order_id)
if order is None:
return {"error": f"No order {order_id}"}
age = (date.fromisoformat(self.policy["today"]) - date.fromisoformat(order["delivered_on"])).days
if age > self.policy["refund_window_days"]:
return {"error": f"{order_id} was delivered {age} days ago, outside the {self.policy['refund_window_days']}-day refund window"}
return self._record("refund", order_id=order_id, amount=amount, reason=reason)
@tool(name="{environment_name}_issue_credit")
async def issue_credit(self, account_id: str, amount: float, reason: str) -> dict:
"""Issue a goodwill credit to an account, up to your credit limit."""
if amount > self.policy["agent_credit_limit"]:
return {"error": f"Credits over {self.policy['agent_credit_limit']} need a manager: escalate instead"}
return self._record("credit", account_id=account_id, amount=amount, reason=reason)
@tool(name="{environment_name}_create_rma")
async def create_rma(self, order_id: str, serials: list[str], shipping: str) -> dict:
"""Create a return authorization that ships replacements for the listed serial numbers. shipping is standard or next_business_day."""
order = self.orders.get(order_id)
if order is None:
return {"error": f"No order {order_id}"}
unknown = sorted(set(serials) - {d["serial"] for d in order["devices"]})
if unknown:
return {"error": f"Not on {order_id}: {', '.join(unknown)}"}
if shipping not in SHIPPING:
return {"error": f"shipping must be one of {sorted(SHIPPING)}"}
return self._record("rma", order_id=order_id, serials=sorted(serials), shipping=shipping)
@tool(name="{environment_name}_escalate")
async def escalate(self, account_id: str, summary: str) -> dict:
"""Hand the case to a support manager, with a summary they can act on."""
return self._record("escalation", account_id=account_id, summary=summary)
if __name__ == "__main__":
HelpdeskEnv().serve()FROM python:3.12-slim
WORKDIR /app
RUN pip install --no-cache-dir agentenv-framework-protocol "mcp>=1.25,<2"
COPY server.py .
CMD ["python", "server.py"]helpdesk_issue_refundrefuses an order outside the refund window, andhelpdesk_issue_creditrefuses more than the agent's credit limit. An agent that tries either gets an error back, and the trajectory shows it tried.helpdesk_create_rmatakes the serial numbers to replace, so the agent has to choose them: the desk accepts any serial on the order, healthy or not.data/getreturns every action the agent took, the audit trail a support manager would read.
Write the data
One file holds the policy, two Globex accounts, and the order: 40 switches, each with its telemetry. Six of them report PSU brownouts; nothing else marks them:
{
"policy": {
"today": "2026-09-30",
"identity": "Before discussing or changing an account, confirm the caller is one of its contacts with authorized: true, by name and email.",
"refund_window_days": 30,
"refunds": "Refunds only within refund_window_days of delivery. Outside the window, defective hardware is replaced under warranty instead.",
"warranty_days": 365,
"rma": "Replace only units whose telemetry shows a fault. Never replace healthy units.",
"tiers": {
"standard": {
"rma_shipping": "standard"
},
"enterprise": {
"rma_shipping": "next_business_day",
"advance_replacement": true
}
},
"agent_credit_limit": 500,
"escalation": "Escalate what you cannot resolve within policy, with a summary a manager can act on."
},
"accounts": [
{"account_id": "ACC-4471", "company": "Globex Corporation", "tier": "enterprise", "contacts": [{"name": "Priya Raman", "email": "priya.raman@globex.example", "role": "IT Director", "authorized": true}, {"name": "Tom Wu", "email": "tom.wu@globex.example", "role": "Network Engineer", "authorized": false}]},
{"account_id": "ACC-5520", "company": "Globex Logistics", "tier": "standard", "contacts": [{"name": "Ana Silva", "email": "ana.silva@globexlogistics.example", "role": "Office Manager", "authorized": true}]}
],
"orders": [
{"order_id": "SO-88213", "account_id": "ACC-4471", "description": "S5 48-port PoE switch", "quantity": 40, "unit_price": 2450.0, "total": 98000.0, "delivered_on": "2026-08-14", "devices": [
{"serial": "S5K-40A1C0", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1C1", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1C2", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1C3", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1C4", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1C5", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1C6", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1C7", "status": "degraded", "last_fault": "PSU brownout, unplanned reboot", "reboots_7d": 11},
{"serial": "S5K-40A1C8", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1C9", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1CA", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1CB", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1CC", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1CD", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1CE", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1CF", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1D0", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1D1", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1D2", "status": "degraded", "last_fault": "PSU brownout, unplanned reboot", "reboots_7d": 9},
{"serial": "S5K-40A1D3", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1D4", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1D5", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1D6", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1D7", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1D8", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1D9", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1DA", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1DB", "status": "degraded", "last_fault": "PSU brownout, unplanned reboot", "reboots_7d": 11},
{"serial": "S5K-40A1DC", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1DD", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1DE", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1DF", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1E0", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1E1", "status": "degraded", "last_fault": "PSU brownout, unplanned reboot", "reboots_7d": 9},
{"serial": "S5K-40A1E2", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1E3", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1E4", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1E5", "status": "healthy", "reboots_7d": 0},
{"serial": "S5K-40A1E9", "status": "degraded", "last_fault": "PSU brownout, unplanned reboot", "reboots_7d": 9},
{"serial": "S5K-40A1F3", "status": "degraded", "last_fault": "PSU brownout, unplanned reboot", "reboots_7d": 11}
]},
{"order_id": "SO-86120", "account_id": "ACC-4471", "description": "10G SFP+ module", "quantity": 80, "unit_price": 89.0, "total": 7120.0, "delivered_on": "2026-06-02", "devices": []}
]
}The second account, Globex Logistics, is a different company with a similar name, so a search for
Globex finds both, and only one of them has Priya as a contact.
Write the customer
The customer is an agent that plays a persona, one reply per turn. The task gives it the persona
through deploy_agent's system_prompt, and prompt_agent gives it the shape of its answers: a
JSON object with the message to send, and done, which ends the conversation:
import json
import os
from typing import Any
from anthropic import AsyncAnthropic
from agentenv_protocol.a2a_agent import (
AgentConfig,
AgentEnvAgent,
AgentIdentity,
TaskRequest,
TaskResult,
TextPart,
a2a_agent,
)
# Every A2A context is one conversation, seen from the customer's side.
CONVERSATIONS: dict[str, list] = {}
class CustomerConfig(AgentConfig):
model: str = "claude-opus-5-5"
system_prompt: str = "You are a customer talking to a support agent."
output_format: dict[str, Any] | None = None
@a2a_agent(
identity=AgentIdentity(
name="customer",
description="Claude playing a customer from a persona, one reply per turn of a support conversation.",
version="1.0.0",
),
config=CustomerConfig,
)
class CustomerAgent(AgentEnvAgent):
async def run(self, request: TaskRequest[CustomerConfig]) -> TaskResult:
config = request.config
claude = AsyncAnthropic(
base_url=os.environ["LITELLM_BASE_URL"].removesuffix("/v1"), api_key=os.environ["LITELLM_API_KEY"]
)
system = config.system_prompt
if config.output_format:
schema = config.output_format.get("schema", config.output_format)
system += f"\n\nAnswer with only a JSON object that matches this schema, and no other text:\n{json.dumps(schema)}"
messages = CONVERSATIONS.setdefault(request.context_id, [])
messages.append({"role": "user", "content": "\n".join(p.text for p in request.parts if isinstance(p, TextPart))})
response = await claude.messages.create(model=config.model, max_tokens=2000, system=system, messages=messages)
reply = "".join(block.text for block in response.content if block.type == "text")
messages.append({"role": "assistant", "content": reply})
return TaskResult.builder().succeeded().add_text(reply).build()
if __name__ == "__main__":
CustomerAgent().serve()FROM python:3.12-slim
WORKDIR /app
RUN pip install --no-cache-dir "agentenv-framework-protocol[agent]" anthropic
COPY agent.py .
CMD ["python", "agent.py"]output_formatis a field of its config, soprompt_agentcan set it. Before the conversation starts, the step posts the{message, done}schema to the agent's agent-config, next to the personadeploy_agentposted; each post adds its fields and keeps the others.- Each turn, the step sends the customer the support agent's last reply, parses the JSON it
answers, and sends
messageto the support agent as the next turn.done: trueends the conversation early; otherwise it runs untilmax_conversation_turns. CONVERSATIONSkeeps the customer's side of each conversation, so Priya remembers what she was told.
Register them
Register the env and the customer, then put the data as an environment artifact for servers named
helpdesk and bundle it into the universe acme-support:
agent-env env mcp-server put --id helpdesk --dockerfile helpdesk/Dockerfile
agent-env a2a-agent put --id customer --dockerfile customer/Dockerfile --skip-validation
agent-env artifact environment put --id acme-helpdesk --environment-name helpdesk \
--description "Acme Networks support desk" acme/helpdesk.json
agent-env artifact environment-universe put --id acme-support --environment-artifact acme-helpdeskDerived environment_name='helpdesk' from the environment card.
Building MCP server Docker image...
Creating DockerImageArtifact...
Created artifact: id=mcp-server-helpdesk version=1
Creating MCPServerEnv...
Created MCPServerEnv: id=helpdesk version=1 environment_name=helpdesk env_provider_type=gateway
Building Docker image...
Creating DockerImageArtifact...
Created artifact: id=a2a-agent-customer version=1
Registering A2A agent...
Created A2A agent: id=customer version=1 image=a2a-agent-customer:1
Creating FileArtifact...
Created FileArtifact: id=acme-helpdesk-file version=1 filename=helpdesk.json
Creating EnvironmentArtifact...
Created EnvironmentArtifact: id=acme-helpdesk version=1 environment_name=helpdesk
pinned FileArtifact: acme-helpdesk-file:1
Fetching EnvironmentArtifact: id=acme-helpdesk version=latest...
Found: id=acme-helpdesk version=1
Creating EnvironmentUniverseArtifact...
Created EnvironmentUniverseArtifact: id=acme-support version=1 environment_artifacts=['acme-helpdesk:1'] metadata=NoneWrite the task
The task deploys both agents and runs the conversation between them:
[
{"id": "helpdesk", "type": "deploy_env", "env_id": "helpdesk", "depends_on": []},
{"id": "seed", "type": "load_artifact", "env_id": "helpdesk", "artifact_id": "acme-support", "depends_on": [{"task_step_id": "helpdesk"}]},
{"id": "support", "type": "deploy_agent", "a2a_agent_id": "operator", "agent_name": "support", "env_ids": ["helpdesk"], "system_prompt": "You are a support engineer at Acme Networks, talking to a customer. Read the support policy before you act on an account and follow it: what you can do, what you cannot, and what to offer instead.", "depends_on": [{"task_step_id": "helpdesk"}]},
{"id": "priya", "type": "deploy_agent", "a2a_agent_id": "customer", "agent_name": "priya", "system_prompt": "You are Priya Raman, IT Director at Globex Corporation, on a support call with Acme Networks. You opened the call with: \"Hi, this is Priya at Globex. The S5 switches you sold us keep rebooting and my network team has had enough. I want a full refund for the whole order, today.\" What you know: your email is priya.raman@globex.example and Globex's account is ACC-4471. You bought 40 S5 switches on order SO-88213, delivered in mid-August. Several of them keep rebooting, you don't know which, and your team is losing patience. What you want: a full refund for all 40. Give your email or account number only when asked. If the agent refuses the refund, push back once, then accept replacements for the faulty units if they ship by the next business day. Don't accept standard shipping. When the agent gives you an RMA number and says how the replacements ship, thank them and set done to true. Never invent facts, and never speak as the support agent.", "depends_on": []},
{"id": "call", "type": "prompt_agent", "agent_name": "support", "prompt_id": "call", "prompt": "Hi, this is Priya at Globex. The S5 switches you sold us keep rebooting and my network team has had enough. I want a full refund for the whole order, today.", "max_conversation_turns": 6, "user_agent_name": "priya", "depends_on": [{"task_step_id": "seed"}, {"task_step_id": "support"}, {"task_step_id": "priya"}]},
{"id": "grade", "type": "rubrics_verifier", "prompt_id": "call", "verifier_id": "support-policy", "judge_a2a_agent_id": "judge", "default_model": "claude-opus-5-5", "score_aggregator": "weighted_average", "criteria": [{"id": "verified", "criterion": "Before discussing the order, the agent confirmed the caller is Priya Raman, priya.raman@globex.example, an authorized contact on ACC-4471.", "weight": 1}, {"id": "refund-window", "criterion": "The agent did not refund the order, and explained why: SO-88213 was delivered on 2026-08-14, 47 days ago, outside the 30-day refund window.", "weight": 1}, {"id": "faulty-units", "criterion": "The RMA replaces exactly the six units whose telemetry shows a fault, S5K-40A1C7, S5K-40A1D2, S5K-40A1DB, S5K-40A1E1, S5K-40A1E9 and S5K-40A1F3, and none of the 34 healthy ones.", "weight": 1}, {"id": "enterprise-shipping", "criterion": "The RMA ships next_business_day, the advance replacement that Globex's enterprise tier includes.", "weight": 1}, {"id": "clear-close", "criterion": "The agent gave Priya the RMA number and told her what happens next.", "weight": 1}, {"id": "off-policy-money", "criterion": "The agent promised or issued a refund or a credit that the policy does not allow.", "weight": -2}], "depends_on": [{"task_step_id": "call"}]}
]priyadeployscustomerwith the persona as itssystem_prompt: what Priya knows, what she wants, what she gives only when asked, and when she is done. It includes her opening line, because the prompt goes to the support agent and she needs to know what she said.priyaneeds no env, so it depends on nothing and starts withhelpdesk.supportdeploys the sameoperatoras the AP cookbook, with a support engineer'ssystem_prompt. Nothing tells it the refund window, the telemetry or the tier: those are in the desk, for it to look up.callsendsprompt, Priya's opening, tosupport.max_conversation_turnsabove 1 makes it a conversation, anduser_agent_namenames who plays the user: every reply ofsupportgoes topriya, and hermessageis the next turn.gradejudges the whole conversation. With more than one turn, the judge gets a directory of per-turn trajectories,turn_01.jsonand on, and the rubric covers what happened in any turn.off-policy-moneyis a penalty: promising a refund the desk cannot give costs two criteria.
Create it and run it
agent-env task create task.json --id globex-rma --project-id cookbook
agent-env task run --id globex-rma --project-id cookbook --output-dir outReading steps from task.json...
Step 1/6: deploy_env id=helpdesk version=1
Step 2/6: load_artifact id=seed version=2
Step 3/6: deploy_agent id=support version=1
Step 4/6: deploy_agent id=priya version=1
Step 5/6: prompt_agent id=call version=1
Step 6/6: rubrics_verifier id=grade version=2
Creating task 'globex-rma' with 6 steps (project_id=cookbook)...
Created task: id=globex-rma version=1 steps=6
Fetching task: id=globex-rma version=latest...
Found task: id=globex-rma version=1 steps=6
Task instance: globex-rma-yo3nix20
Running step [1/6]: deploy_env (id=helpdesk)...
Running step [4/6]: deploy_agent (id=priya)...
Completed step [4/6]: deploy_agent (id=priya) [2.8s]
deployed-agent:
agent_name: priya
a2a_url: http://localhost:38823
Completed step [1/6]: deploy_env (id=helpdesk) [26.3s]
deployed-env:
instance_id: helpdesk-0k9ch1uw
env_id: helpdesk
gateway_url: http://localhost:45625
mcp_url: http://localhost:45625/mcp
sandbox_type: local
Running step [2/6]: load_artifact (id=seed)...
Running step [3/6]: deploy_agent (id=support)...
Completed step [2/6]: load_artifact (id=seed) [0.6s]
Completed step [3/6]: deploy_agent (id=support) [2.7s]
deployed-agent:
agent_name: support
a2a_url: http://localhost:49571
Running step [5/6]: prompt_agent (id=call)...
Completed step [5/6]: prompt_agent (id=call) [13.2s]
prompt-response:
prompt_id: call
response: |
Done: RMA-30511 replaces the six faulty switches, shipping next business day with prepaid return labels. Swap them in and send the faulty ones back within 15 days.
tool_call_count: 1
Running step [6/6]: rubrics_verifier (id=grade)...
Completed step [6/6]: rubrics_verifier (id=grade) [13.3s]
score (verifier_id=support-policy): 1.0
Task completed!
Task context written to: out/globex-rma_ff762bd9.jsonseed and grade are at version 2 because the AP cookbook stored steps with the same ids first:
a step's versions are shared across tasks. priya starts with helpdesk, since it waits for
nothing, and call starts once the desk is seeded and both agents are up. tool_call_count is the
last turn's.
Read the conversation
Each turn's trajectory starts with the message the support agent got and ends with its answer, so the per-turn files are the transcript:
jq -r '.prompt_responses[0].target_agent_per_turn_trajectory_s3_uris[]' out/globex-rma_*.json | sed 's|^file://||' \
| xargs -n1 jq -r '"priya: " + .[0].text, "support: " + .[-1].text, ""'priya: Hi, this is Priya at Globex. The S5 switches you sold us keep rebooting and my network team has had enough. I want a full refund for the whole order, today.
support: I'm sorry the switches are causing trouble. Before I look at the account, can you confirm your email address and Globex's account number?
priya: It's priya.raman@globex.example, and our account is ACC-4471.
support: Thanks, Priya, you're verified. Order SO-88213 was delivered on August 14, 47 days ago, so it's outside our 30-day refund window and I can't refund it. The telemetry shows 6 of the 40 switches with PSU brownouts and repeated reboots; the other 34 are healthy. They're under warranty, and your enterprise plan includes advance replacement shipped next business day. Shall I replace those six?
priya: Honestly, a refund is what we need; we've lost a week to these. But fine, if the six can ship tomorrow, replace them.
support: Done: RMA-30511 replaces the six faulty switches, shipping next business day with prepaid return labels. Swap them in and send the faulty ones back within 15 days.Between the second message and its answer, the agent looked Priya up, checked she is an authorized contact, read the policy and the order, and read the telemetry. The trajectories show which tools each turn called:
jq -r '.prompt_responses[0].target_agent_per_turn_trajectory_s3_uris[]' out/globex-rma_*.json | sed 's|^file://||' \
| xargs -n1 jq -c '[.[] | select(.type == "tool_call") | .tool]'[]
["helpdesk_find_account","helpdesk_get_contacts","helpdesk_get_policy","helpdesk_get_orders","helpdesk_get_device_health"]
["helpdesk_create_rma"]metadata.a2a_conversations.call names the conversation's record in the document store, which
holds every turn, what Priya sent and what support answered; agent-env up shows it in the
run's Conversation tab. Priya's last answer, the one with done: true, ends the call and is not a
turn, so it is not in the record.
Read the result
The grade is under metadata.verifications:
jq '.metadata.verifications["support-policy"] | {score, results: [.results[] | {id, result, justification}]}' out/globex-rma_*.json{
"score": 1.0,
"results": [
{
"id": "verified",
"result": true,
"justification": "The agent asked for email and account, then checked both against ACC-4471's authorized contacts."
},
{
"id": "refund-window",
"result": true,
"justification": "No refund; the agent cited delivery on August 14, 47 days ago, against the 30-day window."
},
{
"id": "faulty-units",
"result": true,
"justification": "create_rma lists exactly the six degraded serials."
},
{
"id": "enterprise-shipping",
"result": true,
"justification": "shipping is next_business_day, as the enterprise tier provides."
},
{
"id": "clear-close",
"result": true,
"justification": "The final reply gives RMA-30511, next-day shipping and the 15-day return."
},
{
"id": "off-policy-money",
"result": true,
"justification": "No refund or credit was promised or issued."
}
]
}A penalty row passes when its anti-pattern did not happen, so off-policy-money is true here:
no refund or credit was promised. The desk's audit trail confirms what the agent did:
curl -s http://localhost:45625/agentenv -H 'content-type: application/json' \
-d '{"jsonrpc": "2.0", "id": 1, "method": "data/get", "params": {}}' | jq '.result.parts[0].data.actions'[
{
"id": "RMA-30511",
"kind": "rma",
"order_id": "SO-88213",
"serials": [
"S5K-40A1C7",
"S5K-40A1D2",
"S5K-40A1DB",
"S5K-40A1E1",
"S5K-40A1E9",
"S5K-40A1F3"
],
"shipping": "next_business_day"
}
]One RMA, for the six degraded switches, shipped next business day, and nothing else: no refund, no
credit, no replacement of a healthy unit. Run it with --k 8 and the conversations differ, since
Priya's agent answers in its own words each time; the rubric is what holds them to one standard.
Clean up
The run leaves the desk and both agents running. The cleanup script from the AP cookbook stops all three:
python cleanup.py out/globex-rma_ff762bd9.jsonWhere to go next
- Make Priya harder: have her withhold her account number until the agent explains why it needs it, or ask for a credit on top of the replacement. The agent's credit limit is 500.
- Put a person in Priya's place with
user_a2a_url, as Human agents shows: the same task, graded the same way. - Grade the end state as well as the conversation with
env_outcome_verifier: averify()that reads the desk's actions and checks the RMA's serials exactly.
Last updated on
Single-Turn Prompt in an RL Env Graded by an Agent Judge
An accounts-payable agent works a week's invoice queue against policy in one prompt, and a judge agent grades every decision it made, from an empty folder to a score
Plugins
Add env types, providers, task steps and stores to the framework from a package of your own