More task steps
Every built-in step besides the important six, grouped by family with its key fields, how to list the registry, and how your own step joins it
The step registry holds 49 built-in step types. Six of them make up most tasks, and
Important Task Steps covers them: deploy_env,
load_artifact, deploy_agent, prompt_agent, rubrics_verifier and collect_artifacts. This
page lists the other 43 in seven families. The
last family, the 20 registration validators, is one you rarely put in a task yourself.
| Family | Steps |
|---|---|
| Provisioning | deploy_sandbox, run_docker_container, install_agent, teardown_sandboxes |
| Env control | reset_env, apply_server_config, modify_env_tool_access, sync_env_clock, register_env_triggers |
| Agents | add_skills, build_mcp_cli, peer_agents, register_agent_triggers |
| Capture | snapshot_env, snapshot_agent_state |
| Code and grading | run_code, env_outcome_verifier, run_container_unit_tests_verifier, verify_sandbox, agent_prompt_response_verifier, aggregate_verifiers |
| Human in the loop | review, deploy_human_agent |
| Registration validators | 20 validators, such as verify_env_card and verify_a2a_agent_card, that validate commands run |
Each table gives a step's key fields, not all of them: required fields first, then the useful
optional ones, with defaults in parentheses. Every step also takes id, type, depends_on,
fail_task_on_error and retry_config, which What is a task step
covers.
Provisioning
These steps build a machine for a task that needs one rather than an env, such as a coding task: a bare sandbox, a container built on it, and an agent installed into the container. The last one stops sandboxes, a machine's or an env's and an agent's.
| Step | What it does | Key fields |
|---|---|---|
deploy_sandbox | Starts a bare sandbox and adds it to deployed_sandboxes under sandbox_name. sandbox_mode: "vm" makes a machine that later steps build and run containers on; "container" runs a single image that listens on port, and fails without both. | sandbox_name, sandbox_mode (required); image, port, exposed_ports, cpu (2.0), memory_mb (8192), disk_size_gb (10), ttl_seconds (7200), sandbox_type, env_vars |
run_docker_container | Builds an image on a vm sandbox and starts it in the background. The build context is a file artifact universe or a .zip at a URL, exactly one of the two. ready_command is retried inside the container for up to 120 seconds. The container is recorded in metadata["deployed_docker_containers"]. | sandbox_name, and docker_context_artifact_id or docker_context_url (required); container_name (task-container), dockerfile_path (Dockerfile), ports, env_vars, build_args, command_override, keep_alive_with_base_command, ready_command, volumes |
install_agent | Installs a registered agent into a running container, or onto the vm sandbox itself when container_name is omitted, with the install commands its card declares in the urn:agentenv:install/v1 extension. It waits for the agent's card and adds the agent to deployed_agents under agent_name. | sandbox_name, a2a_agent_id (required); container_name, agent_name (default-agent), a2a_port, workspace_dir |
teardown_sandboxes | Terminates the sandboxes behind the envs, agents and sandboxes it names, to stop paying for them once nothing needs them. A run never does this on its own, and the local sandbox ignores its TTL. It is best-effort: a sandbox already gone or failing to stop is logged, and the run goes on. The ids go to metadata["torn_down_sandbox_ids"]; the deployed_envs and deployed_agents entries stay. agent-env task create rejects one with fail_task_on_error: true. | env_ids, agent_names or sandbox_names (at least one); fail_task_on_error (false, and must stay so) |
install_agent reaches the agent through the sandbox's tunnel for its A2A port, so that port must
be in the sandbox's exposed_ports and, for a container, in its ports too. A deploy_agent step
with a sandbox_name does something else: it runs the agent in a container of its own on that
sandbox, which needs port 8000 in exposed_ports.
A task without an env
This task runs a coding agent against a repository's tests. box is the machine, repo the
repository's container on it, and coder the agent installed into that container. The verifier,
run_container_unit_tests_verifier, is in Code and grading:
[
{"id": "box", "type": "deploy_sandbox", "sandbox_name": "box", "sandbox_mode": "vm", "exposed_ports": [8000], "depends_on": []},
{"id": "repo", "type": "run_docker_container", "sandbox_name": "box", "container_name": "repo", "docker_context_url": "https://example.com/repo.zip", "ports": [8000], "keep_alive_with_base_command": true, "depends_on": [{"task_step_id": "box"}]},
{"id": "coder", "type": "install_agent", "sandbox_name": "box", "container_name": "repo", "a2a_agent_id": "claude-code", "agent_name": "coder", "depends_on": [{"task_step_id": "repo"}]},
{"id": "fix", "type": "prompt_agent", "agent_name": "coder", "prompt_id": "fix", "prompt": "Make the failing tests pass.", "depends_on": [{"task_step_id": "coder"}]},
{"id": "tests", "type": "run_container_unit_tests_verifier", "sandbox_name": "box", "container_name": "repo", "command": "pytest -q", "verifier_id": "tests", "depends_on": [{"task_step_id": "fix"}]},
{"id": "teardown", "type": "teardown_sandboxes", "sandbox_names": ["box"], "fail_task_on_error": false, "depends_on": [{"task_step_id": "tests"}]}
]keep_alive_with_base_command starts the image's own command in the background and keeps the
container running, so the agent can be installed into it. Port 8000 is the A2A port this example
assumes the agent's install extension names. Preflight looks up only envs and run_code scripts,
so this task is created with nothing registered:
agent-env task create fix-tests.json --id fix-the-tests --project-id <project-id>Reading steps from fix-tests.json...
Step 1/6: deploy_sandbox id=box version=1
Step 2/6: run_docker_container id=repo version=1
Step 3/6: install_agent id=coder version=1
Step 4/6: prompt_agent id=fix version=1
Step 5/6: run_container_unit_tests_verifier id=tests version=1
Step 6/6: teardown_sandboxes id=teardown version=1
Creating task 'fix-the-tests' with 6 steps (project_id=<project-id>)...
Created task: id=fix-the-tests version=1 steps=6Env control
These steps act on an env instance that a deploy_env step started, found by env_id, and run
before the agent is prompted. apply_server_config calls extensions your MCP servers advertise on
their environment cards. modify_env_tool_access, sync_env_clock and register_env_triggers
call extensions the gateway advertises, so they need the
gateway topology, the default.
| Step | What it does | Key fields |
|---|---|---|
reset_env | Calls the env's reset() on its instance, to return it to a clean state. It is best-effort: fail_task_on_error defaults to false, and an env type that does not implement reset(), which none of the built-in types do, is logged and skipped. | env_id (required) |
apply_server_config | Invokes configuration extensions on the env's MCP servers. Each directive names a service (a server's environment_name, or "*" for every server in the env), the uri of an extension that server's card advertises, and the args to send. Applied directives are recorded in metadata["server_config_changes"]. | env_id, directives (required); tolerate_unadvertised (false), timeout_seconds (30) |
modify_env_tool_access | Disables or enables tools for one role on the gateway, the role an agent is deployed with in deploy_agent. The role's state afterwards is recorded in metadata["tool_access_changes"]. | env_id, action (disable or enable), role, tools (all required) |
sync_env_clock | Arms the gateway's virtual clock at virtual_time and sets its speed, then syncs every MCP server whose card advertises urn:agentenv:clock/v1; by default it skips the others. Place it right before prompt_agent, so virtual time starts when the agent does. | env_id, virtual_time (RFC 3339, required); virtual_seconds_per_real_second (1.0; 0 freezes the clock, 86400 is the most), tolerate_missing_sync_time (true) |
register_env_triggers | Registers triggers on the gateway. A trigger's condition is a tool call (action), a check on the state (state) or a virtual time (time); its actions are a plain-language instruction for an executor agent (nl), a tool call (tool), or tools enabled or disabled for a role (permission). | env_id, triggers (required); watch_roles, executor_agent_name, executor_timeout_seconds (120) |
The executor agent, the one that carries out nl actions, must be deployed with a role that is
not in watch_roles; otherwise its own tool calls would fire the triggers, and the step fails.
Two of these steps give reply-to-dana a deadline. The trigger disables send for the role
assistant at 17:00 on Friday, virtual time. The clock starts on Monday at 09:00 and runs 3,600
times faster than real time, so the week passes in under two minutes:
[
{"id": "deadline", "type": "register_env_triggers", "env_id": "email", "triggers": [{"id": "lock-send", "when": {"type": "time", "at": "2026-10-02T17:00:00Z"}, "actions": [{"type": "permission", "action": "disable", "role": "assistant", "tools": ["send"]}]}], "depends_on": [{"task_step_id": "email"}]},
{"id": "clock", "type": "sync_env_clock", "env_id": "email", "virtual_time": "2026-09-28T09:00:00Z", "virtual_seconds_per_real_second": 3600, "depends_on": [{"task_step_id": "inbox"}, {"task_step_id": "assistant"}, {"task_step_id": "deadline"}]}
]In this variant of the task, the assistant step also sets "role": "assistant", so the trigger
applies to its tool calls, and reply depends on clock instead of on inbox and assistant.
Agents
These steps add to an agent that a deploy_agent or install_agent step started, found by
agent_name, through extensions the agent's card advertises. build_mcp_cli is the exception: it
reads an env instance and makes an artifact for an agent to load.
| Step | What it does | Key fields |
|---|---|---|
add_skills | Registers skills on the agent. A skill is inline (name, description, body), a skill directory at an object store URL (s3_url), or a stored skill artifact (skill_artifact_id). The step can also write a skill that points the agent at a CLI or at files that load_artifact gave it, by artifact id. | agent_name (default-agent); skills, cli_artifact_ids, file_artifact_universe_ids |
build_mcp_cli | Generates a command-line tool for a deployed env's MCP tools, shaped by the interface manifest the server serves when it has one, and stores it as a CLI artifact, recorded in metadata["cli_artifact"]. It needs the gateway topology. load_artifact with agent_name and env_id installs the CLI for the agent. agent-env env mcp-server create-cli runs this step. | env_id, command_name (required); cli_artifact_id (cli-<env_id>) |
peer_agents | Sends each source agent a routing table of the agents it may message, so it can reach them over A2A. Every agent named must be deployed, and each source must advertise the peer-agents extension. | peerings (required): a list of {"source_agent_name": ..., "peer_agent_names": [...]} |
register_agent_triggers | Registers triggers on an agent that advertises the triggers extension. A condition of type env_trigger waits for a trigger that a register_env_triggers step registered, and the step checks that trigger exists before it posts. | agent_name, triggers (required) |
Capture
These steps save a run's state as new artifacts, which outlast the sandboxes and which a later
task can load. So does collect_artifacts, which
saves the files an agent made.
| Step | What it does | Key fields |
|---|---|---|
snapshot_env | Reads the state of every MCP server in a deployed env through its gateway, with the data-plane method data/get, and stores it as a new environment universe artifact. The id and version go to metadata["env_snapshotted_universes"][<step id>], where a later load_artifact in the same task finds them by artifact_from_step_id. | env_id, or env_instance_id or env_step_id to pick one deployment; snapshot_id, export_timeout_seconds (600), include_env_trajectory (false) |
snapshot_agent_state | Captures the agent's conversation for a prompt_id through its snapshot extension, as a file artifact universe under artifact_id, recorded in metadata["agent_snapshots"]. A later deploy_agent can start an agent from it. With env_id and universe_artifact_id, it also saves each MCP server's state beside it. | artifact_id, prompt_id (required); agent_name (default-agent), env_id, universe_artifact_id, timeout_seconds (120) |
Without snapshot_id, the universe's id is snapshot-<env_id>- followed by the run's
eight-character suffix, so a retried run adds a version rather than a second universe.
Code and grading
run_code runs your own script. The other five are verifiers, besides rubrics_verifier on the previous page:
each writes a score to metadata["verifications"][<verifier_id>], and all but
aggregate_verifiers write their result rows beside it.
| Step | What it does | Key fields |
|---|---|---|
run_code | Runs the run(input) function of a Python script artifact and stores the JSON it returns in metadata["script_results"][<result_id>]. input is {"args": ..., "results": ...}: the step's args and every earlier script result. With env_id it runs on the env's host; otherwise it runs in the agent's container, where it sees the agent's files. Preflight checks that the script artifact exists. | script_artifact_id (required); entrypoint (run), args, result_id (the step id), env_id, agent_name (default-agent), script_file, timeout_seconds (600) |
env_outcome_verifier | Grades the env's end state rather than the reply: it loads a Python file from a file artifact and awaits its async def verify(mcp_url) with the env instance's MCP URL, in the process that runs the task, with no model. The rows verify() returns are stored unchanged as results. The instance must still be running. | env_id, file_artifact_id (required); file_artifact_version, verifier_id, score_aggregator (all_pass) |
run_container_unit_tests_verifier | Runs a test command in a container that run_docker_container started, after any setup_commands. The score is 1.0 when it exits 0, or the number in the file at reward_path when that is set. stdout and stderr are stored as file artifacts. Without reward_path, a non-zero exit fails the step, after the result is recorded. | sandbox_name, container_name, command (required); verifier_id (the step id), setup_commands, timeout_sec (300), env_vars, result_paths, reward_path |
verify_sandbox | Checks an agent's sandbox, or one deploy_sandbox started, against criteria of four types: probe_file_exists, probe_dir_exists, probe_file_contains and bash_cmd_succeeds, relative to base_dir. It skips criteria of other types, so one list can serve both it and a rubrics_verifier. | criteria; agent_name or sandbox_name, base_dir (/app), score_aggregator (all_pass), verifier_id, shell_timeout_seconds (120) |
agent_prompt_response_verifier | Checks the reply stored under prompt_id, with no model: response_contains passes when every string in needles is in it, response_regex_present when pattern matches. A criterion of another type fails the step. | prompt_id (required); criteria, score_aggregator (all_pass), verifier_id |
aggregate_verifiers | Combines the result rows of earlier verifiers, named in verifier_ids, into one score, leaving out rows they skipped. | verifier_ids; score_aggregator (weighted_average), verifier_id |
score_aggregator is all_pass, any_pass or weighted_average. Where the table gives no default
for verifier_id, an unset one is a random id, so set it whenever something reads the result.
env_outcome_verifier
This verify() reads the sent folder through the data plane, which the instance serves at the same
base URL as /mcp, and passes when an email went to Dana:
from agentenv_protocol import DataPart, client
async def verify(mcp_url: str) -> list[dict]:
state = await client.get_data(mcp_url.removesuffix("/mcp"))
sent = [email for part in state.parts if isinstance(part, DataPart) for email in part.data.get("sent", [])]
replied = any(email["to"] == "dana@example.com" for email in sent)
return [{"id": "replied-to-dana", "description": "An email was sent to dana@example.com", "result": replied}]No command stores a single file as a file artifact, so check-reply is stored from Python, and the
check step of reply-to-dana runs it:
from agent_env.artifact import FileArtifact
FileArtifact.put(id="check-reply", description="Checks that Dana got a reply", file_path="verify.py"){"id": "check", "type": "env_outcome_verifier", "env_id": "email", "file_artifact_id": "check-reply", "verifier_id": "sent", "depends_on": [{"task_step_id": "reply"}]}Human in the loop
These two bring a person into a run: one waits for a decision, the other lets the agent ask.
| Step | What it does | Key fields |
|---|---|---|
review | Pauses the run until an operator records continue or abort for this step, for up to timeout_seconds. abort fails the step, so with the default fail_task_on_error the steps after it don't run; continue is recorded in metadata["review_decisions"]. | label, timeout_seconds (10800), poll_interval_seconds (5) |
deploy_human_agent | Adds a person to the run as an agent: an A2A endpoint at a2a_url, or at [conversations] default_human_a2a_url followed by /instance/<instance id>, added to deployed_agents. peer_agents can then make it a peer that the agent consults. A task can have one. | agent_name (human_agent), a2a_url |
Core ships no screen for a review's decision. Record it from Python with
get_review_store().set_decision(instance_id, step_id, "continue"), or "abort", imported from
agent_env.task_step.review_store. The wait holds no sandbox, so place it after the agent's output
is saved and before collect_artifacts.
Registration validators
The 20 validators check that an env or an agent works with the framework: that it serves its card,
answers the protocol, describes its tools, and handles what the framework sends it. The validate
commands build them into a task, together with steps such as deploy_env, deploy_agent and
prompt_agent, run it, and save the results to the env's or the agent's metadata:
| Command | The validators it runs |
|---|---|
agent-env env mcp-server validate email, or env mcp-server put --validate | verify_env_card, verify_env_core_protocol, verify_mcp_tool_schema, verify_spec_conformance, verify_mcp_env_assessment, validation_gate_aggregator |
agent-env env multi validate --id <id>, env multi put --validate, and env website put unless --skip-validation | verify_env_card |
agent-env env multi validate-universe-compatibility --env-id <id> --universe-artifact-id <id> | verify_universe_load_export_roundtrip, combine_universe_verdicts |
agent-env a2a-agent validate --id claude-code, and a2a-agent put unless --skip-validation | the twelve verify_a2a_* steps; verify_a2a_install runs in a second task, only when the agent's card advertises urn:agentenv:install/v1 |
Each run stores its task as a new version, for example validate-email-v1 for version 1 of email
and validate-a2a-claude-code-v1 for version 1 of claude-code.
| Step | What it checks |
|---|---|
verify_env_card | The environment card the instance serves, saved to the env's metadata |
verify_env_core_protocol | Which data-plane methods the env supports: data/reset, data/add, data/get |
verify_mcp_tool_schema | That every MCP tool has a description, and that its parameters are described |
verify_spec_conformance | The live tools against the OpenAPI spec the server serves; records drift, not a required gate |
verify_mcp_env_assessment | Tool correctness, from an agent's structured answer after it tried the tools |
validation_gate_aggregator | Rolls the verdicts into one validation_gate, which put --validate enforces |
verify_universe_load_export_roundtrip | Loads a universe into a multi env, exports it and compares, to find what a load loses |
combine_universe_verdicts | Merges that comparison with a judge agent's verdict into one compatibility result |
verify_a2a_agent_card | The structure of the agent's card |
verify_a2a_agent_config_identity | That the agent-config extension changes the card the agent serves |
verify_a2a_agent_mcp | That the agent calls MCP tools, from a test prompt and the env's trajectory |
verify_a2a_core_protocol | The A2A methods message/send and tasks/get |
verify_a2a_install | Installing the agent into a container through its install/v1 extension |
verify_a2a_modalities | Which inputs the agent handles: text, images, audio, PDF and video |
verify_a2a_peer_agents | The peer-agents extension: the agent relays a token from a peer |
verify_a2a_role | That a role set through agent-config reads back |
verify_a2a_skill_config | The skill-config extension, from a rubric on answers only the skills give |
verify_a2a_snapshot | The snapshot extension: a second agent recalls a token from the first one's snapshot |
verify_a2a_system_prompt | That a system_prompt set through agent-config reaches the model |
verify_a2a_trajectory | The ways of fetching a trajectory that the agent advertises |
Listing the registry
The registry is what a step's type is looked up in, and it can hold more than the built-ins. To
see what yours holds, read it from Python:
from agent_env.task_step.registry import get_task_step_registry
registry = get_task_step_registry()
print(len(registry))
print(registry["snapshot_env"])
print(sorted(registry)[:4])49
<class 'agent_env.task_step.task_steps.snapshot_env.SnapshotEnvTaskStep'>
['add_skills', 'agent_prompt_response_verifier', 'aggregate_verifiers', 'apply_server_config']get_task_step_registry() builds the registry from the configuration in effect, so run it in the
directory, or with the AGENT_ENV_CONFIG, that your tasks use.
Your own steps
A step type of your own, such as the ReplyMentionsStep that
Your own step writes, joins the same registry in one of two
ways: [task_steps] impls in config.toml names its class as module:Class, or an installed
package registers it under the agent_env.task_steps entry point, named after its type.
[task_steps]
impls = ["mycorp.steps:ReplyMentionsStep"]With that file in effect and mycorp importable, the same script counts one more:
50
<class 'agent_env.task_step.task_steps.snapshot_env.SnapshotEnvTaskStep'>
['add_skills', 'agent_prompt_response_verifier', 'aggregate_verifiers', 'apply_server_config']Every process that reads or runs a task using the step needs the same registration.
Last updated on