# More task steps (https://www.agentenvframework.com/docs/tasks/more-steps)

> Every built-in step besides the important six, grouped by family with its key fields, how to list the registry, and how your own step joins it

The step registry holds 49 built-in step types. Six of them make up most tasks, and
[Important Task Steps](https://www.agentenvframework.com/docs/tasks/important-steps.md) covers them: `deploy_env`,
`load_artifact`, `deploy_agent`, `prompt_agent`, `rubrics_verifier` and `collect_artifacts`. This
page lists the other 43 in seven families. The
last family, the 20 registration validators, is one you rarely put in a task yourself.

1. The registry holds 49 built-in step types, and a step’s `type` string picks its class.
2. Six steps make up most tasks. Important Task Steps covers them, one by one.
3. Provisioning starts a sandbox, builds a container on it, installs an agent inside, and stops sandboxes.
4. Env control resets a running env instance, or sets its config, tool access, clock and triggers.
5. Agent steps give a deployed agent more: skills, a CLI, peers to message, triggers of its own.
6. Capture steps save state as new artifacts: an env’s data, or an agent’s conversation.
7. Code and grading run your script or tests, check an env, a sandbox or a reply, combine scores.
8. Human in the loop pauses a run for an operator, or adds a person the agent can message.
9. 20 validators check an env or agent works with the framework; `validate` commands run them.
10. Your own step joins the registry through `[task_steps] impls` or an entry point: 50 types.

Parts of the scene:

- **The step registry**: Maps each step `type` to a Python class. It holds the 49 built-in steps, then classes from the `agent_env.task_steps` entry point, then `[task_steps] impls`. `get_task_step_registry()` returns it.
- **A lookup by type**: Creating or reading a task looks each step’s `type` up in the registry and builds the step with that class. A type the registry does not hold fails with `Unknown task step type`.
- **Important Task Steps**: The six steps nearly every task uses, from `deploy_env` to `collect_artifacts`. Important Task Steps covers each one, with its fields.
- **Provisioning**: Steps for tasks that need a machine rather than an env server: a sandbox, a container built on it, and an agent installed into it, which find each other by `sandbox_name` and `container_name`. `teardown_sandboxes` stops sandboxes when nothing needs them.
- **Env control**: Steps that act on an env instance that `deploy_env` started, found by `env_id`, before the agent is prompted. Tool access, the clock and triggers need the gateway topology.
- **Agents**: Steps that add to an agent that `deploy_agent` or `install_agent` started, found by `agent_name`, through extensions its card advertises. `build_mcp_cli` builds a CLI from an env for an agent to load.
- **Capture**: Steps that save a run’s state as new artifacts, so it outlasts the sandboxes and a later task can load it.
- **Code and grading**: A step that runs your script, and the verifiers besides `rubrics_verifier`, from `env_outcome_verifier` to `aggregate_verifiers`. Each verifier writes `metadata["verifications"][<verifier_id>]`.
- **Human in the loop**: A step that waits for a person’s decision, and one that adds a person to the run as an agent.
- **Registration validators**: Twenty steps that check an env or an agent works with the framework. `agent-env env mcp-server validate` and `agent-env a2a-agent validate` build them into a task and run it; you rarely write them yourself.
- **Your own steps**: Classes you register join the same registry, under their own `type`. A type that collides with a built-in is rejected.
- **reply_mentions**: A custom step, `mycorp.steps:ReplyMentionsStep`, named in `[task_steps] impls`. With it, the registry holds 50 types.
- **deploy_env**: Deploys a registered env as a new env instance and adds it to `deployed_envs`.
- **load_artifact**: Loads a stored artifact into an env instance through its data plane, or into an agent’s sandbox as files.
- **deploy_agent**: Deploys a registered agent in its own sandbox, gives it the MCP URLs of `env_ids`, and adds it to `deployed_agents`.
- **prompt_agent**: Prompts an agent by `agent_name` and stores its reply and trajectory under `prompt_id`.
- **rubrics_verifier**: Has a judge grade the reply under a `prompt_id` against criteria you write in plain language.
- **collect_artifacts**: Copies the files an agent made out of its sandbox into the object store, as file artifacts.
- **deploy_sandbox**: Starts a bare sandbox, a `vm` or a `container`, and adds it to `deployed_sandboxes` under `sandbox_name`.
- **run_docker_container**: Builds an image on a `vm` sandbox from a file artifact universe or a `.zip` URL, and runs it in the background.
- **install_agent**: Installs a registered agent into a running container, or onto the sandbox itself, with the install commands its card declares.
- **teardown_sandboxes**: Stops the sandboxes behind the envs, agents and sandboxes it names, best-effort. A run never does this on its own.
- **reset_env**: Calls the env’s `reset()` on its instance. Best-effort: the built-in env types don’t implement it, so on them it skips.
- **apply_server_config**: Invokes configuration extensions that the env’s MCP servers advertise on their cards, each with its own `args`.
- **modify_env_tool_access**: Disables or enables a list of tools for one `role` on the env’s gateway.
- **sync_env_clock**: Sets the gateway’s virtual clock and its speed, and syncs every MCP server that supports it. Place it right before `prompt_agent`.
- **register_env_triggers**: Registers triggers on the gateway: on a tool call, a state or a virtual time, run an instruction, a tool call or a permission change.
- **add_skills**: Registers skills on a deployed agent: inline, from an object store URL, from a skill artifact, or for a CLI it was given.
- **build_mcp_cli**: Generates a command-line tool for a deployed env’s MCP tools and stores it as a CLI artifact, which `load_artifact` installs for an agent.
- **peer_agents**: Pushes each source agent a routing table of peers it can message over A2A.
- **register_agent_triggers**: Registers triggers on an agent that advertises the triggers extension; they can wait on env triggers.
- **snapshot_env**: Reads each MCP server’s state through the gateway with `data/get`, and stores it as a new environment universe artifact.
- **snapshot_agent_state**: Captures an agent’s conversation for a `prompt_id` as a file artifact universe, which a later agent can start from.
- **run_code**: Runs a Python script artifact’s `run(input)` on the env’s host or in the agent’s container, and stores the JSON it returns.
- **env_outcome_verifier**: Runs your `verify()` function against an env instance’s MCP URL and scores the env’s end state, with no model.
- **run_container_unit_tests_verifier**: Runs a test command in a container and scores its exit code, or the reward file it writes.
- **verify_sandbox**: Checks files, directories, file contents and shell commands in an agent’s sandbox or a deployed one.
- **agent_prompt_response_verifier**: Checks a reply for strings or a regular expression, with no model.
- **aggregate_verifiers**: Combines the result rows of earlier verifiers into one score.
- **review**: Pauses the run until an operator records continue or abort for this step; abort stops the steps after it.
- **deploy_human_agent**: Adds a person’s A2A endpoint to `deployed_agents`, so `peer_agents` can make it a peer. One per task.
- **verify_env_card**: Fetches the environment card the instance serves and saves it to the env’s metadata.
- **verify_env_core_protocol**: Records which data-plane methods the env supports: `data/reset`, `data/add` and `data/get`.
- **verify_mcp_tool_schema**: Checks that every MCP tool has a description and that its parameters are described.
- **verify_spec_conformance**: Compares the live tools with the OpenAPI spec the server serves, and records the drift; it is not a required gate by default.
- **verify_mcp_env_assessment**: Records tool correctness from an agent’s structured answer after it tried the tools.
- **validation_gate_aggregator**: Rolls the env’s verdicts into one `validation_gate`, which `env mcp-server put --validate` enforces.
- **verify_universe_load_export_roundtrip**: Loads a universe into a `multi` env, exports it, and compares the two to find what the load loses.
- **combine_universe_verdicts**: Merges the round trip’s comparison with a judge agent’s verdict into one compatibility result.
- **verify_a2a_agent_card**: Checks the structure of the agent card the agent serves.
- **verify_a2a_agent_config_identity**: Checks that the agent-config extension changes the card the agent serves.
- **verify_a2a_agent_mcp**: Checks that the agent calls MCP tools, from a test prompt and the env’s trajectory.
- **verify_a2a_core_protocol**: Checks the core A2A methods, `message/send` and `tasks/get`.
- **verify_a2a_install**: Records whether installing the agent into a container through its `install/v1` extension works.
- **verify_a2a_modalities**: Grades which inputs the agent handles: text, images, audio, PDF and video.
- **verify_a2a_peer_agents**: Checks the peer-agents extension: the agent relays a token from a peer.
- **verify_a2a_role**: Checks that a `role` set through the agent-config extension reads back.
- **verify_a2a_skill_config**: Checks the skill-config extension, from a rubric on answers only the skills give.
- **verify_a2a_snapshot**: Checks the snapshot extension: a second agent recalls a token from the first one’s snapshot.
- **verify_a2a_system_prompt**: Checks that a `system_prompt` set through agent-config reaches the model.
- **verify_a2a_trajectory**: Checks the ways of fetching a trajectory that the agent advertises.

| Family                                              | Steps                                                                                                                                              |
| --------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| [Provisioning](#provisioning)                       | `deploy_sandbox`, `run_docker_container`, `install_agent`, `teardown_sandboxes`                                                                    |
| [Env control](#env-control)                         | `reset_env`, `apply_server_config`, `modify_env_tool_access`, `sync_env_clock`, `register_env_triggers`                                            |
| [Agents](#agents)                                   | `add_skills`, `build_mcp_cli`, `peer_agents`, `register_agent_triggers`                                                                            |
| [Capture](#capture)                                 | `snapshot_env`, `snapshot_agent_state`                                                                                                             |
| [Code and grading](#code-and-grading)               | `run_code`, `env_outcome_verifier`, `run_container_unit_tests_verifier`, `verify_sandbox`, `agent_prompt_response_verifier`, `aggregate_verifiers` |
| [Human in the loop](#human-in-the-loop)             | `review`, `deploy_human_agent`                                                                                                                     |
| [Registration validators](#registration-validators) | 20 validators, such as `verify_env_card` and `verify_a2a_agent_card`, that `validate` commands run                                                 |

Each table gives a step's key fields, not all of them: required fields first, then the useful
optional ones, with defaults in parentheses. Every step also takes `id`, `type`, `depends_on`,
`fail_task_on_error` and `retry_config`, which [What is a task step](https://www.agentenvframework.com/docs/tasks/task-steps.md)
covers.

## Provisioning

These steps build a machine for a task that needs one rather than an env, such as a coding task:
a bare sandbox, a container built on it, and an agent installed into the container. The last one
stops sandboxes, a machine's or an env's and an agent's.

| Step                   | What it does                                                                                                                                                                                                                                                                                                                                                                                                                                                                      | Key fields                                                                                                                                                                                                                                                                    |
| ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `deploy_sandbox`       | Starts a bare sandbox and adds it to `deployed_sandboxes` under `sandbox_name`. `sandbox_mode: "vm"` makes a machine that later steps build and run containers on; `"container"` runs a single `image` that listens on `port`, and fails without both.                                                                                                                                                                                                                            | `sandbox_name`, `sandbox_mode` (required); `image`, `port`, `exposed_ports`, `cpu` (2.0), `memory_mb` (8192), `disk_size_gb` (10), `ttl_seconds` (7200), `sandbox_type`, `env_vars`                                                                                           |
| `run_docker_container` | Builds an image on a `vm` sandbox and starts it in the background. The build context is a file artifact universe or a `.zip` at a URL, exactly one of the two. `ready_command` is retried inside the container for up to 120 seconds. The container is recorded in `metadata["deployed_docker_containers"]`.                                                                                                                                                                      | `sandbox_name`, and `docker_context_artifact_id` or `docker_context_url` (required); `container_name` (`task-container`), `dockerfile_path` (`Dockerfile`), `ports`, `env_vars`, `build_args`, `command_override`, `keep_alive_with_base_command`, `ready_command`, `volumes` |
| `install_agent`        | Installs a registered agent into a running container, or onto the `vm` sandbox itself when `container_name` is omitted, with the install commands its card declares in the `urn:agentenv:install/v1` extension. It waits for the agent's card and adds the agent to `deployed_agents` under `agent_name`.                                                                                                                                                                         | `sandbox_name`, `a2a_agent_id` (required); `container_name`, `agent_name` (`default-agent`), `a2a_port`, `workspace_dir`                                                                                                                                                      |
| `teardown_sandboxes`   | Terminates the sandboxes behind the envs, agents and sandboxes it names, to stop paying for them once nothing needs them. A run never does this on its own, and the `local` sandbox ignores its TTL. It is best-effort: a sandbox already gone or failing to stop is logged, and the run goes on. The ids go to `metadata["torn_down_sandbox_ids"]`; the `deployed_envs` and `deployed_agents` entries stay. `agent-env task create` rejects one with `fail_task_on_error: true`. | `env_ids`, `agent_names` or `sandbox_names` (at least one); `fail_task_on_error` (`false`, and must stay so)                                                                                                                                                                  |

`install_agent` reaches the agent through the sandbox's tunnel for its A2A port, so that port must
be in the sandbox's `exposed_ports` and, for a container, in its `ports` too. A `deploy_agent` step
with a `sandbox_name` does something else: it runs the agent in a container of its own on that
sandbox, which needs port 8000 in `exposed_ports`.

### A task without an env

This task runs a coding agent against a repository's tests. `box` is the machine, `repo` the
repository's container on it, and `coder` the agent installed into that container. The verifier,
`run_container_unit_tests_verifier`, is in [Code and grading](#code-and-grading):

```json title="fix-tests.json"
[
  {"id": "box", "type": "deploy_sandbox", "sandbox_name": "box", "sandbox_mode": "vm", "exposed_ports": [8000], "depends_on": []},
  {"id": "repo", "type": "run_docker_container", "sandbox_name": "box", "container_name": "repo", "docker_context_url": "https://example.com/repo.zip", "ports": [8000], "keep_alive_with_base_command": true, "depends_on": [{"task_step_id": "box"}]},
  {"id": "coder", "type": "install_agent", "sandbox_name": "box", "container_name": "repo", "a2a_agent_id": "claude-code", "agent_name": "coder", "depends_on": [{"task_step_id": "repo"}]},
  {"id": "fix", "type": "prompt_agent", "agent_name": "coder", "prompt_id": "fix", "prompt": "Make the failing tests pass.", "depends_on": [{"task_step_id": "coder"}]},
  {"id": "tests", "type": "run_container_unit_tests_verifier", "sandbox_name": "box", "container_name": "repo", "command": "pytest -q", "verifier_id": "tests", "depends_on": [{"task_step_id": "fix"}]},
  {"id": "teardown", "type": "teardown_sandboxes", "sandbox_names": ["box"], "fail_task_on_error": false, "depends_on": [{"task_step_id": "tests"}]}
]
```

`keep_alive_with_base_command` starts the image's own command in the background and keeps the
container running, so the agent can be installed into it. Port 8000 is the A2A port this example
assumes the agent's install extension names. Preflight looks up only envs and `run_code` scripts,
so this task is created with nothing registered:

```bash title="Terminal"
agent-env task create fix-tests.json --id fix-the-tests --project-id <project-id>
```

```text title="Output"
Reading steps from fix-tests.json...
  Step 1/6: deploy_sandbox id=box version=1
  Step 2/6: run_docker_container id=repo version=1
  Step 3/6: install_agent id=coder version=1
  Step 4/6: prompt_agent id=fix version=1
  Step 5/6: run_container_unit_tests_verifier id=tests version=1
  Step 6/6: teardown_sandboxes id=teardown version=1
Creating task 'fix-the-tests' with 6 steps (project_id=<project-id>)...
Created task: id=fix-the-tests version=1 steps=6
```

## Env control

These steps act on an env instance that a `deploy_env` step started, found by `env_id`, and run
before the agent is prompted. `apply_server_config` calls extensions your MCP servers advertise on
their environment cards. `modify_env_tool_access`, `sync_env_clock` and `register_env_triggers`
call extensions the gateway advertises, so they need the
[gateway topology](https://www.agentenvframework.com/docs/environments/topology.md#the-gateway-topology), the default.

| Step                     | What it does                                                                                                                                                                                                                                                                                                                | Key fields                                                                                                                                                            |
| ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `reset_env`              | Calls the env's `reset()` on its instance, to return it to a clean state. It is best-effort: `fail_task_on_error` defaults to `false`, and an env type that does not implement `reset()`, which none of the built-in types do, is logged and skipped.                                                                       | `env_id` (required)                                                                                                                                                   |
| `apply_server_config`    | Invokes configuration extensions on the env's MCP servers. Each directive names a `service` (a server's `environment_name`, or `"*"` for every server in the env), the `uri` of an extension that server's card advertises, and the `args` to send. Applied directives are recorded in `metadata["server_config_changes"]`. | `env_id`, `directives` (required); `tolerate_unadvertised` (`false`), `timeout_seconds` (30)                                                                          |
| `modify_env_tool_access` | Disables or enables tools for one role on the gateway, the `role` an agent is deployed with in `deploy_agent`. The role's state afterwards is recorded in `metadata["tool_access_changes"]`.                                                                                                                                | `env_id`, `action` (`disable` or `enable`), `role`, `tools` (all required)                                                                                            |
| `sync_env_clock`         | Arms the gateway's virtual clock at `virtual_time` and sets its speed, then syncs every MCP server whose card advertises `urn:agentenv:clock/v1`; by default it skips the others. Place it right before `prompt_agent`, so virtual time starts when the agent does.                                                         | `env_id`, `virtual_time` (RFC 3339, required); `virtual_seconds_per_real_second` (1.0; 0 freezes the clock, 86400 is the most), `tolerate_missing_sync_time` (`true`) |
| `register_env_triggers`  | Registers triggers on the gateway. A trigger's condition is a tool call (`action`), a check on the state (`state`) or a virtual time (`time`); its actions are a plain-language instruction for an executor agent (`nl`), a tool call (`tool`), or tools enabled or disabled for a role (`permission`).                     | `env_id`, `triggers` (required); `watch_roles`, `executor_agent_name`, `executor_timeout_seconds` (120)                                                               |

The executor agent, the one that carries out `nl` actions, must be deployed with a `role` that is
not in `watch_roles`; otherwise its own tool calls would fire the triggers, and the step fails.

Two of these steps give `reply-to-dana` a deadline. The trigger disables `send` for the role
`assistant` at 17:00 on Friday, virtual time. The clock starts on Monday at 09:00 and runs 3,600
times faster than real time, so the week passes in under two minutes:

```json title="Two more steps for reply-to-dana"
[
  {"id": "deadline", "type": "register_env_triggers", "env_id": "email", "triggers": [{"id": "lock-send", "when": {"type": "time", "at": "2026-10-02T17:00:00Z"}, "actions": [{"type": "permission", "action": "disable", "role": "assistant", "tools": ["send"]}]}], "depends_on": [{"task_step_id": "email"}]},
  {"id": "clock", "type": "sync_env_clock", "env_id": "email", "virtual_time": "2026-09-28T09:00:00Z", "virtual_seconds_per_real_second": 3600, "depends_on": [{"task_step_id": "inbox"}, {"task_step_id": "assistant"}, {"task_step_id": "deadline"}]}
]
```

In this variant of the task, the `assistant` step also sets `"role": "assistant"`, so the trigger
applies to its tool calls, and `reply` depends on `clock` instead of on `inbox` and `assistant`.

## Agents

These steps add to an agent that a `deploy_agent` or `install_agent` step started, found by
`agent_name`, through extensions the agent's card advertises. `build_mcp_cli` is the exception: it
reads an env instance and makes an artifact for an agent to load.

| Step                      | What it does                                                                                                                                                                                                                                                                                                                                                                   | Key fields                                                                                 |
| ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------ |
| `add_skills`              | Registers skills on the agent. A skill is inline (`name`, `description`, `body`), a skill directory at an object store URL (`s3_url`), or a stored skill artifact (`skill_artifact_id`). The step can also write a skill that points the agent at a CLI or at files that `load_artifact` gave it, by artifact id.                                                              | `agent_name` (`default-agent`); `skills`, `cli_artifact_ids`, `file_artifact_universe_ids` |
| `build_mcp_cli`           | Generates a command-line tool for a deployed env's MCP tools, shaped by the interface manifest the server serves when it has one, and stores it as a CLI artifact, recorded in `metadata["cli_artifact"]`. It needs the gateway topology. `load_artifact` with `agent_name` and `env_id` installs the CLI for the agent. `agent-env env mcp-server create-cli` runs this step. | `env_id`, `command_name` (required); `cli_artifact_id` (`cli-<env_id>`)                    |
| `peer_agents`             | Sends each source agent a routing table of the agents it may message, so it can reach them over A2A. Every agent named must be deployed, and each source must advertise the peer-agents extension.                                                                                                                                                                             | `peerings` (required): a list of `{"source_agent_name": ..., "peer_agent_names": [...]}`   |
| `register_agent_triggers` | Registers triggers on an agent that advertises the triggers extension. A condition of type `env_trigger` waits for a trigger that a `register_env_triggers` step registered, and the step checks that trigger exists before it posts.                                                                                                                                          | `agent_name`, `triggers` (required)                                                        |

## Capture

These steps save a run's state as new artifacts, which outlast the sandboxes and which a later
task can load. So does [`collect_artifacts`](https://www.agentenvframework.com/docs/tasks/important-steps.md#collect_artifacts), which
saves the files an agent made.

| Step                   | What it does                                                                                                                                                                                                                                                                                                                               | Key fields                                                                                                                                                |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `snapshot_env`         | Reads the state of every MCP server in a deployed env through its gateway, with the data-plane method `data/get`, and stores it as a new environment universe artifact. The id and version go to `metadata["env_snapshotted_universes"][<step id>]`, where a later `load_artifact` in the same task finds them by `artifact_from_step_id`. | `env_id`, or `env_instance_id` or `env_step_id` to pick one deployment; `snapshot_id`, `export_timeout_seconds` (600), `include_env_trajectory` (`false`) |
| `snapshot_agent_state` | Captures the agent's conversation for a `prompt_id` through its snapshot extension, as a file artifact universe under `artifact_id`, recorded in `metadata["agent_snapshots"]`. A later `deploy_agent` can start an agent from it. With `env_id` and `universe_artifact_id`, it also saves each MCP server's state beside it.              | `artifact_id`, `prompt_id` (required); `agent_name` (`default-agent`), `env_id`, `universe_artifact_id`, `timeout_seconds` (120)                          |

Without `snapshot_id`, the universe's id is `snapshot-<env_id>-` followed by the run's
eight-character suffix, so a retried run adds a version rather than a second universe.

## Code and grading

`run_code` runs your own script. The other five are verifiers, besides `rubrics_verifier` on the previous page:
each writes a `score` to `metadata["verifications"][<verifier_id>]`, and all but
`aggregate_verifiers` write their result rows beside it.

| Step                                | What it does                                                                                                                                                                                                                                                                                                                                                                                                   | Key fields                                                                                                                                                                 |
| ----------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `run_code`                          | Runs the `run(input)` function of a Python script artifact and stores the JSON it returns in `metadata["script_results"][<result_id>]`. `input` is `{"args": ..., "results": ...}`: the step's `args` and every earlier script result. With `env_id` it runs on the env's host; otherwise it runs in the agent's container, where it sees the agent's files. Preflight checks that the script artifact exists. | `script_artifact_id` (required); `entrypoint` (`run`), `args`, `result_id` (the step id), `env_id`, `agent_name` (`default-agent`), `script_file`, `timeout_seconds` (600) |
| `env_outcome_verifier`              | Grades the env's end state rather than the reply: it loads a Python file from a file artifact and awaits its `async def verify(mcp_url)` with the env instance's MCP URL, in the process that runs the task, with no model. The rows `verify()` returns are stored unchanged as `results`. The instance must still be running.                                                                                 | `env_id`, `file_artifact_id` (required); `file_artifact_version`, `verifier_id`, `score_aggregator` (`all_pass`)                                                           |
| `run_container_unit_tests_verifier` | Runs a test command in a container that `run_docker_container` started, after any `setup_commands`. The score is 1.0 when it exits 0, or the number in the file at `reward_path` when that is set. stdout and stderr are stored as file artifacts. Without `reward_path`, a non-zero exit fails the step, after the result is recorded.                                                                        | `sandbox_name`, `container_name`, `command` (required); `verifier_id` (the step id), `setup_commands`, `timeout_sec` (300), `env_vars`, `result_paths`, `reward_path`      |
| `verify_sandbox`                    | Checks an agent's sandbox, or one `deploy_sandbox` started, against criteria of four types: `probe_file_exists`, `probe_dir_exists`, `probe_file_contains` and `bash_cmd_succeeds`, relative to `base_dir`. It skips criteria of other types, so one list can serve both it and a `rubrics_verifier`.                                                                                                          | `criteria`; `agent_name` or `sandbox_name`, `base_dir` (`/app`), `score_aggregator` (`all_pass`), `verifier_id`, `shell_timeout_seconds` (120)                             |
| `agent_prompt_response_verifier`    | Checks the reply stored under `prompt_id`, with no model: `response_contains` passes when every string in `needles` is in it, `response_regex_present` when `pattern` matches. A criterion of another type fails the step.                                                                                                                                                                                     | `prompt_id` (required); `criteria`, `score_aggregator` (`all_pass`), `verifier_id`                                                                                         |
| `aggregate_verifiers`               | Combines the result rows of earlier verifiers, named in `verifier_ids`, into one score, leaving out rows they skipped.                                                                                                                                                                                                                                                                                         | `verifier_ids`; `score_aggregator` (`weighted_average`), `verifier_id`                                                                                                     |

`score_aggregator` is `all_pass`, `any_pass` or `weighted_average`. Where the table gives no default
for `verifier_id`, an unset one is a random id, so set it whenever something reads the result.

### `env_outcome_verifier`

This `verify()` reads the sent folder through the data plane, which the instance serves at the same
base URL as `/mcp`, and passes when an email went to Dana:

```python title="verify.py"
from agentenv_protocol import DataPart, client


async def verify(mcp_url: str) -> list[dict]:
    state = await client.get_data(mcp_url.removesuffix("/mcp"))
    sent = [email for part in state.parts if isinstance(part, DataPart) for email in part.data.get("sent", [])]
    replied = any(email["to"] == "dana@example.com" for email in sent)
    return [{"id": "replied-to-dana", "description": "An email was sent to dana@example.com", "result": replied}]
```

No command stores a single file as a file artifact, so `check-reply` is stored from Python, and the
`check` step of `reply-to-dana` runs it:

```python title="put_check.py"
from agent_env.artifact import FileArtifact

FileArtifact.put(id="check-reply", description="Checks that Dana got a reply", file_path="verify.py")
```

```json title="The check step"
{"id": "check", "type": "env_outcome_verifier", "env_id": "email", "file_artifact_id": "check-reply", "verifier_id": "sent", "depends_on": [{"task_step_id": "reply"}]}
```

## Human in the loop

These two bring a person into a run: one waits for a decision, the other lets the agent ask.

| Step                 | What it does                                                                                                                                                                                                                                                             | Key fields                                                      |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------- |
| `review`             | Pauses the run until an operator records `continue` or `abort` for this step, for up to `timeout_seconds`. `abort` fails the step, so with the default `fail_task_on_error` the steps after it don't run; `continue` is recorded in `metadata["review_decisions"]`.      | `label`, `timeout_seconds` (10800), `poll_interval_seconds` (5) |
| `deploy_human_agent` | Adds a person to the run as an agent: an A2A endpoint at `a2a_url`, or at `[conversations] default_human_a2a_url` followed by `/instance/<instance id>`, added to `deployed_agents`. `peer_agents` can then make it a peer that the agent consults. A task can have one. | `agent_name` (`human_agent`), `a2a_url`                         |

Core ships no screen for a review's decision. Record it from Python with
`get_review_store().set_decision(instance_id, step_id, "continue")`, or `"abort"`, imported from
`agent_env.task_step.review_store`. The wait holds no sandbox, so place it after the agent's output
is saved and before `collect_artifacts`.

## Registration validators

The 20 validators check that an env or an agent works with the framework: that it serves its card,
answers the protocol, describes its tools, and handles what the framework sends it. The validate
commands build them into a task, together with steps such as `deploy_env`, `deploy_agent` and
`prompt_agent`, run it, and save the results to the env's or the agent's metadata:

| Command                                                                                                                | The validators it runs                                                                                                                                        |
| ---------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `agent-env env mcp-server validate email`, or `env mcp-server put --validate`                                          | `verify_env_card`, `verify_env_core_protocol`, `verify_mcp_tool_schema`, `verify_spec_conformance`, `verify_mcp_env_assessment`, `validation_gate_aggregator` |
| `agent-env env multi validate --id <id>`, `env multi put --validate`, and `env website put` unless `--skip-validation` | `verify_env_card`                                                                                                                                             |
| `agent-env env multi validate-universe-compatibility --env-id <id> --universe-artifact-id <id>`                        | `verify_universe_load_export_roundtrip`, `combine_universe_verdicts`                                                                                          |
| `agent-env a2a-agent validate --id claude-code`, and `a2a-agent put` unless `--skip-validation`                        | the twelve `verify_a2a_*` steps; `verify_a2a_install` runs in a second task, only when the agent's card advertises `urn:agentenv:install/v1`                  |

Each run stores its task as a new version, for example `validate-email-v1` for version 1 of `email`
and `validate-a2a-claude-code-v1` for version 1 of `claude-code`.

| Step                                    | What it checks                                                                                |
| --------------------------------------- | --------------------------------------------------------------------------------------------- |
| `verify_env_card`                       | The environment card the instance serves, saved to the env's metadata                         |
| `verify_env_core_protocol`              | Which data-plane methods the env supports: `data/reset`, `data/add`, `data/get`               |
| `verify_mcp_tool_schema`                | That every MCP tool has a description, and that its parameters are described                  |
| `verify_spec_conformance`               | The live tools against the OpenAPI spec the server serves; records drift, not a required gate |
| `verify_mcp_env_assessment`             | Tool correctness, from an agent's structured answer after it tried the tools                  |
| `validation_gate_aggregator`            | Rolls the verdicts into one `validation_gate`, which `put --validate` enforces                |
| `verify_universe_load_export_roundtrip` | Loads a universe into a `multi` env, exports it and compares, to find what a load loses       |
| `combine_universe_verdicts`             | Merges that comparison with a judge agent's verdict into one compatibility result             |
| `verify_a2a_agent_card`                 | The structure of the agent's card                                                             |
| `verify_a2a_agent_config_identity`      | That the agent-config extension changes the card the agent serves                             |
| `verify_a2a_agent_mcp`                  | That the agent calls MCP tools, from a test prompt and the env's trajectory                   |
| `verify_a2a_core_protocol`              | The A2A methods `message/send` and `tasks/get`                                                |
| `verify_a2a_install`                    | Installing the agent into a container through its `install/v1` extension                      |
| `verify_a2a_modalities`                 | Which inputs the agent handles: text, images, audio, PDF and video                            |
| `verify_a2a_peer_agents`                | The peer-agents extension: the agent relays a token from a peer                               |
| `verify_a2a_role`                       | That a `role` set through agent-config reads back                                             |
| `verify_a2a_skill_config`               | The skill-config extension, from a rubric on answers only the skills give                     |
| `verify_a2a_snapshot`                   | The snapshot extension: a second agent recalls a token from the first one's snapshot          |
| `verify_a2a_system_prompt`              | That a `system_prompt` set through agent-config reaches the model                             |
| `verify_a2a_trajectory`                 | The ways of fetching a trajectory that the agent advertises                                   |

## Listing the registry

The registry is what a step's `type` is looked up in, and it can hold more than the built-ins. To
see what yours holds, read it from Python:

```python title="registry.py"
from agent_env.task_step.registry import get_task_step_registry

registry = get_task_step_registry()
print(len(registry))
print(registry["snapshot_env"])
print(sorted(registry)[:4])
```

```text title="Output"
49
<class 'agent_env.task_step.task_steps.snapshot_env.SnapshotEnvTaskStep'>
['add_skills', 'agent_prompt_response_verifier', 'aggregate_verifiers', 'apply_server_config']
```

`get_task_step_registry()` builds the registry from the configuration in effect, so run it in the
directory, or with the `AGENT_ENV_CONFIG`, that your tasks use.

## Your own steps

A step type of your own, such as the `ReplyMentionsStep` that
[Your own step](https://www.agentenvframework.com/docs/tasks/task-steps.md#your-own-step) writes, joins the same registry in one of two
ways: `[task_steps] impls` in `config.toml` names its class as `module:Class`, or an installed
package registers it under the `agent_env.task_steps` entry point, named after its `type`.

```toml title=".agentenv/config.toml"
[task_steps]
impls = ["mycorp.steps:ReplyMentionsStep"]
```

With that file in effect and `mycorp` importable, the same script counts one more:

```text title="Output"
50
<class 'agent_env.task_step.task_steps.snapshot_env.SnapshotEnvTaskStep'>
['add_skills', 'agent_prompt_response_verifier', 'aggregate_verifiers', 'apply_server_config']
```

Every process that reads or runs a task using the step needs the same registration.