# Important Task Steps (https://www.agentenvframework.com/docs/tasks/important-steps)

> The Task Steps we have found most valuable

Most tasks are built from the same six steps, used in the same order. A task deploys an env, loads
its data and deploys an agent next to it, prompts the agent, grades the reply, and keeps what the
agent made:

1. [`deploy_env`](#deploy_env) starts an env instance.
2. [`load_artifact`](#load_artifact) loads data into the instance, while
   [`deploy_agent`](#deploy_agent) starts the agent and connects it to the instance.
3. [`prompt_agent`](#prompt_agent) gives the agent its task.
4. [`rubrics_verifier`](#rubrics_verifier) has a judge grade the reply.
5. [`collect_artifacts`](#collect_artifacts) copies files out of the agent's sandbox.

Each step finds what an earlier step made in the run's context by a name: an instance by `env_id`,
an agent by `agent_name`, a reply by `prompt_id`. This page covers the six in that order, with
`reply-to-dana`, a rubric named `tone` in place of its `check`, and a `collect` step:

```json title="task.json"
[
  {"id": "email", "type": "deploy_env", "env_id": "email", "depends_on": []},
  {"id": "inbox", "type": "load_artifact", "env_id": "email", "artifact_id": "inbox", "depends_on": [{"task_step_id": "email"}]},
  {"id": "assistant", "type": "deploy_agent", "env_ids": ["email"], "a2a_agent_id": "claude-code", "agent_name": "assistant", "depends_on": [{"task_step_id": "email"}]},
  {"id": "reply", "type": "prompt_agent", "agent_name": "assistant", "prompt_id": "reply", "prompt": "Reply to Dana.", "depends_on": [{"task_step_id": "inbox"}, {"task_step_id": "assistant"}]},
  {"id": "tone", "type": "rubrics_verifier", "prompt_id": "reply", "verifier_id": "tone", "criteria": [{"id": "polite", "criterion": "The email to Dana is polite and signed.", "weight": 1}, {"id": "explains", "criterion": "The email tells Dana what was done with the Q3 numbers.", "weight": 1}], "depends_on": [{"task_step_id": "reply"}]},
  {"id": "collect", "type": "collect_artifacts", "agent_name": "assistant", "depends_on": [{"task_step_id": "reply"}]}
]
```

`inbox` and `assistant` run at the same time, as do `tone` and `collect`. Each section opens with a
run of its step, then lists the fields a task author sets, required first, then the useful optional
ones with their defaults, then names any fields it leaves out. The fields every step has, such as
`depends_on` and `fail_task_on_error`, are on [What is a task step](https://www.agentenvframework.com/docs/tasks/task-steps.md).

## `deploy_env`

1. `email` checks that no step deployed `email` yet, then reads the latest version of the env from the document store.
2. The sandbox provider of `sandbox_type`, `local` Docker here, starts a sandbox for the instance, with the gateway’s port open.
3. Into the sandbox go a docker-compose file and three containers: the `email` MCP server, its state store and the gateway.
4. The step polls the gateway’s `/.well-known/agent-env.json` until it answers, for up to 300 seconds.
5. The instance gets an id, `email-` and eight characters, and a record in the document store’s `env_instances`.
6. A `DeployedEnv` goes into `deployed_envs`, and every later step that names `email` finds the instance, and its MCP URL, there.

Parts of the scene:

- **env_id**: Required. The registered env to deploy, and the name later steps find the instance by. A second `deploy_env` of `email` in one run fails with `Env 'email' is already deployed`.
- **env_version**: Omitted, the latest version when the step runs, so a new version put between two runs changes what the next one deploys. Pin it when runs must be comparable.
- **sandbox_type**: Where the instance runs. Omitted, it is `[sandbox] default`, `local` Docker unless you change it; `agent-env task run --env-sandbox` overrides it for one run.
- **ttl_seconds**: How long the instance may run: 7200 seconds unless set. A remote sandbox stops when it runs out; the `local` sandbox does not enforce it.
- **The email step**: The `deploy_env` step with the id `email`. It awaits the env’s own `deploy()`, which picks the environment provider, the gateway by default, and the sandbox provider.
- **The document store**: Holds the env’s versions in `envs` and one record per deployed instance in `env_instances`. SQLite unless `[stores.document]` names another.
- **The sandbox**: One machine for the whole instance. The local provider runs it in Docker; Modal and E2B run it remotely. The gateway’s port, 18765, is the one the instance is reached on.
- **The gateway**: The instance’s front door: it serves the env card at `/.well-known/agent-env.json` and the MCP endpoint, and passes each tool call to the MCP server behind it.
- **TaskStepContext**: The run’s ledger, one per run. `deploy_env` reads `deployed_envs` to refuse a second deploy of the same env, and adds the new instance to it.
- **deployed_envs**: One `DeployedEnv` per env instance the run deployed. `load_artifact` and `deploy_agent` find `email` here by `env_id`.
- **instance_id**: The instance’s id: the env id and eight random characters. `agent-env env get-instance` and `Env.from_instance_id` find the instance by it.
- **mcp_url**: The gateway’s MCP endpoint, at the path the env card names. `deploy_agent` gives it to the agent, and verifiers call it.

`deploy_env` deploys a registered env as a new env instance, as
[`agent-env env deploy`](https://www.agentenvframework.com/docs/environments/deploying.md#deploy-it) does. Each run of the task deploys
its own instance.

| Field                              | Default             | Meaning                                                                                                           |
| ---------------------------------- | ------------------- | ----------------------------------------------------------------------------------------------------------------- |
| `env_id`                           | required            | The env to deploy                                                                                                 |
| `env_version`                      | the latest          | The version to deploy. A bare id resolves each time the task runs, so pin it when runs must be comparable         |
| `ttl_seconds`                      | `7200`              | How long the instance may run. A remote sandbox stops when it runs out; the `local` sandbox does not enforce it   |
| `sandbox_type`                     | `[sandbox] default` | The sandbox to run in, such as `local` or `modal_vm`. `agent-env task run --env-sandbox` overrides it for one run |
| `disk_size_gb`, `cpu`, `memory_mb` | `10`, the sandbox's | The instance's disk, CPU and memory                                                                               |
| `gateway_mode`                     | `performance`       | The gateway's mode, `performance` or `consistent`                                                                 |

Omitted: `env_state_type`, `env_state_instance_id`, `priority` and `metadata`.

`deploy_env` reads nothing from the context. It adds a `DeployedEnv` to `deployed_envs`: the
`env_id`, the version it deployed, the instance's `instance_id` and its `mcp_url`. Every later step
that names the env by `env_id` finds the instance there. A run deploys an env once: a second
`deploy_env` with the same `env_id` fails with `Env 'email' is already deployed`.

```json title="The email step"
{"id": "email", "type": "deploy_env", "env_id": "email", "depends_on": []}
```

## `load_artifact`

1. `inbox` finds the `email` instance in `deployed_envs` by `env_id`, and in its env card the URL of its data plane.
2. It reads the environment artifact `inbox` from the document store and checks that its `environment_name` is the env’s, `email`.
3. The artifact’s file lives in the object store. The step copies it into the `email` server’s container, as `/data/inbox.json`.
4. Over the instance’s data plane, `data/reset` empties the server: whatever an earlier load or agent left is gone.
5. `data/add` hands the server the file, and the server loads it: Dana’s email is in the inbox, for every run alike.
6. The data is in the instance, so the step adds nothing to the context. `reply` can start once `assistant` is up too.

Parts of the scene:

- **env_id**: The deployed env to load into, found in `deployed_envs`. The step fails with `Env 'email' not found in context.deployed_envs` when no step deployed it; `agent_name` loads files into an agent’s sandbox instead.
- **artifact_id**: The artifact to load, here the environment artifact `inbox`. `artifacts` takes a list instead, and `artifact_from_step_id` an artifact an earlier step made.
- **artifact_version**: Omitted, the latest version of the artifact when the step runs. Pin it to load the same data in every run.
- **The inbox step**: The `load_artifact` step with the id `inbox`. For an environment artifact loaded into an env, it goes through the env’s data plane, which every agent-env server serves.
- **The document store**: Holds every artifact’s versions in `artifacts`. The environment artifact names the file artifact that holds its data, and the step reads both.
- **The object store**: Holds the artifact’s bytes, here `inbox.json`. The filesystem store unless `[stores.object]` names another.
- **The email instance**: The instance `email` deployed. Its gateway serves the data plane, where `data/reset` and `data/add` arrive as JSON-RPC calls, and the `email` server holds the data.
- **TaskStepContext**: The run’s ledger. `load_artifact` reads the instance from `deployed_envs`. Loading into an env writes nothing; files copied into an agent are listed in `metadata["loaded_file_artifact_universes"]`.
- **deployed_envs**: Where `email` put its `DeployedEnv`. The step finds the instance here by `env_id`; the env card stored with it gives the URL of the instance’s data plane.

`load_artifact` loads a stored artifact into a running env instance or into an agent's sandbox, so
every run starts from the same data. An environment artifact loaded into an env goes through the
env's data plane: the step calls `data/reset` and then `data/add` with the artifact's file. Its
`environment_name` must match the env's, as [The name must match](https://www.agentenvframework.com/docs/artifacts/with-environments.md#the-name-must-match) shows.
A file artifact, or a file artifact universe, loaded into an agent is copied into its sandbox.

| Field                   | Default                | Meaning                                                                                                    |
| ----------------------- | ---------------------- | ---------------------------------------------------------------------------------------------------------- |
| `artifact_id`           | one source is required | The artifact to load                                                                                       |
| `artifact_version`      | the latest             | Its version                                                                                                |
| `artifacts`             | none                   | A list of `{"id": ..., "version": ...}` to load, instead of `artifact_id`                                  |
| `artifact_from_step_id` | none                   | Load the artifact an earlier `collect_artifacts` or `snapshot_env` step recorded, by that step's id        |
| `urls`                  | none                   | Files to download into an agent's or a sandbox's filesystem                                                |
| `env_id`                | one target is required | The deployed env to load into                                                                              |
| `agent_name`            | none                   | The deployed agent whose sandbox to copy files into                                                        |
| `destination_path`      | `/tmp/file_artifacts`  | Where files land in an agent's sandbox; an environment artifact staged into an agent lands in `/app/files` |

Omitted: `collected_artifacts_step_id`, the older name of `artifact_from_step_id`, still accepted;
`env_step_id`, `sandbox_name`, `container_name` and `snapshot_after_load`.

`load_artifact` reads `deployed_envs` by `env_id`, or `deployed_agents` by `agent_name`. Loading an
environment artifact into an env writes nothing to the context: the data is in the instance. Files
copied into an agent are listed, with the agent's name and folder, in
`metadata["loaded_file_artifact_universes"]`, which is how a judge learns what it can open.

```json title="The inbox step"
{"id": "inbox", "type": "load_artifact", "env_id": "email", "artifact_id": "inbox", "depends_on": [{"task_step_id": "email"}]}
```

## `deploy_agent`

1. `assistant` finds each env in `env_ids` in `deployed_envs` before it deploys anything, and checks no agent is named `assistant` yet.
2. It reads the registered agent `claude-code` from the document store, and with it the image the agent runs.
3. The agent sandbox provider pulls the image and runs it on port 8000, with the run’s model key in `LITELLM_API_KEY`.
4. The step polls `/.well-known/agent.json` until the agent serves its card, for up to 300 seconds. The card lists its extensions.
5. For each env, it posts the instance’s MCP URL to the agent’s `/ext/mcp-config`. Now the agent has `email_search` and `send`.
6. A `DeployedAgent` named `assistant` goes into `deployed_agents`, and the run’s `default_agent_model` is set from `[model.roles] agent`.

Parts of the scene:

- **env_ids**: The deployed envs whose MCP URL the agent gets, each found in `deployed_envs`. They are all resolved before a sandbox starts, so a missing one fails the step early.
- **a2a_agent_id**: The registered agent to deploy, here `claude-code`. Omitted, it is `[agents] default_a2a_agent_id`, `a2a-default` unless you set it; `agent-env task run --a2a-agent-id` overrides it for one run.
- **agent_name**: The name later steps find the agent by: `prompt_agent`, `collect_artifacts` and a judging `rubrics_verifier`. A second agent with the same name fails with `Agent with name 'assistant' is already deployed`.
- **sandbox_type**: Where the agent runs. Omitted, it is `[sandbox] agent_default`, `local` Docker unless you change it; `agent-env task run --agent-sandbox` overrides it for one run.
- **ttl_seconds**: How long the agent’s sandbox may run: 7200 seconds unless set. A remote sandbox stops when it runs out; the `local` sandbox does not enforce it.
- **The assistant step**: The `deploy_agent` step with the id `assistant`. It runs at the same time as `inbox`: both wait only for `email`.
- **The document store**: Holds the registered agents’ versions in `a2a_agents`, and one record per deployed agent in `a2a_agent_instances`.
- **The image store**: Holds the agent’s Docker image. The local sandbox pulls it from here; an OCI registry at `localhost:5000` unless `[stores.image]` names another.
- **The agent’s sandbox**: Its own sandbox, apart from the env’s. The agent serves A2A on port 8000, and the step reaches it through the sandbox’s URL for that port.
- **The agent card**: Served at `/.well-known/agent.json`. Its extensions say what the step can set: `urn:agentenv:mcp-config/v1` for envs, skill-config for `skills`, agent-config for `system_prompt`.
- **/ext/mcp-config**: The MCP-config extension. The step posts `{"url": <mcp_url>, "name": "email"}` for each env in `env_ids`, and the agent connects to the instance’s MCP server and lists its tools.
- **TaskStepContext**: The run’s ledger. `deploy_agent` reads the envs and the run’s model key, and adds the agent. `prompt_agent` finds it there by `agent_name`.
- **deployed_envs**: Where `email` put its `DeployedEnv`, with the MCP URL the agent gets.
- **deployed_agents**: One `DeployedAgent` per agent: its `agent_name`, A2A URL, sandbox, card and `instance_id`, which is the agent id and eight random characters.
- **user_overrides**: The run’s options. `litellm_api_key` goes into the agent’s `LITELLM_API_KEY`; `a2a_agent_id` and `agent_sandbox` win over the step’s own fields.
- **default_agent_model**: Set from `[model.roles] agent`, or `[model] default`, unless an earlier step set it. A `prompt_agent` step without a `model` asks the agent for this one.

`deploy_agent` deploys a registered A2A agent in its own sandbox and gives it a name for the rest of
the run. For each env in `env_ids`, it posts the instance's MCP URL to the agent's MCP-config
extension, which is when the agent gets the env's tools, `email_search` and `send`. It then
registers the step's `skills` and sets its `system_prompt`, if it has them.

| Field                              | Default                                        | Meaning                                                                                                        |
| ---------------------------------- | ---------------------------------------------- | -------------------------------------------------------------------------------------------------------------- |
| `a2a_agent_id`                     | `[agents] default_a2a_agent_id`, `a2a-default` | The registered agent to deploy. `agent-env task run --a2a-agent-id` overrides it for one run                   |
| `a2a_agent_version`                | the latest                                     | Its version                                                                                                    |
| `agent_name`                       | `default-agent`                                | The name later steps find the agent by                                                                         |
| `env_ids`                          | `[]`                                           | The deployed envs whose MCP URL the agent gets                                                                 |
| `system_prompt`                    | none                                           | A system prompt, sent when the agent's card accepts one. `<key>` placeholders are replaced from the run's seed |
| `role`                             | none                                           | A role, sent to the agent with its name; the env's triggers and tool access act on it                          |
| `skills`                           | `[]`                                           | Skills to register on the agent                                                                                |
| `env_vars`                         | `{}`                                           | Environment variables for the agent's container                                                                |
| `sandbox_type`                     | `[sandbox] agent_default`                      | The sandbox to run in. `agent-env task run --agent-sandbox` overrides it for one run                           |
| `ttl_seconds`                      | `7200`                                         | How long the sandbox may run; not enforced by `local`                                                          |
| `cpu`, `memory_mb`, `disk_size_gb` | the sandbox's                                  | Its resources                                                                                                  |

Omitted: `env_step_id`, `agent_description`, `litellm_base_url`, `priority`, `network_policy`,
`sandbox_name`, `enable_docker`, `metadata`, and the snapshot and changelog fields that restore an
agent's earlier state.

`deploy_agent` reads `deployed_envs` by each id in `env_ids` before it deploys anything. It adds a
`DeployedAgent` to `deployed_agents`: the `agent_name`, the agent's A2A URL, its sandbox and its
`instance_id`. It also sets the run's default agent model from `[model.roles] agent` or
`[model] default`, unless an earlier step set it. A second agent with the same `agent_name` fails
with `Agent with name 'assistant' is already deployed`.

```json title="The assistant step"
{"id": "assistant", "type": "deploy_agent", "env_ids": ["email"], "a2a_agent_id": "claude-code", "agent_name": "assistant", "depends_on": [{"task_step_id": "email"}]}
```

## `prompt_agent`

**Single-turn**

1. `reply` finds `assistant` in `deployed_agents` by `agent_name`, and takes the run’s `default_agent_model`, as the step sets no `model`.
2. It posts the model, and the other settings the agent’s card accepts, to the agent’s `/ext/agent-config`.
3. `message/send` gives the agent “Reply to Dana.” over A2A, and the agent answers at once with a task id.
4. The agent calls the env’s tools: `email_search` finds Dana’s email, `send` replies. The step polls `tasks/get` every 2 seconds.
5. The A2A task completes. Its message carries the reply’s text, and a data part with the tool-call count.
6. The agent’s `/ext/trajectory` hands over every turn and tool call, stored in the object store under the `prompt_id`.
7. A `PromptResponse` with `prompt_id: reply` goes into `prompt_responses`: the reply, the trajectory’s URI, the count and the model.

**Multi-turn**

1. `reply` finds `assistant` by `agent_name`, and `user` by `user_agent_name`: an agent a `deploy_agent` step deployed to play the user.
2. It posts the model to `assistant`’s agent-config, and an `output_format` to `user`’s: every answer is `{message, done}`.
3. The step opens a conversation record and sends “Reply to Dana.” to `assistant`, which drafts the email and asks before sending it.
4. `assistant`’s reply goes to `user`, which answers `{message: "Send it, and cc Sam.", done: false}`: that message is the next prompt.
5. `assistant` sends the email with Sam in cc and says so. That reply goes to `user` too.
6. `user` answers with `done: true`, so the conversation closes after 2 of its 4 turns. Every turn is in the conversation record.
7. The `PromptResponse` holds the last reply and a trajectory per turn; `metadata.a2a_conversations.reply` names the conversation record.

**Human in the loop**

1. `reply` finds `assistant`. The user is no agent of the run but `user_a2a_url`: a person’s A2A endpoint, such as the one on Human agents.
2. The step opens a conversation record and sends “Reply to Dana.” to `assistant`, which drafts the email and asks before sending it.
3. `assistant`’s reply goes to the person as an A2A message, and the step waits up to `user_agent_timeout_seconds`, 600, for the answer.
4. The person types “Mention the chart.” Their task completes, and their text, as it is, is `assistant`’s next prompt.
5. `assistant` sends the email and says so. The person enters an empty line, so their task fails and the conversation closes.
6. The `PromptResponse` holds the last reply and a trajectory per turn; `metadata.a2a_conversations.reply` names the conversation record.

Parts of the scene:

- **agent_name**: The deployed agent to prompt, found in `deployed_agents`. The step fails with `Agent with name 'assistant' not found in context.deployed_agents` when no step deployed it.
- **prompt_id**: The key the reply is stored under, and the one a verifier asks for. Omitted, it is a random hex string, so set it whenever a verifier reads the reply.
- **prompt**: The text the agent gets. `<key>` placeholders in it are replaced from the run’s seed; `parts` sends a list of A2A parts, such as text and files, instead.
- **timeout_seconds**: How long to wait for the agent to finish: 600 seconds unless set. After that, the step fails with a `TimeoutError`.
- **model**: Unset here, so the agent gets the run’s `default_agent_model`. A step’s `model` wins over that, and `agent-env task run --agent-model` over both.
- **The reply step**: The `prompt_agent` step with the id `reply`. It waits for `inbox` and `assistant`, so the agent starts with Dana’s email in the inbox and the env’s tools in hand.
- **The assistant agent**: The agent `assistant` deployed, reached over A2A at its sandbox’s URL. It runs the task as its harness does, here Claude Code with the env’s MCP tools.
- **The email instance**: The env the agent was given. The agent’s tool calls go to the instance’s MCP URL, and the instance’s state changes with them: the reply lands in the sent folder.
- **The object store**: Holds the trajectory, under `prompt_agent_trajectories/prompt_id=reply/`, one file per turn. A judge reads it from here.
- **TaskStepContext**: The run’s ledger. `prompt_agent` reads the agent and the model, and adds the reply. `rubrics_verifier` finds it there by `prompt_id`.
- **deployed_agents**: Where `assistant` put its `DeployedAgent`, with the A2A URL the step sends to.
- **default_agent_model**: The model `deploy_agent` set for the run. The agent is asked for it because neither the step nor `--agent-model` names one.
- **prompt_responses**: One `PromptResponse` per prompt: the reply’s text, the prompt, the trajectory’s URI, the tool-call count, the model, and the error if the agent failed. A failed agent still leaves one, and then the step fails.
- **max_conversation_turns**: `1` unless set: one prompt, one reply. Above 1, the step runs a conversation, with each reply going to the user and the user’s answer coming back as the next turn, for up to this many turns.
- **user_agent_name**: The deployed agent that plays the user: `human_agent` unless set. The step asks it for `{message, done}` answers, and `done: true` ends the conversation early.
- **user_a2a_url**: An A2A endpoint that plays the user, such as a person answering from a terminal. Its text is the next prompt as it is. Without it or a user agent, a conversation goes to `[conversations] default_human_a2a_url`.
- **user_agent_timeout_seconds**: How long the user has to answer each turn: 600 seconds unless set. A user that runs out of it closes the conversation, and the step keeps the agent’s last reply.
- **The user agent**: An agent that plays the user, deployed by a `deploy_agent` step as `user`, with a system prompt saying who it is. It gets the step’s `output_format`, and its `message` is the next prompt.
- **The person**: A person behind an A2A endpoint: an SDK agent whose `run()` waits for what they type. To the step it is an agent like any other, only slower.
- **The conversation record**: One record per conversation in the document store’s `agent_env_a2a_conversations`, with every turn, the agent’s and the user’s. `agent-env up` shows it in the run’s Conversation tab.
- **a2a_conversations**: The id of each step’s conversation record, keyed by the step’s id, here `reply`.

`prompt_agent` sends a prompt to a deployed agent over A2A and waits for it to finish. It stores
the agent's reply, where the agent's trajectory is stored, and the number of tool calls it made.
That is a single turn, the default. With `max_conversation_turns` above 1, it runs a conversation
instead: each reply goes to the user, and the user's answer comes back as the next turn. The user is
the agent named `user_agent_name`, which answers `{message, done}` and ends the conversation with
`done: true`, or a person at `user_a2a_url`, as [Human agents](https://www.agentenvframework.com/docs/agents/human-agents.md) shows,
whose text is the next prompt as it is. Every turn lands in the conversation store.

| Field                        | Default                     | Meaning                                                                                                                                                                                     |
| ---------------------------- | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `prompt`                     | this or `parts` is required | The prompt's text. `<key>` placeholders are replaced from the run's seed                                                                                                                    |
| `parts`                      | none                        | A list of A2A parts, such as text and files, instead of `prompt`                                                                                                                            |
| `agent_name`                 | `default-agent`             | The deployed agent to prompt                                                                                                                                                                |
| `prompt_id`                  | a random hex string         | The key the reply is stored under. Set it whenever a verifier reads the reply                                                                                                               |
| `timeout_seconds`            | `600`                       | How long to wait for the agent before the step fails                                                                                                                                        |
| `model`                      | none                        | The agent's model. [`agent-env task run --agent-model`](https://www.agentenvframework.com/docs/tasks/running.md#task-run-options) wins over it, and the run's `default_agent_model`, which `deploy_agent` sets, fills in for it |
| `output_format`              | none                        | A JSON schema for a structured reply                                                                                                                                                        |
| `max_conversation_turns`     | `1`                         | The most turns in a conversation; above 1, each reply goes to the user                                                                                                                      |
| `user_agent_name`            | `human_agent`               | The deployed agent that plays the user in a conversation                                                                                                                                    |
| `user_a2a_url`               | none                        | An A2A endpoint that plays the user instead, such as a person's. Without it or a user agent, a conversation goes to `[conversations] default_human_a2a_url`                                 |
| `user_agent_timeout_seconds` | `600`                       | How long the user has to answer each turn before the conversation closes                                                                                                                    |

Omitted: `system_prompt`, `max_turns`, `model_params`, `effort`, `max_thinking_tokens`, `harness`,
`agentenv_tools`, `context_id`, `poll_interval_seconds`, `trajectory_output_prefix`,
`snapshot_config`, and the other `user_*` fields of a conversation.

`prompt_agent` reads `deployed_agents` by `agent_name`, and fails with
`Agent with name 'assistant' not found in context.deployed_agents` when no step deployed it. It adds
a `PromptResponse` with its `prompt_id` to `prompt_responses`: the reply's text, the prompt, where
the trajectory is stored, the tool-call count, the model, and the error, if the agent failed. A
failed agent still leaves its `PromptResponse`, and then the step fails.

```json title="The reply step"
{"id": "reply", "type": "prompt_agent", "agent_name": "assistant", "prompt_id": "reply", "prompt": "Reply to Dana.", "depends_on": [{"task_step_id": "inbox"}, {"task_step_id": "assistant"}]}
```

## `rubrics_verifier`

1. `tone` finds the `PromptResponse` under `prompt_id: reply`: the prompt, the reply and where the trajectory is stored.
2. With no `agent_name`, it deploys the default agent as its judge, in a sandbox of its own, and copies the trajectory into it.
3. The judge gets the criteria, as `c1` and `c2`, the prompt and the reply in one message, and the path of the trajectory.
4. The judge answers with a score and a justification per criterion, which the step maps back to `polite` and `explains`.
5. The judge’s sandbox is stopped. `all_pass` makes the rows one score: 1.0, as both criteria passed.
6. The rows, the score and the judge’s trajectory go into `metadata.verifications.tone`, the key `verifier_id` names.

Parts of the scene:

- **prompt_id**: Required. The reply to grade, found in `prompt_responses`. When that `PromptResponse` carries an error, the step skips the judge and stores a score of 0 with one row, `prompt_error`.
- **verifier_id**: The key the result is stored under, `verifications.tone`. Omitted, it is a random hex string, so set it to find the score later.
- **criteria**: Required. A list of `{id, criterion, weight}` in plain language, with unique ids. A negative weight marks a penalty, which only `weighted_average` counts.
- **score_aggregator**: `all_pass` unless set: 1.0 when every row passed, else 0. `any_pass` needs one; `weighted_average` averages the rows’ scores by weight, less the penalties.
- **use_agent_judge**: `true` unless set: an agent grades, and can open the trajectory and files. `false` calls the judge model directly, with no agent and no sandbox.
- **The tone step**: The `rubrics_verifier` step with the id `tone`. It waits only for `reply`, so it runs at the same time as `collect`.
- **The judge**: An agent deployed for the grading only, the one `judge_a2a_agent_id` names or `[agents] default_a2a_agent_id`, and stopped afterwards. With `agent_name`, a deployed agent grades instead.
- **The object store**: Holds the trajectory `reply` stored. The step copies it into the judge’s sandbox, so the judge can check what the agent did, not only what it said.
- **TaskStepContext**: The run’s ledger. `rubrics_verifier` reads the reply from `prompt_responses` and writes its result under `metadata.verifications`.
- **prompt_responses**: Where `reply` put its `PromptResponse`: the prompt, the reply and the trajectory’s URI, which is what the judge reads.
- **verifications.tone**: What the step wrote: `format`, `results` with one row per criterion and its `result`, `score` and `justification`, the `score`, and where the judge’s trajectory is stored.

`rubrics_verifier` grades a reply against criteria you write in plain language. A judge reads the
prompt, the reply and, by default, the trajectory, and marks each criterion passed or failed; the
step folds those rows into one score from 0 to 1. By default the judge is an agent: with no
`agent_name`, the step deploys one for the grading and stops it afterwards.

| Field                   | Default                                                                 | Meaning                                                                                                       |
| ----------------------- | ----------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| `prompt_id`             | required                                                                | The reply to grade                                                                                            |
| `criteria`              | required                                                                | A list of `{"id": ..., "criterion": ..., "weight": ...}`. The ids must be unique, non-empty strings           |
| `verifier_id`           | a random hex string                                                     | The key the result is stored under                                                                            |
| `score_aggregator`      | `all_pass`                                                              | How rows become the score: `all_pass`, `any_pass` or `weighted_average`, where a negative weight is a penalty |
| `agent_name`            | none                                                                    | Grade with an agent a `deploy_agent` step deployed, instead of deploying a judge                              |
| `judge_a2a_agent_id`    | `[agents] default_a2a_agent_id`                                         | The agent the step deploys as its judge when `agent_name` is not set                                          |
| `use_agent_judge`       | `true`                                                                  | `false` calls the judge model directly, with no agent and no sandbox                                          |
| `use_trajectory`        | `true`                                                                  | Give the judge the trajectory as well as the reply                                                            |
| `default_model`         | `[model.roles] judge`, else `[model] default`, else `claude-sonnet-4-6` | The judge's model, resolved when the task is created and stored with the step                                 |
| `output_format`         | `rubric_binary`                                                         | The rows' shape: `rubric_binary`, pass or fail per criterion, or `rubric_partial`                             |
| `judge_timeout_seconds` | `1000`                                                                  | How long the judge may take                                                                                   |

Omitted: `trajectory_filter`, `grading_policy_prompt`, `context_facts_for_judge`, `effort`,
`max_thinking_tokens`, `judge_sandbox_type`, `judge_priority`, `default_model_api_base`, and the
other values of `output_format`.

`rubrics_verifier` reads `prompt_responses` by `prompt_id`, and `deployed_agents` by `agent_name`
when it has one. It writes `metadata["verifications"][<verifier_id>]`: `format`; `results`, one row
per criterion with the judge's `result`, `score` and `justification`; the `score`; and where the
judge's own trajectory is stored. When the `PromptResponse` carries an error, the step skips the
judge and stores a score of 0 with one row, `prompt_error`.

```json title="The tone step"
{"id": "tone", "type": "rubrics_verifier", "prompt_id": "reply", "verifier_id": "tone", "criteria": [
  {"id": "polite", "criterion": "The email to Dana is polite and signed.", "weight": 1},
  {"id": "explains", "criterion": "The email tells Dana what was done with the Q3 numbers.", "weight": 1}
], "depends_on": [{"task_step_id": "reply"}]}
```

With `all_pass`, `tone` scores 1.0 only when the judge passes both criteria. A judge you name in
`agent_name` can be any registered agent, and its prompt lists the files a `load_artifact` step
copied into it, as the [`collect_artifacts`](#collect_artifacts) example shows.

## `collect_artifacts`

1. `collect` finds `analyst` in `deployed_agents` by `agent_name`, and checks that its sandbox still answers.
2. `q3.png` is relative, so it is read from `base_path`: `/app/artifact/q3.png`. The step takes its size first.
3. The file comes out of the sandbox as base64 and is checked against that size, so a truncated copy is never uploaded.
4. It goes to the object store under `collected_artifacts`, the run’s instance id, and a version that is the time in seconds.
5. It is registered as a file artifact, and a file artifact universe named after the run holds it: both are in the document store.
6. `metadata.collected_artifacts.collect` records the file’s URL and the universe’s id and version, which a later step loads.

Parts of the scene:

- **agent_name**: The deployed agent whose sandbox to read, found in `deployed_agents`. Omitted, it is `default-agent`. A sandbox past its TTL fails the step: it is gone.
- **artifact_paths**: The files to collect, relative to `base_path` or absolute. Left empty, the step collects every file under `base_path`; named files of which none can be read fail it.
- **base_path**: `/app/artifact` unless set: the folder relative paths are read from, where the agent has to write the files it wants kept.
- **The collect step**: chart-q3’s `collect_artifacts` step with the id `collect`. It waits for `chart`, the prompt in which the analyst drew the chart.
- **The analyst’s sandbox**: Where the analyst agent ran and saved `q3.png`. It is gone when its sandbox stops, which is why the step copies it out.
- **The object store**: Holds the file’s bytes, under `artifacts/collected_artifacts/<run instance id>/<version>/q3.png`.
- **The document store**: Holds the new file artifact, `<run instance id>-q3.png`, and the file artifact universe that lists it, named after the run.
- **TaskStepContext**: The run’s ledger. `collect_artifacts` reads the agent and the run’s instance id, and records what it saved under `metadata.collected_artifacts`.
- **deployed_agents**: Where `analyst`, and `judge`, put their `DeployedAgent`s. The step reads the analyst’s sandbox from its entry.
- **instance_id**: The run’s task instance id. The universe is named after it, so each run of the task saves its files apart.
- **collected_artifacts.collect**: Under the step’s id: `artifacts`, each file’s object URL, and `file_artifact_universe`, the universe’s `{id, version}`. A `load_artifact` with `artifact_from_step_id: collect` loads that universe.

`collect_artifacts` copies files out of a deployed agent's sandbox into the object store, registers
each as a file artifact, and puts them together in one file artifact universe, so they outlast the
sandbox. With no `artifact_paths`, it collects every file under `base_path`, and an empty folder
collects nothing without failing. When you name files and none of them can be read, the step fails.

| Field                | Default                            | Meaning                                                                                                       |
| -------------------- | ---------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| `agent_name`         | `default-agent`                    | The deployed agent whose sandbox to read                                                                      |
| `artifact_paths`     | `[]`, every file under `base_path` | The files to collect, relative to `base_path` or absolute                                                     |
| `base_path`          | `/app/artifact`                    | The folder relative paths are read from                                                                       |
| `artifacts_key`      | `expected_artifacts`               | A seed key whose comma-separated file list replaces `artifact_paths` for the run                              |
| `manifest_step_id`   | none                               | Take the files from the `artifacts` that a `prompt_agent` step's structured reply declared, by that step's id |
| `exclude_basenames`  | `[]`                               | File names to skip when collecting a whole folder                                                             |
| `universe_id_suffix` | none                               | Added to the universe's id, so two collect steps in one run make two universes                                |

Omitted: `env_id`, `sandbox_name` and `container_name`, which collect from a desktop env's
controller, a sandbox's host or a container instead of an agent.

`collect_artifacts` reads `deployed_agents` by `agent_name`. It writes
`metadata["collected_artifacts"][<step id>]`: each file's object URL under `artifacts`, and the
`{id, version}` of the universe under `file_artifact_universe`. A `load_artifact` step whose
`artifact_from_step_id` is the collect step's id loads that universe, which is how one agent hands
its files to another.

In `chart-q3`, the first example on [Core concepts](https://www.agentenvframework.com/docs/core-concepts.md#tasks), an analyst agent
saves a chart as `/app/artifact/q3.png`. These three steps collect it, copy it into the sandbox of a
judge agent deployed beside the analyst, and have that judge grade the analyst's reply:

```json title="task.json (three steps of chart-q3)"
[
  {"id": "collect", "type": "collect_artifacts", "agent_name": "analyst", "artifact_paths": ["q3.png"], "depends_on": [{"task_step_id": "chart"}]},
  {"id": "handoff", "type": "load_artifact", "agent_name": "judge", "artifact_from_step_id": "collect", "depends_on": [{"task_step_id": "collect"}, {"task_step_id": "judge"}]},
  {"id": "grade", "type": "rubrics_verifier", "agent_name": "judge", "prompt_id": "chart", "verifier_id": "chart", "criteria": [{"id": "bars", "criterion": "q3.png is a bar chart with one bar per month of Q3.", "weight": 1}], "depends_on": [{"task_step_id": "handoff"}]}
]
```

`handoff` copies `q3.png` into the judge's `/tmp/file_artifacts`, and `grade`'s prompt tells the
judge the file is there, so it checks the chart itself rather than the analyst's description of it.

These six steps make up most tasks. [More task steps](https://www.agentenvframework.com/docs/tasks/more-steps.md) covers the other
built-in steps, from `env_outcome_verifier`, which checks an env's end state with your code, and
`teardown_sandboxes`, which stops what a run deployed, to the steps that set an env's clock and
triggers.