Important Task Steps
The Task Steps we have found most valuable
Most tasks are built from the same six steps, used in the same order. A task deploys an env, loads its data and deploys an agent next to it, prompts the agent, grades the reply, and keeps what the agent made:
deploy_envstarts an env instance.load_artifactloads data into the instance, whiledeploy_agentstarts the agent and connects it to the instance.prompt_agentgives the agent its task.rubrics_verifierhas a judge grade the reply.collect_artifactscopies files out of the agent's sandbox.
Each step finds what an earlier step made in the run's context by a name: an instance by env_id,
an agent by agent_name, a reply by prompt_id. This page covers the six in that order, with
reply-to-dana, a rubric named tone in place of its check, and a collect step:
[
{"id": "email", "type": "deploy_env", "env_id": "email", "depends_on": []},
{"id": "inbox", "type": "load_artifact", "env_id": "email", "artifact_id": "inbox", "depends_on": [{"task_step_id": "email"}]},
{"id": "assistant", "type": "deploy_agent", "env_ids": ["email"], "a2a_agent_id": "claude-code", "agent_name": "assistant", "depends_on": [{"task_step_id": "email"}]},
{"id": "reply", "type": "prompt_agent", "agent_name": "assistant", "prompt_id": "reply", "prompt": "Reply to Dana.", "depends_on": [{"task_step_id": "inbox"}, {"task_step_id": "assistant"}]},
{"id": "tone", "type": "rubrics_verifier", "prompt_id": "reply", "verifier_id": "tone", "criteria": [{"id": "polite", "criterion": "The email to Dana is polite and signed.", "weight": 1}, {"id": "explains", "criterion": "The email tells Dana what was done with the Q3 numbers.", "weight": 1}], "depends_on": [{"task_step_id": "reply"}]},
{"id": "collect", "type": "collect_artifacts", "agent_name": "assistant", "depends_on": [{"task_step_id": "reply"}]}
]inbox and assistant run at the same time, as do tone and collect. Each section opens with a
run of its step, then lists the fields a task author sets, required first, then the useful optional
ones with their defaults, then names any fields it leaves out. The fields every step has, such as
depends_on and fail_task_on_error, are on What is a task step.
deploy_env
deploy_env deploys a registered env as a new env instance, as
agent-env env deploy does. Each run of the task deploys
its own instance.
| Field | Default | Meaning |
|---|---|---|
env_id | required | The env to deploy |
env_version | the latest | The version to deploy. A bare id resolves each time the task runs, so pin it when runs must be comparable |
ttl_seconds | 7200 | How long the instance may run. A remote sandbox stops when it runs out; the local sandbox does not enforce it |
sandbox_type | [sandbox] default | The sandbox to run in, such as local or modal_vm. agent-env task run --env-sandbox overrides it for one run |
disk_size_gb, cpu, memory_mb | 10, the sandbox's | The instance's disk, CPU and memory |
gateway_mode | performance | The gateway's mode, performance or consistent |
Omitted: env_state_type, env_state_instance_id, priority and metadata.
deploy_env reads nothing from the context. It adds a DeployedEnv to deployed_envs: the
env_id, the version it deployed, the instance's instance_id and its mcp_url. Every later step
that names the env by env_id finds the instance there. A run deploys an env once: a second
deploy_env with the same env_id fails with Env 'email' is already deployed.
{"id": "email", "type": "deploy_env", "env_id": "email", "depends_on": []}load_artifact
load_artifact loads a stored artifact into a running env instance or into an agent's sandbox, so
every run starts from the same data. An environment artifact loaded into an env goes through the
env's data plane: the step calls data/reset and then data/add with the artifact's file. Its
environment_name must match the env's, as The name must match shows.
A file artifact, or a file artifact universe, loaded into an agent is copied into its sandbox.
| Field | Default | Meaning |
|---|---|---|
artifact_id | one source is required | The artifact to load |
artifact_version | the latest | Its version |
artifacts | none | A list of {"id": ..., "version": ...} to load, instead of artifact_id |
artifact_from_step_id | none | Load the artifact an earlier collect_artifacts or snapshot_env step recorded, by that step's id |
urls | none | Files to download into an agent's or a sandbox's filesystem |
env_id | one target is required | The deployed env to load into |
agent_name | none | The deployed agent whose sandbox to copy files into |
destination_path | /tmp/file_artifacts | Where files land in an agent's sandbox; an environment artifact staged into an agent lands in /app/files |
Omitted: collected_artifacts_step_id, the older name of artifact_from_step_id, still accepted;
env_step_id, sandbox_name, container_name and snapshot_after_load.
load_artifact reads deployed_envs by env_id, or deployed_agents by agent_name. Loading an
environment artifact into an env writes nothing to the context: the data is in the instance. Files
copied into an agent are listed, with the agent's name and folder, in
metadata["loaded_file_artifact_universes"], which is how a judge learns what it can open.
{"id": "inbox", "type": "load_artifact", "env_id": "email", "artifact_id": "inbox", "depends_on": [{"task_step_id": "email"}]}deploy_agent
deploy_agent deploys a registered A2A agent in its own sandbox and gives it a name for the rest of
the run. For each env in env_ids, it posts the instance's MCP URL to the agent's MCP-config
extension, which is when the agent gets the env's tools, email_search and send. It then
registers the step's skills and sets its system_prompt, if it has them.
| Field | Default | Meaning |
|---|---|---|
a2a_agent_id | [agents] default_a2a_agent_id, a2a-default | The registered agent to deploy. agent-env task run --a2a-agent-id overrides it for one run |
a2a_agent_version | the latest | Its version |
agent_name | default-agent | The name later steps find the agent by |
env_ids | [] | The deployed envs whose MCP URL the agent gets |
system_prompt | none | A system prompt, sent when the agent's card accepts one. <key> placeholders are replaced from the run's seed |
role | none | A role, sent to the agent with its name; the env's triggers and tool access act on it |
skills | [] | Skills to register on the agent |
env_vars | {} | Environment variables for the agent's container |
sandbox_type | [sandbox] agent_default | The sandbox to run in. agent-env task run --agent-sandbox overrides it for one run |
ttl_seconds | 7200 | How long the sandbox may run; not enforced by local |
cpu, memory_mb, disk_size_gb | the sandbox's | Its resources |
Omitted: env_step_id, agent_description, litellm_base_url, priority, network_policy,
sandbox_name, enable_docker, metadata, and the snapshot and changelog fields that restore an
agent's earlier state.
deploy_agent reads deployed_envs by each id in env_ids before it deploys anything. It adds a
DeployedAgent to deployed_agents: the agent_name, the agent's A2A URL, its sandbox and its
instance_id. It also sets the run's default agent model from [model.roles] agent or
[model] default, unless an earlier step set it. A second agent with the same agent_name fails
with Agent with name 'assistant' is already deployed.
{"id": "assistant", "type": "deploy_agent", "env_ids": ["email"], "a2a_agent_id": "claude-code", "agent_name": "assistant", "depends_on": [{"task_step_id": "email"}]}prompt_agent
prompt_agent sends a prompt to a deployed agent over A2A and waits for it to finish. It stores
the agent's reply, where the agent's trajectory is stored, and the number of tool calls it made.
That is a single turn, the default. With max_conversation_turns above 1, it runs a conversation
instead: each reply goes to the user, and the user's answer comes back as the next turn. The user is
the agent named user_agent_name, which answers {message, done} and ends the conversation with
done: true, or a person at user_a2a_url, as Human agents shows,
whose text is the next prompt as it is. Every turn lands in the conversation store.
| Field | Default | Meaning |
|---|---|---|
prompt | this or parts is required | The prompt's text. <key> placeholders are replaced from the run's seed |
parts | none | A list of A2A parts, such as text and files, instead of prompt |
agent_name | default-agent | The deployed agent to prompt |
prompt_id | a random hex string | The key the reply is stored under. Set it whenever a verifier reads the reply |
timeout_seconds | 600 | How long to wait for the agent before the step fails |
model | none | The agent's model. agent-env task run --agent-model wins over it, and the run's default_agent_model, which deploy_agent sets, fills in for it |
output_format | none | A JSON schema for a structured reply |
max_conversation_turns | 1 | The most turns in a conversation; above 1, each reply goes to the user |
user_agent_name | human_agent | The deployed agent that plays the user in a conversation |
user_a2a_url | none | An A2A endpoint that plays the user instead, such as a person's. Without it or a user agent, a conversation goes to [conversations] default_human_a2a_url |
user_agent_timeout_seconds | 600 | How long the user has to answer each turn before the conversation closes |
Omitted: system_prompt, max_turns, model_params, effort, max_thinking_tokens, harness,
agentenv_tools, context_id, poll_interval_seconds, trajectory_output_prefix,
snapshot_config, and the other user_* fields of a conversation.
prompt_agent reads deployed_agents by agent_name, and fails with
Agent with name 'assistant' not found in context.deployed_agents when no step deployed it. It adds
a PromptResponse with its prompt_id to prompt_responses: the reply's text, the prompt, where
the trajectory is stored, the tool-call count, the model, and the error, if the agent failed. A
failed agent still leaves its PromptResponse, and then the step fails.
{"id": "reply", "type": "prompt_agent", "agent_name": "assistant", "prompt_id": "reply", "prompt": "Reply to Dana.", "depends_on": [{"task_step_id": "inbox"}, {"task_step_id": "assistant"}]}rubrics_verifier
rubrics_verifier grades a reply against criteria you write in plain language. A judge reads the
prompt, the reply and, by default, the trajectory, and marks each criterion passed or failed; the
step folds those rows into one score from 0 to 1. By default the judge is an agent: with no
agent_name, the step deploys one for the grading and stops it afterwards.
| Field | Default | Meaning |
|---|---|---|
prompt_id | required | The reply to grade |
criteria | required | A list of {"id": ..., "criterion": ..., "weight": ...}. The ids must be unique, non-empty strings |
verifier_id | a random hex string | The key the result is stored under |
score_aggregator | all_pass | How rows become the score: all_pass, any_pass or weighted_average, where a negative weight is a penalty |
agent_name | none | Grade with an agent a deploy_agent step deployed, instead of deploying a judge |
judge_a2a_agent_id | [agents] default_a2a_agent_id | The agent the step deploys as its judge when agent_name is not set |
use_agent_judge | true | false calls the judge model directly, with no agent and no sandbox |
use_trajectory | true | Give the judge the trajectory as well as the reply |
default_model | [model.roles] judge, else [model] default, else claude-sonnet-4-6 | The judge's model, resolved when the task is created and stored with the step |
output_format | rubric_binary | The rows' shape: rubric_binary, pass or fail per criterion, or rubric_partial |
judge_timeout_seconds | 1000 | How long the judge may take |
Omitted: trajectory_filter, grading_policy_prompt, context_facts_for_judge, effort,
max_thinking_tokens, judge_sandbox_type, judge_priority, default_model_api_base, and the
other values of output_format.
rubrics_verifier reads prompt_responses by prompt_id, and deployed_agents by agent_name
when it has one. It writes metadata["verifications"][<verifier_id>]: format; results, one row
per criterion with the judge's result, score and justification; the score; and where the
judge's own trajectory is stored. When the PromptResponse carries an error, the step skips the
judge and stores a score of 0 with one row, prompt_error.
{"id": "tone", "type": "rubrics_verifier", "prompt_id": "reply", "verifier_id": "tone", "criteria": [
{"id": "polite", "criterion": "The email to Dana is polite and signed.", "weight": 1},
{"id": "explains", "criterion": "The email tells Dana what was done with the Q3 numbers.", "weight": 1}
], "depends_on": [{"task_step_id": "reply"}]}With all_pass, tone scores 1.0 only when the judge passes both criteria. A judge you name in
agent_name can be any registered agent, and its prompt lists the files a load_artifact step
copied into it, as the collect_artifacts example shows.
collect_artifacts
collect_artifacts copies files out of a deployed agent's sandbox into the object store, registers
each as a file artifact, and puts them together in one file artifact universe, so they outlast the
sandbox. With no artifact_paths, it collects every file under base_path, and an empty folder
collects nothing without failing. When you name files and none of them can be read, the step fails.
| Field | Default | Meaning |
|---|---|---|
agent_name | default-agent | The deployed agent whose sandbox to read |
artifact_paths | [], every file under base_path | The files to collect, relative to base_path or absolute |
base_path | /app/artifact | The folder relative paths are read from |
artifacts_key | expected_artifacts | A seed key whose comma-separated file list replaces artifact_paths for the run |
manifest_step_id | none | Take the files from the artifacts that a prompt_agent step's structured reply declared, by that step's id |
exclude_basenames | [] | File names to skip when collecting a whole folder |
universe_id_suffix | none | Added to the universe's id, so two collect steps in one run make two universes |
Omitted: env_id, sandbox_name and container_name, which collect from a desktop env's
controller, a sandbox's host or a container instead of an agent.
collect_artifacts reads deployed_agents by agent_name. It writes
metadata["collected_artifacts"][<step id>]: each file's object URL under artifacts, and the
{id, version} of the universe under file_artifact_universe. A load_artifact step whose
artifact_from_step_id is the collect step's id loads that universe, which is how one agent hands
its files to another.
In chart-q3, the first example on Core concepts, an analyst agent
saves a chart as /app/artifact/q3.png. These three steps collect it, copy it into the sandbox of a
judge agent deployed beside the analyst, and have that judge grade the analyst's reply:
[
{"id": "collect", "type": "collect_artifacts", "agent_name": "analyst", "artifact_paths": ["q3.png"], "depends_on": [{"task_step_id": "chart"}]},
{"id": "handoff", "type": "load_artifact", "agent_name": "judge", "artifact_from_step_id": "collect", "depends_on": [{"task_step_id": "collect"}, {"task_step_id": "judge"}]},
{"id": "grade", "type": "rubrics_verifier", "agent_name": "judge", "prompt_id": "chart", "verifier_id": "chart", "criteria": [{"id": "bars", "criterion": "q3.png is a bar chart with one bar per month of Q3.", "weight": 1}], "depends_on": [{"task_step_id": "handoff"}]}
]handoff copies q3.png into the judge's /tmp/file_artifacts, and grade's prompt tells the
judge the file is there, so it checks the chart itself rather than the analyst's description of it.
These six steps make up most tasks. More task steps covers the other
built-in steps, from env_outcome_verifier, which checks an env's end state with your code, and
teardown_sandboxes, which stops what a run deployed, to the steps that set an env's clock and
triggers.
Last updated on