Skip to content
AgentEnv Framework

Important Task Steps

The Task Steps we have found most valuable

Most tasks are built from the same six steps, used in the same order. A task deploys an env, loads its data and deploys an agent next to it, prompts the agent, grades the reply, and keeps what the agent made:

  1. deploy_env starts an env instance.
  2. load_artifact loads data into the instance, while deploy_agent starts the agent and connects it to the instance.
  3. prompt_agent gives the agent its task.
  4. rubrics_verifier has a judge grade the reply.
  5. collect_artifacts copies files out of the agent's sandbox.

Each step finds what an earlier step made in the run's context by a name: an instance by env_id, an agent by agent_name, a reply by prompt_id. This page covers the six in that order, with reply-to-dana, a rubric named tone in place of its check, and a collect step:

task.json
[
  {"id": "email", "type": "deploy_env", "env_id": "email", "depends_on": []},
  {"id": "inbox", "type": "load_artifact", "env_id": "email", "artifact_id": "inbox", "depends_on": [{"task_step_id": "email"}]},
  {"id": "assistant", "type": "deploy_agent", "env_ids": ["email"], "a2a_agent_id": "claude-code", "agent_name": "assistant", "depends_on": [{"task_step_id": "email"}]},
  {"id": "reply", "type": "prompt_agent", "agent_name": "assistant", "prompt_id": "reply", "prompt": "Reply to Dana.", "depends_on": [{"task_step_id": "inbox"}, {"task_step_id": "assistant"}]},
  {"id": "tone", "type": "rubrics_verifier", "prompt_id": "reply", "verifier_id": "tone", "criteria": [{"id": "polite", "criterion": "The email to Dana is polite and signed.", "weight": 1}, {"id": "explains", "criterion": "The email tells Dana what was done with the Q3 numbers.", "weight": 1}], "depends_on": [{"task_step_id": "reply"}]},
  {"id": "collect", "type": "collect_artifacts", "agent_name": "assistant", "depends_on": [{"task_step_id": "reply"}]}
]

inbox and assistant run at the same time, as do tone and collect. Each section opens with a run of its step, then lists the fields a task author sets, required first, then the useful optional ones with their defaults, then names any fields it leaves out. The fields every step has, such as depends_on and fail_task_on_error, are on What is a task step.

deploy_env

deploy_env deploys a registered env as a new env instance, as agent-env env deploy does. Each run of the task deploys its own instance.

FieldDefaultMeaning
env_idrequiredThe env to deploy
env_versionthe latestThe version to deploy. A bare id resolves each time the task runs, so pin it when runs must be comparable
ttl_seconds7200How long the instance may run. A remote sandbox stops when it runs out; the local sandbox does not enforce it
sandbox_type[sandbox] defaultThe sandbox to run in, such as local or modal_vm. agent-env task run --env-sandbox overrides it for one run
disk_size_gb, cpu, memory_mb10, the sandbox'sThe instance's disk, CPU and memory
gateway_modeperformanceThe gateway's mode, performance or consistent

Omitted: env_state_type, env_state_instance_id, priority and metadata.

deploy_env reads nothing from the context. It adds a DeployedEnv to deployed_envs: the env_id, the version it deployed, the instance's instance_id and its mcp_url. Every later step that names the env by env_id finds the instance there. A run deploys an env once: a second deploy_env with the same env_id fails with Env 'email' is already deployed.

The email step
{"id": "email", "type": "deploy_env", "env_id": "email", "depends_on": []}

load_artifact

load_artifact loads a stored artifact into a running env instance or into an agent's sandbox, so every run starts from the same data. An environment artifact loaded into an env goes through the env's data plane: the step calls data/reset and then data/add with the artifact's file. Its environment_name must match the env's, as The name must match shows. A file artifact, or a file artifact universe, loaded into an agent is copied into its sandbox.

FieldDefaultMeaning
artifact_idone source is requiredThe artifact to load
artifact_versionthe latestIts version
artifactsnoneA list of {"id": ..., "version": ...} to load, instead of artifact_id
artifact_from_step_idnoneLoad the artifact an earlier collect_artifacts or snapshot_env step recorded, by that step's id
urlsnoneFiles to download into an agent's or a sandbox's filesystem
env_idone target is requiredThe deployed env to load into
agent_namenoneThe deployed agent whose sandbox to copy files into
destination_path/tmp/file_artifactsWhere files land in an agent's sandbox; an environment artifact staged into an agent lands in /app/files

Omitted: collected_artifacts_step_id, the older name of artifact_from_step_id, still accepted; env_step_id, sandbox_name, container_name and snapshot_after_load.

load_artifact reads deployed_envs by env_id, or deployed_agents by agent_name. Loading an environment artifact into an env writes nothing to the context: the data is in the instance. Files copied into an agent are listed, with the agent's name and folder, in metadata["loaded_file_artifact_universes"], which is how a judge learns what it can open.

The inbox step
{"id": "inbox", "type": "load_artifact", "env_id": "email", "artifact_id": "inbox", "depends_on": [{"task_step_id": "email"}]}

deploy_agent

deploy_agent deploys a registered A2A agent in its own sandbox and gives it a name for the rest of the run. For each env in env_ids, it posts the instance's MCP URL to the agent's MCP-config extension, which is when the agent gets the env's tools, email_search and send. It then registers the step's skills and sets its system_prompt, if it has them.

FieldDefaultMeaning
a2a_agent_id[agents] default_a2a_agent_id, a2a-defaultThe registered agent to deploy. agent-env task run --a2a-agent-id overrides it for one run
a2a_agent_versionthe latestIts version
agent_namedefault-agentThe name later steps find the agent by
env_ids[]The deployed envs whose MCP URL the agent gets
system_promptnoneA system prompt, sent when the agent's card accepts one. <key> placeholders are replaced from the run's seed
rolenoneA role, sent to the agent with its name; the env's triggers and tool access act on it
skills[]Skills to register on the agent
env_vars{}Environment variables for the agent's container
sandbox_type[sandbox] agent_defaultThe sandbox to run in. agent-env task run --agent-sandbox overrides it for one run
ttl_seconds7200How long the sandbox may run; not enforced by local
cpu, memory_mb, disk_size_gbthe sandbox'sIts resources

Omitted: env_step_id, agent_description, litellm_base_url, priority, network_policy, sandbox_name, enable_docker, metadata, and the snapshot and changelog fields that restore an agent's earlier state.

deploy_agent reads deployed_envs by each id in env_ids before it deploys anything. It adds a DeployedAgent to deployed_agents: the agent_name, the agent's A2A URL, its sandbox and its instance_id. It also sets the run's default agent model from [model.roles] agent or [model] default, unless an earlier step set it. A second agent with the same agent_name fails with Agent with name 'assistant' is already deployed.

The assistant step
{"id": "assistant", "type": "deploy_agent", "env_ids": ["email"], "a2a_agent_id": "claude-code", "agent_name": "assistant", "depends_on": [{"task_step_id": "email"}]}

prompt_agent

prompt_agent sends a prompt to a deployed agent over A2A and waits for it to finish. It stores the agent's reply, where the agent's trajectory is stored, and the number of tool calls it made. That is a single turn, the default. With max_conversation_turns above 1, it runs a conversation instead: each reply goes to the user, and the user's answer comes back as the next turn. The user is the agent named user_agent_name, which answers {message, done} and ends the conversation with done: true, or a person at user_a2a_url, as Human agents shows, whose text is the next prompt as it is. Every turn lands in the conversation store.

FieldDefaultMeaning
promptthis or parts is requiredThe prompt's text. <key> placeholders are replaced from the run's seed
partsnoneA list of A2A parts, such as text and files, instead of prompt
agent_namedefault-agentThe deployed agent to prompt
prompt_ida random hex stringThe key the reply is stored under. Set it whenever a verifier reads the reply
timeout_seconds600How long to wait for the agent before the step fails
modelnoneThe agent's model. agent-env task run --agent-model wins over it, and the run's default_agent_model, which deploy_agent sets, fills in for it
output_formatnoneA JSON schema for a structured reply
max_conversation_turns1The most turns in a conversation; above 1, each reply goes to the user
user_agent_namehuman_agentThe deployed agent that plays the user in a conversation
user_a2a_urlnoneAn A2A endpoint that plays the user instead, such as a person's. Without it or a user agent, a conversation goes to [conversations] default_human_a2a_url
user_agent_timeout_seconds600How long the user has to answer each turn before the conversation closes

Omitted: system_prompt, max_turns, model_params, effort, max_thinking_tokens, harness, agentenv_tools, context_id, poll_interval_seconds, trajectory_output_prefix, snapshot_config, and the other user_* fields of a conversation.

prompt_agent reads deployed_agents by agent_name, and fails with Agent with name 'assistant' not found in context.deployed_agents when no step deployed it. It adds a PromptResponse with its prompt_id to prompt_responses: the reply's text, the prompt, where the trajectory is stored, the tool-call count, the model, and the error, if the agent failed. A failed agent still leaves its PromptResponse, and then the step fails.

The reply step
{"id": "reply", "type": "prompt_agent", "agent_name": "assistant", "prompt_id": "reply", "prompt": "Reply to Dana.", "depends_on": [{"task_step_id": "inbox"}, {"task_step_id": "assistant"}]}

rubrics_verifier

rubrics_verifier grades a reply against criteria you write in plain language. A judge reads the prompt, the reply and, by default, the trajectory, and marks each criterion passed or failed; the step folds those rows into one score from 0 to 1. By default the judge is an agent: with no agent_name, the step deploys one for the grading and stops it afterwards.

FieldDefaultMeaning
prompt_idrequiredThe reply to grade
criteriarequiredA list of {"id": ..., "criterion": ..., "weight": ...}. The ids must be unique, non-empty strings
verifier_ida random hex stringThe key the result is stored under
score_aggregatorall_passHow rows become the score: all_pass, any_pass or weighted_average, where a negative weight is a penalty
agent_namenoneGrade with an agent a deploy_agent step deployed, instead of deploying a judge
judge_a2a_agent_id[agents] default_a2a_agent_idThe agent the step deploys as its judge when agent_name is not set
use_agent_judgetruefalse calls the judge model directly, with no agent and no sandbox
use_trajectorytrueGive the judge the trajectory as well as the reply
default_model[model.roles] judge, else [model] default, else claude-sonnet-4-6The judge's model, resolved when the task is created and stored with the step
output_formatrubric_binaryThe rows' shape: rubric_binary, pass or fail per criterion, or rubric_partial
judge_timeout_seconds1000How long the judge may take

Omitted: trajectory_filter, grading_policy_prompt, context_facts_for_judge, effort, max_thinking_tokens, judge_sandbox_type, judge_priority, default_model_api_base, and the other values of output_format.

rubrics_verifier reads prompt_responses by prompt_id, and deployed_agents by agent_name when it has one. It writes metadata["verifications"][<verifier_id>]: format; results, one row per criterion with the judge's result, score and justification; the score; and where the judge's own trajectory is stored. When the PromptResponse carries an error, the step skips the judge and stores a score of 0 with one row, prompt_error.

The tone step
{"id": "tone", "type": "rubrics_verifier", "prompt_id": "reply", "verifier_id": "tone", "criteria": [
  {"id": "polite", "criterion": "The email to Dana is polite and signed.", "weight": 1},
  {"id": "explains", "criterion": "The email tells Dana what was done with the Q3 numbers.", "weight": 1}
], "depends_on": [{"task_step_id": "reply"}]}

With all_pass, tone scores 1.0 only when the judge passes both criteria. A judge you name in agent_name can be any registered agent, and its prompt lists the files a load_artifact step copied into it, as the collect_artifacts example shows.

collect_artifacts

collect_artifacts copies files out of a deployed agent's sandbox into the object store, registers each as a file artifact, and puts them together in one file artifact universe, so they outlast the sandbox. With no artifact_paths, it collects every file under base_path, and an empty folder collects nothing without failing. When you name files and none of them can be read, the step fails.

FieldDefaultMeaning
agent_namedefault-agentThe deployed agent whose sandbox to read
artifact_paths[], every file under base_pathThe files to collect, relative to base_path or absolute
base_path/app/artifactThe folder relative paths are read from
artifacts_keyexpected_artifactsA seed key whose comma-separated file list replaces artifact_paths for the run
manifest_step_idnoneTake the files from the artifacts that a prompt_agent step's structured reply declared, by that step's id
exclude_basenames[]File names to skip when collecting a whole folder
universe_id_suffixnoneAdded to the universe's id, so two collect steps in one run make two universes

Omitted: env_id, sandbox_name and container_name, which collect from a desktop env's controller, a sandbox's host or a container instead of an agent.

collect_artifacts reads deployed_agents by agent_name. It writes metadata["collected_artifacts"][<step id>]: each file's object URL under artifacts, and the {id, version} of the universe under file_artifact_universe. A load_artifact step whose artifact_from_step_id is the collect step's id loads that universe, which is how one agent hands its files to another.

In chart-q3, the first example on Core concepts, an analyst agent saves a chart as /app/artifact/q3.png. These three steps collect it, copy it into the sandbox of a judge agent deployed beside the analyst, and have that judge grade the analyst's reply:

task.json (three steps of chart-q3)
[
  {"id": "collect", "type": "collect_artifacts", "agent_name": "analyst", "artifact_paths": ["q3.png"], "depends_on": [{"task_step_id": "chart"}]},
  {"id": "handoff", "type": "load_artifact", "agent_name": "judge", "artifact_from_step_id": "collect", "depends_on": [{"task_step_id": "collect"}, {"task_step_id": "judge"}]},
  {"id": "grade", "type": "rubrics_verifier", "agent_name": "judge", "prompt_id": "chart", "verifier_id": "chart", "criteria": [{"id": "bars", "criterion": "q3.png is a bar chart with one bar per month of Q3.", "weight": 1}], "depends_on": [{"task_step_id": "handoff"}]}
]

handoff copies q3.png into the judge's /tmp/file_artifacts, and grade's prompt tells the judge the file is there, so it checks the chart itself rather than the analyst's description of it.

These six steps make up most tasks. More task steps covers the other built-in steps, from env_outcome_verifier, which checks an env's end state with your code, and teardown_sandboxes, which stops what a run deployed, to the steps that set an env's clock and triggers.

Last updated on

Ask a question · Report an issue

On this page