What is a task step
A task step is a node in a graph of computation
Every RL task can be expressed as a directed acyclic graph (DAG) of computation. An env or an agent is deployed. An agent is connected to an env. The agent is prompted. Artifacts are collected from the agent. An agent judge verifier generates rubric scores. And so on. Each of these is a node in that DAG, a task step. This lets you create RL tasks without writing code, and makes it simple to see what any task is doing.
Each tab below runs one of the tasks from Core Concepts, from the start of its DAG, and shows what each step adds to the run's context as it finishes.
A step is one JSON object
You write a task as a JSON list, and each step is one object in it. The rest of this page follows
the first tab's task, reply-to-dana: it deploys the email env from
Environments, loads Dana's email into it, deploys an agent, asks it to reply,
and checks what it sent. Creating a task shows all five of its
steps. This is the fourth, reply, which prompts the agent:
{
"id": "reply",
"type": "prompt_agent",
"agent_name": "assistant",
"prompt_id": "reply",
"prompt": "Reply to Dana.",
"depends_on": [{"task_step_id": "inbox"}, {"task_step_id": "assistant"}]
}A step has two kinds of fields. A few exist on every step, whatever its type:
| Field | Default | Meaning |
|---|---|---|
id | required | Unique within the task. depends_on, retry_config and per-run step overrides name a step by its id |
type | required | The kind of step, which selects the class that runs it |
depends_on | omitted, every earlier step | The steps this one waits for, as {"task_step_id": "<id>"} edges; see depends_on |
fail_task_on_error | true | Whether a failure of this step fails the run; see When a step fails |
retry_config | none | Where to roll back to and run again from when this step fails, and how often |
version | set by the store | The step's own version. agent-env task create stores each step as a document of its own, versioned under its id; see Versions |
The rest are the fields of the step's type: agent_name, prompt_id and prompt belong to
prompt_agent. A field you leave out takes the class's default, such as timeout_seconds: 600.
agent-env task create passes every field except type and retry_config to the class's
constructor as keyword arguments, so a field the class does not take stops the create before
anything is saved, not in the middle of a run.
The type picks the class
A step's type is looked up in the step registry, which maps each type to a Python class, a
subclass of TaskStep. prompt_agent is PromptAgentTaskStep. The registry is built from three
sources, in order:
- The built-in steps, nearly 50 of them, which More task steps lists.
- Classes that installed packages register under the
agent_env.task_stepsentry point. - Classes that
[task_steps] implsnames in yourconfig.toml.
What a step does when it runs
When a step's turn comes, the run calls its async execute(context). context is the run's
TaskStepContext: one object for the whole run, handed to every step. What execute returns is
ignored. A step publishes what it made by adding it to the context, and a later step finds it there.
Steps share results through four typed lists and a dictionary on the context. Each list entry carries a name, and steps find each other's results by those names:
| Field | Entries | Added by | Found by |
|---|---|---|---|
deployed_envs | DeployedEnv: an env instance, with its mcp_url | deploy_env | env_id |
deployed_agents | DeployedAgent: an agent, with its A2A URL | deploy_agent and a few others | agent_name |
deployed_sandboxes | DeployedSandbox: a sandbox | deploy_sandbox | sandbox_name |
prompt_responses | PromptResponse: an agent's reply, with its trajectory | prompt_agent | prompt_id |
metadata | Anything else | Any step | Its key |
metadata is free-form. Verifiers write metadata["verifications"][<verifier_id>] with results
and a score, the run records each step that raised in failed_steps, and a run's caller puts
its per-run overrides in user_overrides.
depends_on
depends_on lists the steps a step waits for. It has three forms:
depends_on | The step waits for | In reply-to-dana |
|---|---|---|
| omitted | Every step listed before it | Without it, reply would wait for email, inbox and assistant |
[] | Nothing; it starts in the first wave | email |
[{"task_step_id": "<id>"}, ...] | Exactly those steps | reply waits for inbox and assistant |
A task whose steps all omit depends_on runs them one after another, in the order they are listed.
An edge may only point to a step listed earlier, and agent-env task create rejects one that does
not, before anything is saved.
A step starts as soon as every step it waits for has finished, so steps with no path between them
run concurrently, and no step waits for anything but its own depends_on. In this longer
reply-to-dana, a slow rubric and a quick check start together once the agent has replied, and the
snapshot after the check starts long before the rubric ends:
Running a task runs this task.
When a step fails
The run records every step that raises in metadata["failed_steps"], with its id, type and
error. With fail_task_on_error: true, the default, no further step starts, the steps already
running finish, and the run fails with the step's error. With false, the steps that wait for this
one run anyway, whether or not it produced what they need.
retry_config runs part of the task again when the step fails, before fail_task_on_error
applies. It names where to go back to, retry_from_step_id, the step itself or one of its
ancestors, and how many times, max_retries. Here reply has
{"retry_from_step_id": "email", "max_retries": 1}: when it fails, the run rolls the whole task
back and runs it again on a fresh env instance:
Task run failures shows what a failed run leaves behind.
Your own step
A step type is a Python class: a subclass of TaskStep with its own type and an
async execute(context). It also implements from_dict, which rebuilds the step from its stored
document, and to_dict, which must add the class's own fields. This one checks that the reply
under a prompt_id mentions a piece of text, and writes a verification:
from agent_env.task_step.context import TaskStepContext
from agent_env.task_step.task_step import TaskStep
class ReplyMentionsStep(TaskStep):
type = "reply_mentions"
def __init__(self, id, version, prompt_id, text, verifier_id,
depends_on=None, fail_task_on_error=True):
super().__init__(id, version, depends_on=depends_on,
fail_task_on_error=fail_task_on_error)
self.prompt_id = prompt_id
self.text = text
self.verifier_id = verifier_id
def to_dict(self) -> dict:
return {**super().to_dict(), "prompt_id": self.prompt_id,
"text": self.text, "verifier_id": self.verifier_id}
@classmethod
def from_dict(cls, data: dict) -> "ReplyMentionsStep":
return cls(**cls._base_from_dict(data), prompt_id=data["prompt_id"],
text=data["text"], verifier_id=data["verifier_id"])
async def execute(self, context: TaskStepContext) -> TaskStepContext:
reply = next((r for r in context.prompt_responses
if r.prompt_id == self.prompt_id), None)
if reply is None:
raise RuntimeError(f"No response with prompt_id '{self.prompt_id}'")
passed = self.text.lower() in reply.response.lower()
context.metadata.setdefault("verifications", {})[self.verifier_id] = {
"results": [{"criterion": f"The reply mentions {self.text!r}", "result": passed}],
"score": 1.0 if passed else 0.0,
}
return contextRegister it in config.toml:
[task_steps]
impls = ["mycorp.steps:ReplyMentionsStep"]A package you install can register it instead, with an entry point named after the type:
[project.entry-points."agent_env.task_steps"]
reply_mentions = "mycorp.steps:ReplyMentionsStep"With the step understood, Creating a task puts the five steps of
reply-to-dana together and stores them as a task.
Last updated on