# What is a task step (https://www.agentenvframework.com/docs/tasks/task-steps)

> A task step is a node in a graph of computation

Every RL task can be expressed as a directed acyclic graph (DAG) of computation. An env or an agent
is deployed. An agent is connected to an env. The agent is prompted. Artifacts are collected from the
agent. An agent judge verifier generates rubric scores. And so on. Each of these is a node in that
DAG, a task step. This lets you create RL tasks without writing code, and makes it simple to see what
any task is doing.

Each tab below runs one of the tasks from [Core Concepts](https://www.agentenvframework.com/docs/core-concepts.md#tasks), from the start
of its DAG, and shows what each step adds to the run's context as it finishes.

1. reply-to-dana: The run starts a task instance; no step has written to its `TaskStepContext` yet.
2. reply-to-dana: `email` deploys the env and adds its instance to `deployed_envs`.
3. reply-to-dana: `inbox` loads Dana’s email into the instance, not the context; `assistant` joins `deployed_agents`.
4. reply-to-dana: `reply` finds `assistant` by `agent_name` and adds its answer to `prompt_responses`.
5. reply-to-dana: `check` finds `email` by `env_id`, grades the instance and writes `verifications.sent`.
6. reply-to-dana: The final context is the run’s result, kept in the instance’s record and in the output file.
7. chart-q3: The run starts a task instance; no step has written to its `TaskStepContext` yet.
8. chart-q3: Neither agent depends on anything, so both deploy at once into `deployed_agents`.
9. chart-q3: `data` copies `q3.csv` into the analyst’s sandbox, noted in `loaded_file_artifact_universes`.
10. chart-q3: `chart` prompts the analyst, and its reply joins `prompt_responses` under `chart`.
11. chart-q3: `collect` copies `q3.png` out of the analyst’s sandbox into `collected_artifacts`.
12. chart-q3: `handoff` loads what `collect` stored into the judge’s sandbox: a second loaded-files entry.
13. chart-q3: `grade` has the judge grade the analyst’s work, with `q3.png` at hand, into `verifications`.
14. chart-q3: The final context is the run’s result, kept in the instance’s record and in the output file.
15. talk-to-dana: The run starts a task instance; no step has written to its `TaskStepContext` yet.
16. talk-to-dana: `dana` needs no env, so it deploys beside `email`: one agent and one env instance.
17. talk-to-dana: `inbox` fills the instance’s inbox, and `assistant` joins `deployed_agents`.
18. talk-to-dana: `talk` runs up to six turns between the two agents, and records the reply and the conversation.
19. talk-to-dana: `grade` scores the whole conversation and writes `verifications.conversation`.
20. talk-to-dana: The final context is the run’s result, kept in the instance’s record and in the output file.
21. q3-to-sam: The run starts a task instance; no step has written to its `TaskStepContext` yet.
22. q3-to-sam: Two `deploy_env` steps with nothing between them start together: two entries in `deployed_envs`.
23. q3-to-sam: Both loads fill their instances, and `writer` and `planner` join `deployed_agents`.
24. q3-to-sam: `peer` gives the planner a route to the writer, recorded in `agent_peerings`.
25. q3-to-sam: `ask` prompts the planner, which hands the email to the writer; the reply joins `prompt_responses`.
26. q3-to-sam: `check` finds the email to Sam in `email`’s sent folder and writes its score.
27. q3-to-sam: The final context is the run’s result, kept in the instance’s record and in the output file.
28. inbox-week: The run starts a task instance; no step has written to its `TaskStepContext` yet.
29. inbox-week: `email` deploys the env and adds its instance to `deployed_envs`.
30. inbox-week: `assistant` joins `deployed_agents`; `deadline` records its trigger in `env_trigger_registrations`.
31. inbox-week: `clock` starts Monday 09:00 at 3,600× and records it in `clock_configurations`.
32. inbox-week: `week` prompts the assistant; its reply and the triggers’ final state join the context.
33. inbox-week: `check` writes its score, and `snapshot` records the new universe in `env_snapshotted_universes`.
34. inbox-week: The final context is the run’s result, kept in the instance’s record and in the output file.
35. support-world: The run starts a task instance; no step has written to its `TaskStepContext` yet.
36. support-world: The 20 customers need no env, so they deploy beside `office`: 20 agents and one instance.
37. support-world: `company` loads its universe into `office`; the support and world agents join `deployed_agents`.
38. support-world: `peer` records the support agents’ routes, and `rules` the world’s triggers.
39. support-world: `clock` starts the week at 3,600× and records it in `clock_configurations`.
40. support-world: Twenty conversations run at once: 20 replies, 20 conversation ids and the triggers’ state.
41. support-world: Twenty rubric scores and the end-state check land in `verifications`.
42. support-world: The final context is the run’s result, kept in the instance’s record and in the output file.

Parts of the scene:

- **TaskStepContext**: One object for the whole run. The run hands it to every step’s `execute(context)`; a step adds what it made, and a later step finds it by name. After each step, the run writes what changed to the instance’s record.
- **deployed_envs**: A `DeployedEnv` for each `deploy_env` step: the env id and version, the instance id and the instance’s `mcp_url`. Later steps find it by `env_id`.
- **deployed_agents**: A `DeployedAgent` for each `deploy_agent` step, with its A2A URL and its sandbox. Later steps find it by `agent_name`.
- **deployed_sandboxes**: A `DeployedSandbox` for each `deploy_sandbox` step, found by `sandbox_name`. None of these tasks deploys a bare sandbox.
- **prompt_responses**: A `PromptResponse` for each `prompt_agent` step: the prompt, the agent’s reply and where its trajectory is stored. Verifiers find it by `prompt_id`.
- **metadata**: Everything else, under one key per kind of result. The panel shows only the keys this task’s steps write; `agent-env task run` also starts it with `run_group_id`, `task_id` and `project_id`.
- **verifications**: One entry per verifier, under its `verifier_id`, with its `results` and `score`. A verifier without a `verifier_id` gets a random one, so these tasks name each.
- **loaded_file_artifact_universes**: Each set of files `load_artifact` copied into a sandbox, with the agent it went to. `rubrics_verifier` lists a judge’s files in its prompt.
- **loaded_environment_universes**: Each environment universe `load_artifact` loaded into an env instance. A single environment artifact, such as `inbox`, loads with no entry.
- **collected_artifacts**: The files each `collect_artifacts` step copied out of a sandbox, under the step’s id, with the file artifact universe that holds them.
- **a2a_conversations**: A multi-turn `prompt_agent` records its conversation’s id under its step id.
- **agent_peerings**: Each route `peer_agents` pushed: the agent that got it and the peers it names.
- **env_trigger_registrations**: Each `register_env_triggers` step: its env, the trigger ids the gateway added and the executor agent, if any.
- **clock_configurations**: Each `sync_env_clock` step: the virtual time, the rate and which servers synced to it.
- **env_trigger_state**: After its prompt, `prompt_agent` captures the trigger state of each env with registered triggers, under the env’s id.
- **env_snapshotted_universes**: The id and version of the universe each `snapshot_env` step created, under the step’s id. A later `load_artifact` finds it with `artifact_from_step_id`.

## A step is one JSON object

You write a task as a JSON list, and each step is one object in it. The rest of this page follows
the first tab's task, `reply-to-dana`: it deploys the `email` env from
[Environments](https://www.agentenvframework.com/docs/environments.md), loads Dana's email into it, deploys an agent, asks it to reply,
and checks what it sent. [Creating a task](https://www.agentenvframework.com/docs/tasks/creating.md) shows all five of its
steps. This is the fourth, `reply`, which prompts the agent:

```json title="The reply step"
{
  "id": "reply",
  "type": "prompt_agent",
  "agent_name": "assistant",
  "prompt_id": "reply",
  "prompt": "Reply to Dana.",
  "depends_on": [{"task_step_id": "inbox"}, {"task_step_id": "assistant"}]
}
```

A step has two kinds of fields. A few exist on every step, whatever its type:

| Field                | Default                     | Meaning                                                                                                                                                          |
| -------------------- | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `id`                 | required                    | Unique within the task. `depends_on`, `retry_config` and per-run step overrides name a step by its id                                                            |
| `type`               | required                    | The kind of step, which selects the class that runs it                                                                                                           |
| `depends_on`         | omitted, every earlier step | The steps this one waits for, as `{"task_step_id": "<id>"}` edges; see [depends\_on](#depends_on)                                                                |
| `fail_task_on_error` | `true`                      | Whether a failure of this step fails the run; see [When a step fails](#when-a-step-fails)                                                                        |
| `retry_config`       | none                        | Where to roll back to and run again from when this step fails, and how often                                                                                     |
| `version`            | set by the store            | The step's own version. `agent-env task create` stores each step as a document of its own, versioned under its id; see [Versions](https://www.agentenvframework.com/docs/tasks/creating.md#versions) |

The rest are the fields of the step's type: `agent_name`, `prompt_id` and `prompt` belong to
`prompt_agent`. A field you leave out takes the class's default, such as `timeout_seconds: 600`.
`agent-env task create` passes every field except `type` and `retry_config` to the class's
constructor as keyword arguments, so a field the class does not take stops the create before
anything is saved, not in the middle of a run.

## The type picks the class

A step's `type` is looked up in the step registry, which maps each type to a Python class, a
subclass of `TaskStep`. `prompt_agent` is `PromptAgentTaskStep`. The registry is built from three
sources, in order:

1. The built-in steps, nearly 50 of them, which [More task steps](https://www.agentenvframework.com/docs/tasks/more-steps.md) lists.
2. Classes that installed packages register under the `agent_env.task_steps` entry point.
3. Classes that `[task_steps] impls` names in your `config.toml`.

## What a step does when it runs

When a step's turn comes, the run calls its `async execute(context)`. `context` is the run's
`TaskStepContext`: one object for the whole run, handed to every step. What `execute` returns is
ignored. A step publishes what it made by adding it to the context, and a later step finds it there.

Steps share results through four typed lists and a dictionary on the context. Each list entry
carries a name, and steps find each other's results by those names:

| Field                | Entries                                                 | Added by                        | Found by       |
| -------------------- | ------------------------------------------------------- | ------------------------------- | -------------- |
| `deployed_envs`      | `DeployedEnv`: an env instance, with its `mcp_url`      | `deploy_env`                    | `env_id`       |
| `deployed_agents`    | `DeployedAgent`: an agent, with its A2A URL             | `deploy_agent` and a few others | `agent_name`   |
| `deployed_sandboxes` | `DeployedSandbox`: a sandbox                            | `deploy_sandbox`                | `sandbox_name` |
| `prompt_responses`   | `PromptResponse`: an agent's reply, with its trajectory | `prompt_agent`                  | `prompt_id`    |
| `metadata`           | Anything else                                           | Any step                        | Its key        |

`metadata` is free-form. Verifiers write `metadata["verifications"][<verifier_id>]` with `results`
and a `score`, the run records each step that raised in `failed_steps`, and a run's caller puts
its per-run overrides in `user_overrides`.

## `depends_on`

`depends_on` lists the steps a step waits for. It has three forms:

| `depends_on`                      | The step waits for                   | In `reply-to-dana`                                                  |
| --------------------------------- | ------------------------------------ | ------------------------------------------------------------------- |
| omitted                           | Every step listed before it          | Without it, `reply` would wait for `email`, `inbox` and `assistant` |
| `[]`                              | Nothing; it starts in the first wave | `email`                                                             |
| `[{"task_step_id": "<id>"}, ...]` | Exactly those steps                  | `reply` waits for `inbox` and `assistant`                           |

A task whose steps all omit `depends_on` runs them one after another, in the order they are listed.
An edge may only point to a step listed earlier, and `agent-env task create` rejects one that does
not, before anything is saved.

A step starts as soon as every step it waits for has finished, so steps with no path between them
run concurrently, and no step waits for anything but its own `depends_on`. In this longer
`reply-to-dana`, a slow rubric and a quick check start together once the agent has replied, and the
snapshot after the check starts long before the rubric ends:

1. `email` starts first: it is the only step with nothing to wait for.
2. `email` finishes at 43 s, and `inbox` and `assistant` both start.
3. `inbox` is done in 2 s, but `reply` also waits for `assistant`.
4. `assistant` finishes at 67 s, so `reply` starts.
5. `reply` finishes at 130 s, and `tone` and `check` start together.
6. `check` is done in under a second, so `snapshot` starts while `tone` is still grading.
7. `snapshot` finishes at 135 s; `teardown` waits for `tone` as well.
8. At 180 s `tone` is still grading, the only step running; `teardown` is waiting for it.
9. `tone` finishes at 225 s, and `teardown` starts at once.
10. Done at 227 s. No step waited for anything but its own `depends_on`.

Parts of the scene:

- **The timeline**: Each bar is one step, from when it started to when it finished. The first five durations are those of the run on Running a task; the other three are an example.
- **email**: `deploy_env`, 43 s. Its `depends_on` is empty, so it starts with the run.
- **inbox**: `load_artifact`, 2 s. It waits only for `email`.
- **assistant**: `deploy_agent`, 24 s. It waits only for `email`, and runs beside `inbox`.
- **reply**: `prompt_agent`, 63 s. It waits for `inbox` and `assistant`, so it starts when the slower of the two, `assistant`, finishes.
- **tone**: `rubrics_verifier`, 95 s: it deploys a judge agent, then has it grade `reply` against the rubric. The slowest step, on a branch of its own.
- **check**: `env_outcome_verifier`, under a second: it runs `verify()` against the `email` instance’s MCP URL.
- **snapshot**: `snapshot_env`, 4 s. It depends only on `check`, so it starts at 130 s, 95 s before `tone` ends.
- **teardown**: `teardown_sandboxes`, 2 s. It depends on `tone` and `snapshot`, so it starts when the later of the two finishes.

[Running a task](https://www.agentenvframework.com/docs/tasks/running.md#running-a-task) runs this task.

## When a step fails

The run records every step that raises in `metadata["failed_steps"]`, with its id, type and
error. With `fail_task_on_error: true`, the default, no further step starts, the steps already
running finish, and the run fails with the step's error. With `false`, the steps that wait for this
one run anyway, whether or not it produced what they need.

`retry_config` runs part of the task again when the step fails, before `fail_task_on_error`
applies. It names where to go back to, `retry_from_step_id`, the step itself or one of its
ancestors, and how many times, `max_retries`. Here `reply` has
`{"retry_from_step_id": "email", "max_retries": 1}`: when it fails, the run rolls the whole task
back and runs it again on a fresh env instance:

1. Attempt 1: `email`, `inbox` and `assistant` have run, and `reply` prompts the agent.
2. `reply` raises: the agent’s A2A task failed. The failure goes to `failed_steps`.
3. `reply`’s `retry_config` goes back to `email`, so the span is `email` and every step after it.
4. The run undoes what the span wrote: the env instance and the agent leave the context.
5. Attempt 2 runs the span again: a fresh `email` instance, the inbox loaded again, a new agent.
6. `reply` succeeds this time, and `check` writes its score.
7. The first failure stays in `failed_steps`, marked `retried`. A second would go to `fail_task_on_error`.

Parts of the scene:

- **retry_config**: Set on `reply`: when the step fails, go back to `retry_from_step_id` and run from there again, at most `max_retries` times, 1 by default. `retry_from_step_id` must be the step itself or one of its ancestors; `agent-env task create` rejects any other.
- **Attempts**: Attempt 2 is the first retry. The task instance stays the same; only the span runs again, under the same instance id.
- **The span**: The step `retry_from_step_id` names and every step that depends on it, directly or not: here the whole task. A retry is refused, and the run fails, if a step outside the span ran at the same time as one inside it.
- **The rollback**: Everything the span wrote is undone, in the instance’s record and in the live context, whatever it was. The env instance and the agent from attempt 1 are not stopped: `retry_resets` lists their sandboxes.
- **failed_steps**: Each step that raised: its id, type and error, when it started, how long it ran and whether it was fatal. A failure a retry recovered from stays, with `retried: true`.
- **retry_resets**: One entry per retry: the attempt, the failed step, where it went back to, the steps it ran again, and the sandboxes the rollback left running, for cleanup.

[Task run failures](https://www.agentenvframework.com/docs/tasks/running.md#task-run-failures) shows what a failed run leaves behind.

## Your own step

A step type is a Python class: a subclass of `TaskStep` with its own `type` and an
`async execute(context)`. It also implements `from_dict`, which rebuilds the step from its stored
document, and `to_dict`, which must add the class's own fields. This one checks that the reply
under a `prompt_id` mentions a piece of text, and writes a verification:

```python title="mycorp/steps.py"
from agent_env.task_step.context import TaskStepContext
from agent_env.task_step.task_step import TaskStep


class ReplyMentionsStep(TaskStep):
    type = "reply_mentions"

    def __init__(self, id, version, prompt_id, text, verifier_id,
                 depends_on=None, fail_task_on_error=True):
        super().__init__(id, version, depends_on=depends_on,
                         fail_task_on_error=fail_task_on_error)
        self.prompt_id = prompt_id
        self.text = text
        self.verifier_id = verifier_id

    def to_dict(self) -> dict:
        return {**super().to_dict(), "prompt_id": self.prompt_id,
                "text": self.text, "verifier_id": self.verifier_id}

    @classmethod
    def from_dict(cls, data: dict) -> "ReplyMentionsStep":
        return cls(**cls._base_from_dict(data), prompt_id=data["prompt_id"],
                   text=data["text"], verifier_id=data["verifier_id"])

    async def execute(self, context: TaskStepContext) -> TaskStepContext:
        reply = next((r for r in context.prompt_responses
                      if r.prompt_id == self.prompt_id), None)
        if reply is None:
            raise RuntimeError(f"No response with prompt_id '{self.prompt_id}'")
        passed = self.text.lower() in reply.response.lower()
        context.metadata.setdefault("verifications", {})[self.verifier_id] = {
            "results": [{"criterion": f"The reply mentions {self.text!r}", "result": passed}],
            "score": 1.0 if passed else 0.0,
        }
        return context
```

Register it in `config.toml`:

```toml title=".agentenv/config.toml"
[task_steps]
impls = ["mycorp.steps:ReplyMentionsStep"]
```

A package you install can register it instead, with an entry point named after the type:

```toml title="pyproject.toml"
[project.entry-points."agent_env.task_steps"]
reply_mentions = "mycorp.steps:ReplyMentionsStep"
```

With the step understood, [Creating a task](https://www.agentenvframework.com/docs/tasks/creating.md) puts the five steps of
`reply-to-dana` together and stores them as a task.