Skip to content
AgentEnv Framework

What is a task step

A task step is a node in a graph of computation

Every RL task can be expressed as a directed acyclic graph (DAG) of computation. An env or an agent is deployed. An agent is connected to an env. The agent is prompted. Artifacts are collected from the agent. An agent judge verifier generates rubric scores. And so on. Each of these is a node in that DAG, a task step. This lets you create RL tasks without writing code, and makes it simple to see what any task is doing.

Each tab below runs one of the tasks from Core Concepts, from the start of its DAG, and shows what each step adds to the run's context as it finishes.

A step is one JSON object

You write a task as a JSON list, and each step is one object in it. The rest of this page follows the first tab's task, reply-to-dana: it deploys the email env from Environments, loads Dana's email into it, deploys an agent, asks it to reply, and checks what it sent. Creating a task shows all five of its steps. This is the fourth, reply, which prompts the agent:

The reply step
{
  "id": "reply",
  "type": "prompt_agent",
  "agent_name": "assistant",
  "prompt_id": "reply",
  "prompt": "Reply to Dana.",
  "depends_on": [{"task_step_id": "inbox"}, {"task_step_id": "assistant"}]
}

A step has two kinds of fields. A few exist on every step, whatever its type:

FieldDefaultMeaning
idrequiredUnique within the task. depends_on, retry_config and per-run step overrides name a step by its id
typerequiredThe kind of step, which selects the class that runs it
depends_onomitted, every earlier stepThe steps this one waits for, as {"task_step_id": "<id>"} edges; see depends_on
fail_task_on_errortrueWhether a failure of this step fails the run; see When a step fails
retry_confignoneWhere to roll back to and run again from when this step fails, and how often
versionset by the storeThe step's own version. agent-env task create stores each step as a document of its own, versioned under its id; see Versions

The rest are the fields of the step's type: agent_name, prompt_id and prompt belong to prompt_agent. A field you leave out takes the class's default, such as timeout_seconds: 600. agent-env task create passes every field except type and retry_config to the class's constructor as keyword arguments, so a field the class does not take stops the create before anything is saved, not in the middle of a run.

The type picks the class

A step's type is looked up in the step registry, which maps each type to a Python class, a subclass of TaskStep. prompt_agent is PromptAgentTaskStep. The registry is built from three sources, in order:

  1. The built-in steps, nearly 50 of them, which More task steps lists.
  2. Classes that installed packages register under the agent_env.task_steps entry point.
  3. Classes that [task_steps] impls names in your config.toml.

What a step does when it runs

When a step's turn comes, the run calls its async execute(context). context is the run's TaskStepContext: one object for the whole run, handed to every step. What execute returns is ignored. A step publishes what it made by adding it to the context, and a later step finds it there.

Steps share results through four typed lists and a dictionary on the context. Each list entry carries a name, and steps find each other's results by those names:

FieldEntriesAdded byFound by
deployed_envsDeployedEnv: an env instance, with its mcp_urldeploy_envenv_id
deployed_agentsDeployedAgent: an agent, with its A2A URLdeploy_agent and a few othersagent_name
deployed_sandboxesDeployedSandbox: a sandboxdeploy_sandboxsandbox_name
prompt_responsesPromptResponse: an agent's reply, with its trajectoryprompt_agentprompt_id
metadataAnything elseAny stepIts key

metadata is free-form. Verifiers write metadata["verifications"][<verifier_id>] with results and a score, the run records each step that raised in failed_steps, and a run's caller puts its per-run overrides in user_overrides.

depends_on

depends_on lists the steps a step waits for. It has three forms:

depends_onThe step waits forIn reply-to-dana
omittedEvery step listed before itWithout it, reply would wait for email, inbox and assistant
[]Nothing; it starts in the first waveemail
[{"task_step_id": "<id>"}, ...]Exactly those stepsreply waits for inbox and assistant

A task whose steps all omit depends_on runs them one after another, in the order they are listed. An edge may only point to a step listed earlier, and agent-env task create rejects one that does not, before anything is saved.

A step starts as soon as every step it waits for has finished, so steps with no path between them run concurrently, and no step waits for anything but its own depends_on. In this longer reply-to-dana, a slow rubric and a quick check start together once the agent has replied, and the snapshot after the check starts long before the rubric ends:

Running a task runs this task.

When a step fails

The run records every step that raises in metadata["failed_steps"], with its id, type and error. With fail_task_on_error: true, the default, no further step starts, the steps already running finish, and the run fails with the step's error. With false, the steps that wait for this one run anyway, whether or not it produced what they need.

retry_config runs part of the task again when the step fails, before fail_task_on_error applies. It names where to go back to, retry_from_step_id, the step itself or one of its ancestors, and how many times, max_retries. Here reply has {"retry_from_step_id": "email", "max_retries": 1}: when it fails, the run rolls the whole task back and runs it again on a fresh env instance:

Task run failures shows what a failed run leaves behind.

Your own step

A step type is a Python class: a subclass of TaskStep with its own type and an async execute(context). It also implements from_dict, which rebuilds the step from its stored document, and to_dict, which must add the class's own fields. This one checks that the reply under a prompt_id mentions a piece of text, and writes a verification:

mycorp/steps.py
from agent_env.task_step.context import TaskStepContext
from agent_env.task_step.task_step import TaskStep


class ReplyMentionsStep(TaskStep):
    type = "reply_mentions"

    def __init__(self, id, version, prompt_id, text, verifier_id,
                 depends_on=None, fail_task_on_error=True):
        super().__init__(id, version, depends_on=depends_on,
                         fail_task_on_error=fail_task_on_error)
        self.prompt_id = prompt_id
        self.text = text
        self.verifier_id = verifier_id

    def to_dict(self) -> dict:
        return {**super().to_dict(), "prompt_id": self.prompt_id,
                "text": self.text, "verifier_id": self.verifier_id}

    @classmethod
    def from_dict(cls, data: dict) -> "ReplyMentionsStep":
        return cls(**cls._base_from_dict(data), prompt_id=data["prompt_id"],
                   text=data["text"], verifier_id=data["verifier_id"])

    async def execute(self, context: TaskStepContext) -> TaskStepContext:
        reply = next((r for r in context.prompt_responses
                      if r.prompt_id == self.prompt_id), None)
        if reply is None:
            raise RuntimeError(f"No response with prompt_id '{self.prompt_id}'")
        passed = self.text.lower() in reply.response.lower()
        context.metadata.setdefault("verifications", {})[self.verifier_id] = {
            "results": [{"criterion": f"The reply mentions {self.text!r}", "result": passed}],
            "score": 1.0 if passed else 0.0,
        }
        return context

Register it in config.toml:

.agentenv/config.toml
[task_steps]
impls = ["mycorp.steps:ReplyMentionsStep"]

A package you install can register it instead, with an entry point named after the type:

pyproject.toml
[project.entry-points."agent_env.task_steps"]
reply_mentions = "mycorp.steps:ReplyMentionsStep"

With the step understood, Creating a task puts the five steps of reply-to-dana together and stores them as a task.

Last updated on

Ask a question · Report an issue

On this page