# Running a task (https://www.agentenvframework.com/docs/tasks/running)

> A task, once designed, can be run many times

## Running a task

`agent-env task run` loads a stored task and runs it in the terminal's own process:

```bash title="Terminal"
agent-env task run --id reply-to-dana --output-dir out --project-id <project-id>
```

```text title="Output (banner and most step results left out)"
Fetching task: id=reply-to-dana version=latest...
Found task: id=reply-to-dana version=2 steps=5
Task instance: reply-to-dana-gbgs1yub
Running step [1/5]: deploy_env (id=email)...
Completed step [1/5]: deploy_env (id=email) [42.9s]

Running step [2/5]: load_artifact (id=inbox)...

Running step [3/5]: deploy_agent (id=assistant)...
Completed step [2/5]: load_artifact (id=inbox) [1.9s]
Completed step [3/5]: deploy_agent (id=assistant) [23.7s]

Running step [4/5]: prompt_agent (id=reply)...
Completed step [4/5]: prompt_agent (id=reply) [63.3s]

Running step [5/5]: env_outcome_verifier (id=check)...
Completed step [5/5]: env_outcome_verifier (id=check) [0.4s]
score (verifier_id=sent): 1.0

Task completed!
Task context written to: out/reply-to-dana_811f14b9.json
```

The run loads the latest version unless you pass `--version`. A step starts as soon as the steps in
its `depends_on` have finished. Runs of the same task go at once with `--k`, each with its own task
instance and context:

1. `--k 3` starts three runs of `reply-to-dana` at once. Each registers its own task instance and deploys its own `email`.
2. Each run follows the DAG on its own clock: run 3’s `inbox` and `assistant` start together while run 2’s env is still deploying.
3. Runs 1 and 3 are prompting their agents while run 2’s agent is still starting. No run waits for another.
4. Run 3 finishes first: its `check` finds its own env instance in its own context and scores it 1.0.
5. Run 1 finishes too. Its agent answered without emailing Dana, so `check` scores 0.0, and the run still completes.
6. Run 2 finishes, and the command prints `All 3 runs completed!`: three instances, three contexts, three files.

Parts of the scene:

- **--k 3**: Starts three runs of one task version together, in one process, with no limit on how many. Each line the command prints carries a `[run N]` tag.
- **A run’s task instance**: Each run registers its own task instance, with its own id and record. The record keeps the run’s status, its steps and the context they left.
- **A run’s DAG**: Every run executes the same steps, and in each a step starts once the steps in its `depends_on` have finished: `inbox` and `assistant` together, after `email`. The runs keep their own pace.
- **A run’s context**: Each run has its own `TaskStepContext`: its own env instance in `deployed_envs`, its own agent in `deployed_agents`, its own reply and score. No step reads another run’s.
- **The score**: What `check` stored under `metadata.verifications.sent` for this run. A score of 0.0 is a finished run that failed its check, not a failed run.
- **A run’s file**: Each run writes its own context file to `--output-dir`, `<task id>_<8 hex>.json`, once it completes.

## The `TaskStepContext`

1. `task run` makes one `TaskStepContext` for the run: empty lists, and `metadata` with the task id and the run’s options.
2. `email` deploys the env and adds its `DeployedEnv` to `deployed_envs`: the `env_id`, the instance and its MCP URL.
3. Both find `email` by its `env_id`. `inbox` adds nothing; `assistant` adds its `DeployedAgent` and the default model.
4. `reply` finds its agent by `agent_name` and adds the agent’s answer to `prompt_responses` under `prompt_id: reply`.
5. `check` reads `email`’s MCP URL, runs `verify()` there and writes its rows and score to `metadata.verifications.sent`.
6. The final context goes to `out/reply-to-dana_811f14b9.json`, with `litellm_api_key` and the other secrets removed.

Parts of the scene:

- **agent-env task run**: Runs the task once in this process. `--litellm-api-key` goes into `metadata.user_overrides`, where the steps that call a model read it; `--output-dir` is where the final context goes.
- **default_agent_model**: Set by `deploy_agent` from `[model.roles] agent` or `[model] default`. A `prompt_agent` step without a `model`, and without `--agent-model`, asks for this one.
- **instance_id**: The id of the task instance this run is: the task id and eight random characters. Its record is in the document store; see The task run instance.
- **user_overrides**: The run’s options that steps read when they run: the model key, the agent and sandboxes, the env state. Keys such as `litellm_api_key` never reach the record or the file.
- **The output file**: The final context as JSON, named `<task id>_<8 hex>.json`, in `--output-dir` or a new directory under `/tmp`. `litellm_api_key`, `judge_litellm_api_key`, `usersim_api_key`, `remote_tokens` and `cf_access_client_secret` are removed at every depth.

The `TaskStepContext` is a ledger of what has happened during a task run. When the run completes,
the command writes it to `--output-dir` as `<task id>_<8 hex>.json`, with secrets such as
`litellm_api_key` removed.

## The task run instance

Each run is a task instance, with a record in the document store: the task id and version,
`status`, and `current_step` of `total_steps`. After each step the run writes to the record only
what that step changed, so steps that finish together don't overwrite each other.
`agent-env task get-instance --id reply-to-dana-gbgs1yub` reads it back, during the run or after it:

1. `task run` registers a task instance in `task_instances`: its id, the task version, `status: running` and 0 of 5 steps.
2. When `email` finishes, the record gets only what it changed: its `DeployedEnv`, its entry in `completed_steps`, 1 of 5.
3. `inbox` and `assistant` finish together. Each writes only its own paths, so neither overwrites the other: 3 of 5.
4. `reply` adds the agent’s answer to the record’s context, and the count reaches 4 of 5.
5. At 5 of 5 the record’s `status` becomes `completed`, and `completed_at_utc` is set.
6. `agent-env task get-instance` reads the record from the document store, during the run or after it.
7. `--k 3` registers three instances at once, each with its own id, record and file.

Parts of the scene:

- **The terminal**: What `agent-env task run` prints: the instance id once the run registers it, and a line as each step finishes. The number in brackets is the step’s place in the task, not the order it ran in.
- **The task instance record**: One document in the `task_instances` collection of the document store, `[stores.document]` in `config.toml`: SQLite on disk by default. It outlives the run and the process that ran it.
- **instance_id**: The task id and eight random lowercase letters and digits. Every run gets a new one, a resumed run included.
- **task_version**: The version this run loaded: the latest unless `--version` pins one. A later create doesn’t change a record’s version.
- **status**: `running` from the start, `completed` once every step is recorded, `failed` when a step fails with `fail_task_on_error: true`, and then `error` holds its message.
- **current_step and total_steps**: `current_step` counts the entries in `completed_steps`; `total_steps` is the task’s step count. `get-instance` prints them as `Progress: 5/5`.
- **completed_steps**: Each finished step’s id and status, `success` or `failure`, in the order they finished. A resume records the steps before `--start-step` as `success` without running them.
- **A step’s write**: After a step, the run diffs the context against the one before it and writes only the paths that changed, with the step’s entry and the count. Two steps that finish together touch different paths.
- **context**: The run’s context as the steps leave it, without secrets. After a failure it is the context at the moment of the failure, which a resume starts from.
- **agent-env task get-instance**: Prints a record: its id, task and version, status, progress, times and the whole context as JSON. `Task.get_instance(id)` reads it from Python.
- **--k 3**: Three runs of one task version, started together in one process. Each is its own task instance with its own env instance and agent, record and output file.

## Task run options

| Option                             | Default                      | Meaning                                                                               |
| ---------------------------------- | ---------------------------- | ------------------------------------------------------------------------------------- |
| `--version`                        | the latest                   | The task version to run                                                               |
| `--output-dir`                     | a new directory under `/tmp` | Where the context file goes                                                           |
| `--k`                              | `1`                          | How many runs to start at once                                                        |
| `--agent-model`                    | none                         | The model every `prompt_agent` step asks for, over the step's `model`                 |
| `--a2a-agent-id`                   | none                         | The agent every `deploy_agent` step deploys, over its `a2a_agent_id`                  |
| `--env-sandbox`, `--agent-sandbox` | none                         | The sandbox for every `deploy_env` or `deploy_agent` step, over its `sandbox_type`    |
| `--start-step`, `--context-json`   | `0`, none                    | The 0-based index of the first step to run, and a JSON file to start the context from |
| `--litellm-api-key`                | none                         | The model API key for this run                                                        |

`agent-env task run-batch --id <task id> --seeds seeds.csv` runs a task once per row of a CSV file,
and each row's columns fill the task's [placeholders](https://www.agentenvframework.com/docs/tasks/creating.md#placeholders) of the same
names: a `sender` column fills `<sender>` in `Reply to <sender>.`. Without a sandbox option or `sandbox_type`, envs
and agents run in `[sandbox] default` and `[sandbox] agent_default`, both `local` Docker unless
`config.toml` says otherwise. `agent-env task run --help` lists the
rest.

## Task run failures

A step that raises adds an entry to `metadata.failed_steps`. With `fail_task_on_error: true`, the
default, no further step starts and the record ends `failed`, with the error:

```text title="Output (last lines)"
Running step [5/5]: env_outcome_verifier (id=check)...
Step check failed (fatal): Artifact check-reply not found
Error: Artifact check-reply not found
```

With `false`, the steps that wait for it run anyway, so read `failed_steps` before you trust the
score. A step with a `retry_config` first runs part of the task again; see
[When a step fails](https://www.agentenvframework.com/docs/tasks/task-steps.md#when-a-step-fails).

## Resume a task run

A failed run leaves its env instance and agent running, so you can pick it up at the failed step.
Fix what failed, then save the failed run's context from its record, without its `failed_steps`:

```python title="save_context.py"
import json

from agent_env.task import Task

context = Task.get_instance("reply-to-dana-dvbt79er").context
context["metadata"].pop("failed_steps", None)
with open("context.json", "w") as f:
    json.dump(context, f, indent=2)
```

Then run from `check`, index 4 of the step list:

```bash title="Terminal"
agent-env task run --id reply-to-dana --start-step 4 --context-json context.json --output-dir out --project-id <project-id>
```

The steps before the index are recorded as succeeded without running, and every step from it on
runs again. The resumed run is a new task instance.