Running a task
A task, once designed, can be run many times
Running a task
agent-env task run loads a stored task and runs it in the terminal's own process:
agent-env task run --id reply-to-dana --output-dir out --project-id <project-id>Fetching task: id=reply-to-dana version=latest...
Found task: id=reply-to-dana version=2 steps=5
Task instance: reply-to-dana-gbgs1yub
Running step [1/5]: deploy_env (id=email)...
Completed step [1/5]: deploy_env (id=email) [42.9s]
Running step [2/5]: load_artifact (id=inbox)...
Running step [3/5]: deploy_agent (id=assistant)...
Completed step [2/5]: load_artifact (id=inbox) [1.9s]
Completed step [3/5]: deploy_agent (id=assistant) [23.7s]
Running step [4/5]: prompt_agent (id=reply)...
Completed step [4/5]: prompt_agent (id=reply) [63.3s]
Running step [5/5]: env_outcome_verifier (id=check)...
Completed step [5/5]: env_outcome_verifier (id=check) [0.4s]
score (verifier_id=sent): 1.0
Task completed!
Task context written to: out/reply-to-dana_811f14b9.jsonThe run loads the latest version unless you pass --version. A step starts as soon as the steps in
its depends_on have finished. Runs of the same task go at once with --k, each with its own task
instance and context:
The TaskStepContext
The TaskStepContext is a ledger of what has happened during a task run. When the run completes,
the command writes it to --output-dir as <task id>_<8 hex>.json, with secrets such as
litellm_api_key removed.
The task run instance
Each run is a task instance, with a record in the document store: the task id and version,
status, and current_step of total_steps. After each step the run writes to the record only
what that step changed, so steps that finish together don't overwrite each other.
agent-env task get-instance --id reply-to-dana-gbgs1yub reads it back, during the run or after it:
Task run options
| Option | Default | Meaning |
|---|---|---|
--version | the latest | The task version to run |
--output-dir | a new directory under /tmp | Where the context file goes |
--k | 1 | How many runs to start at once |
--agent-model | none | The model every prompt_agent step asks for, over the step's model |
--a2a-agent-id | none | The agent every deploy_agent step deploys, over its a2a_agent_id |
--env-sandbox, --agent-sandbox | none | The sandbox for every deploy_env or deploy_agent step, over its sandbox_type |
--start-step, --context-json | 0, none | The 0-based index of the first step to run, and a JSON file to start the context from |
--litellm-api-key | none | The model API key for this run |
agent-env task run-batch --id <task id> --seeds seeds.csv runs a task once per row of a CSV file,
and each row's columns fill the task's placeholders of the same
names: a sender column fills <sender> in Reply to <sender>.. Without a sandbox option or sandbox_type, envs
and agents run in [sandbox] default and [sandbox] agent_default, both local Docker unless
config.toml says otherwise. agent-env task run --help lists the
rest.
Task run failures
A step that raises adds an entry to metadata.failed_steps. With fail_task_on_error: true, the
default, no further step starts and the record ends failed, with the error:
Running step [5/5]: env_outcome_verifier (id=check)...
Step check failed (fatal): Artifact check-reply not found
Error: Artifact check-reply not foundWith false, the steps that wait for it run anyway, so read failed_steps before you trust the
score. A step with a retry_config first runs part of the task again; see
When a step fails.
Resume a task run
A failed run leaves its env instance and agent running, so you can pick it up at the failed step.
Fix what failed, then save the failed run's context from its record, without its failed_steps:
import json
from agent_env.task import Task
context = Task.get_instance("reply-to-dana-dvbt79er").context
context["metadata"].pop("failed_steps", None)
with open("context.json", "w") as f:
json.dump(context, f, indent=2)Then run from check, index 4 of the step list:
agent-env task run --id reply-to-dana --start-step 4 --context-json context.json --output-dir out --project-id <project-id>The steps before the index are recorded as succeeded without running, and every step from it on runs again. The resumed run is a new task instance.
Last updated on