Skip to content
AgentEnv Framework

Running a task

A task, once designed, can be run many times

Running a task

agent-env task run loads a stored task and runs it in the terminal's own process:

Terminal
agent-env task run --id reply-to-dana --output-dir out --project-id <project-id>
Output (banner and most step results left out)
Fetching task: id=reply-to-dana version=latest...
Found task: id=reply-to-dana version=2 steps=5
Task instance: reply-to-dana-gbgs1yub
Running step [1/5]: deploy_env (id=email)...
Completed step [1/5]: deploy_env (id=email) [42.9s]

Running step [2/5]: load_artifact (id=inbox)...

Running step [3/5]: deploy_agent (id=assistant)...
Completed step [2/5]: load_artifact (id=inbox) [1.9s]
Completed step [3/5]: deploy_agent (id=assistant) [23.7s]

Running step [4/5]: prompt_agent (id=reply)...
Completed step [4/5]: prompt_agent (id=reply) [63.3s]

Running step [5/5]: env_outcome_verifier (id=check)...
Completed step [5/5]: env_outcome_verifier (id=check) [0.4s]
score (verifier_id=sent): 1.0

Task completed!
Task context written to: out/reply-to-dana_811f14b9.json

The run loads the latest version unless you pass --version. A step starts as soon as the steps in its depends_on have finished. Runs of the same task go at once with --k, each with its own task instance and context:

The TaskStepContext

The TaskStepContext is a ledger of what has happened during a task run. When the run completes, the command writes it to --output-dir as <task id>_<8 hex>.json, with secrets such as litellm_api_key removed.

The task run instance

Each run is a task instance, with a record in the document store: the task id and version, status, and current_step of total_steps. After each step the run writes to the record only what that step changed, so steps that finish together don't overwrite each other. agent-env task get-instance --id reply-to-dana-gbgs1yub reads it back, during the run or after it:

Task run options

OptionDefaultMeaning
--versionthe latestThe task version to run
--output-dira new directory under /tmpWhere the context file goes
--k1How many runs to start at once
--agent-modelnoneThe model every prompt_agent step asks for, over the step's model
--a2a-agent-idnoneThe agent every deploy_agent step deploys, over its a2a_agent_id
--env-sandbox, --agent-sandboxnoneThe sandbox for every deploy_env or deploy_agent step, over its sandbox_type
--start-step, --context-json0, noneThe 0-based index of the first step to run, and a JSON file to start the context from
--litellm-api-keynoneThe model API key for this run

agent-env task run-batch --id <task id> --seeds seeds.csv runs a task once per row of a CSV file, and each row's columns fill the task's placeholders of the same names: a sender column fills <sender> in Reply to <sender>.. Without a sandbox option or sandbox_type, envs and agents run in [sandbox] default and [sandbox] agent_default, both local Docker unless config.toml says otherwise. agent-env task run --help lists the rest.

Task run failures

A step that raises adds an entry to metadata.failed_steps. With fail_task_on_error: true, the default, no further step starts and the record ends failed, with the error:

Output (last lines)
Running step [5/5]: env_outcome_verifier (id=check)...
Step check failed (fatal): Artifact check-reply not found
Error: Artifact check-reply not found

With false, the steps that wait for it run anyway, so read failed_steps before you trust the score. A step with a retry_config first runs part of the task again; see When a step fails.

Resume a task run

A failed run leaves its env instance and agent running, so you can pick it up at the failed step. Fix what failed, then save the failed run's context from its record, without its failed_steps:

save_context.py
import json

from agent_env.task import Task

context = Task.get_instance("reply-to-dana-dvbt79er").context
context["metadata"].pop("failed_steps", None)
with open("context.json", "w") as f:
    json.dump(context, f, indent=2)

Then run from check, index 4 of the step list:

Terminal
agent-env task run --id reply-to-dana --start-step 4 --context-json context.json --output-dir out --project-id <project-id>

The steps before the index are recorded as succeeded without running, and every step from it on runs again. The resumed run is a new task instance.

Last updated on

Ask a question · Report an issue

On this page