Introducing AgentEnv: An Open-Source Framework for Building RL Environments
Agents learn by practicing in environments that behave like the real world: real apps, realistic data, and tasks that range from a single question to a full simulated workplace. Scale AI built AgentEnv Framework to create these environments, and it’s now open source.
Why we built AgentEnv Framework
The problem
Creating realistic RL environments requires collaboration between researchers, engineers, and domain experts across many dimensions: artifacts, environment tools, dynamism of the environment, reproducibility, and more. There is no open source framework for building these environments effectively. Until now.
The approach
Most frameworks start from the task. AgentEnv Framework starts from the world, so one Environment can host as many Tasks and as many different agents as you like.
Can my agent… handle refund requests?
Illustrative
Open source
Build your environments with the same tools we use to build ours. The framework is on GitHub under the Apache-2.0 license.
uv tool install agentenv-frameworkagent-env run helloDesigned for interoperability
AgentEnv Framework keeps each piece a separate, versioned primitive: apps, data, the composed environment, tasks, agents and verifiers.
Each user works on their piece independently, so no one is blocked, and every task pins its versions so a rerun reproduces the original run.
Agent Agnostic
Environments expose their tools over MCP, REST and a generated CLI, and agents connect through A2A, so the same environment and task run unchanged against any agent.
Sandboxes are pluggable. We have day-one support for four sandbox providers (local Docker, Modal containers, and Modal or E2B virtual machines), and you can register your own by name. This allows us to fairly and consistently test a variety of agents against the same task in the same environment, agnostic of the underlying infrastructure.
Infrastructure Agnostic
An environment an engineer deploys on a laptop runs unchanged on a shared platform. AgentEnv Framework ships a registry for environments, tasks, data and agents, and connects to your infrastructure through a single config file. AWS and Google Cloud are supported from day one, and we partnered with Modal, the default sandbox on our own platform.
Ours runs on our internal platform, where we export hundreds of tasks at a time as self-contained bundles that other teams run on their own infrastructure.
One config file points your install at its stores. Every install that points at the same ones sees the same environments.
Dynamic and Composable
Every app, like Slack or email, is its own Environment. Snap several together and the result is still one Environment, from a support desk to an entire workplace.
Apps are built once and shared: an accounting firm, a tech company and a support desk can all use the same Slack, each loaded with its own data. Build your own with the AgentEnv Framework SDK.
Defining the Rules of the World
The real world is ever-changing, occasionally obfuscated and fundamentally non-deterministic. RL environments need to replicate these behavioral primitives while still keeping the mechanical integrity and determinism of RL training.
Virtual Clock
The Virtual Clock is the controller of time in the environment. A task sets it at the start of a run, and it can run faster than real time, up to one virtual day per second.
Controlling time means we can set the same “tomorrow” across unique runs, and an email reply that takes days finishes in seconds.
Same date, same time, in every app and for the agent.
Triggers
A trigger is a rule: when something happens in the world, something else follows. It can fire at a set time, on an agent’s action, or when the world reaches a given state.
This means we can set service “rate limits”, dynamically change the agent’s action space, have predicate-based conversations and more!
Fires when the virtual clock reaches a time you set, once or on a repeat, for work that arrives on a schedule, like a morning rush or a deadline.
EXAMPLE
The clock reaches 09:30.
A second complaint lands in the help desk.
Fires when an agent makes a matching tool call, like emailing one particular person, so people answer when they’re contacted.
EXAMPLE
The agent emails Dana.
“As Dana, reply that you were charged twice.” A world agent writes her reply.
Fires when a check on the world comes back true, however the world got there, so steps unlock once the world is ready.
EXAMPLE
Dana’s order is marked verified.
The agent can now use issue_refund.
Lives on a simulated person instead of the gateway and fires when what the agent says to them matches, so people react to what they’re told.
EXAMPLE
The agent’s message mentions store credit.
Dana says she’d rather have her money back.
RBAC
RBAC (role-based access control) gives each role in the environment its own set of tools, the way one teammate has GitHub access and another doesn’t.
On our support desk, a support agent works tickets but can’t issue refunds; a billing lead can. Tools outside a role never appear in its tool list, and any call to them is refused.
Support desk · 6 tools
- list_tickets
- read_inbox
- send_email
- post_message
- lookup_account
- issue_refund
- list_tickets
- read_inbox
- send_email
- post_message
- list_tickets
- read_inbox
- issue_refund
Blue: only the billing lead has it. The rest is shared with the support agent.
Build any task, step by step
Tasks are DAGs of primitive steps. The framework ships 49 built-in step types, from deploying an environment to grading a trajectory, so most tasks need no code. Plugins add more.
Our tasks range from a single prompt to an entire simulated workplace. The cookbooks build two from an empty folder.
Give an agent a brief and a sandbox with Blender, and it builds and renders a 3D scene. No environment or grading needed: just infrastructure and an agent.
Resolve Dana’s refund in five steps: start the support desk, load its data, bring in the agent, hand it the ticket and grade the result against a rubric.
A world that moves on its own: the clock runs, triggers change the world mid-run, tool access shifts over time and the end state is saved as a snapshot.
Anything is an Environment with Plugins
Plugins allow you to extend and build on top of AgentEnv and share with the community, turning anything an agent can act on into an environment. Watch agents play Civilization III in the OpenCiv3 RL Env and drive a real iPhone in the iOS Mobile RL Env.
Come build your first environment with AgentEnv Framework today!
No lock-in: any agent, any sandbox, any cloud
$ uv tool install agentenv-framework$ agent-env run helloPython 3.11+. Without installing: uvx --from agentenv-framework agent-env run hello. With pip: pip install agentenv-framework.