Skip to content
AgentEnv Framework

Sandbox Provider Plugins

Run envs and agents on compute of your own, here GPU VMs with a display for the OpenCiv3 client

A sandbox provider is where envs and agents run: it creates the machine, runs commands on it and tears it down. The OpenCiv3 plugin ships none; this page writes gpu_vm, which creates GPU VMs with a display from an HTTP VM service.

Why the Live Client Needs a GPU VM

The client that agent-env openciv3 setup --client adds renders on the CPU, at 0.8 to 1.3 s a turn, so the live view skips turns. The openciv3_client env from Environment Plugins boots a GPU VM image, and no built-in (local, modal, modal_vm, e2b) boots a named VM image with a GPU: local and modal_vm ignore create_vm(image=...), e2b raises on it, and modal has no create_vm.

So the env never deploys on [sandbox] default. On local, its environment provider would run systemctl and write under /etc on your own machine.

The Provider

SandboxProvider has one abstract method, create_sandbox. create_vm and get_sandbox raise NotImplementedError until you override them, and gpu_vm overrides both: the client's environment provider calls create_vm, and teardown calls get_sandbox.

src/agentenv_openciv3/gpu_sandbox.py
class GpuVmSandboxProvider(SandboxProvider):
    async def create_vm(self, *, image: Optional[str] = None, boot_mode: Optional[str] = None, cpu: float = 1.0,
                        memory: int = 8192, disk_size_gb: float = 10, timeout: int = 3600 * 2,
                        exposed_ports: Optional[list[int]] = None, setup_for_gateway: bool = True,
                        attribution: Optional[Attribution] = None, priority: Optional[int] = None,
                        network_policy: Optional[NetworkPolicy] = None) -> GpuVmSandbox:
        refuse_unenforceable_policy(self, network_policy)
        response = await self._api.request("POST", "/v1/vms", json={
            "image": image, "cpu": cpu, "memory_mb": memory, "disk_gb": disk_size_gb, "ttl_seconds": timeout,
            "ports": list(exposed_ports or []), "labels": dict(attribution or {}), "priority": priority,
        })
        vm = await self._until_running(response.json())
        sandbox = GpuVmSandbox(self._api, vm, self.effective_network_policy(network_policy))
        try:
            if setup_for_gateway:
                await sandbox.setup_vm_for_gateway(exposed_ports)
        except BaseException:
            await sandbox.terminate()
            raise
        return sandbox

    async def create_sandbox(self, *, image_name: str, port: int, env: dict[str, str], cpu: float = 1.0,
                             memory: int = 8192, disk_size_gb: float = 10, timeout: int = 3600 * 2,
                             attribution: Optional[Attribution] = None, priority: Optional[int] = None,
                             network_policy: Optional[NetworkPolicy] = None) -> GpuVmSandbox:
        return await self.create_vm(cpu=cpu, memory=memory, disk_size_gb=disk_size_gb, timeout=timeout,
                                    exposed_ports=[port], attribution=attribution, priority=priority,
                                    network_policy=network_policy)

    async def get_sandbox(self, sandbox_id: str) -> GpuVmSandbox:
        response = await self._api.request("GET", f"/v1/vms/{sandbox_id}")
        return GpuVmSandbox(self._api, response.json())

_until_running polls the VM every two seconds, and deletes it and raises once it reports failed or outlives boot_timeout_seconds. The provider can't enforce an egress allowlist, so refuse_unenforceable_policy fails before anything is created when a caller asks for one.

The provider defines the VM service's API; every request carries Authorization: Bearer <api_token>. Each request opens its own HTTP client, because core builds a new provider on every use and never calls close().

RequestBody or paramsResponse
POST /v1/vms{image, cpu, memory_mb, disk_gb, ttl_seconds, ports, labels, priority}201 {id, status, image, ports: {"<port>": url}, vnc_url, labels}
GET /v1/vms/{id}the same shape; status is booting, running or failed, with an optional error
POST /v1/vms/{id}/exec{"argv": [...]}{exit_code, stdout, stderr}
PUT /v1/vms/{id}/files?path=raw bytes204
DELETE /v1/vms/{id}204; a 404 counts as already gone

Every VM on this page came from a local stand-in for that service, a real HTTP server on loopback that boots nothing. It has no GPU, and nothing served a display at vnc_url.

The Sandbox

Sandbox has one abstract method, terminate, and exec raises NotImplementedError until you override it. VmSandbox builds exec_script and write_host_file on top of exec.

src/agentenv_openciv3/gpu_sandbox.py
class GpuVmSandbox(VmSandbox):
    type = "gpu_vm"

    def __init__(self, api: VmServiceApi, vm: dict[str, Any], network_policy: NetworkPolicy | None = None) -> None:
        self._api = api
        self.sandbox_id = vm["id"]
        self.mode = SANDBOX_MODE_VM
        self.tunnel_urls = {int(port): url for port, url in vm["ports"].items()}
        self.vnc_url = vm.get("vnc_url")
        self.network_policy = network_policy

    async def terminate(self) -> None:
        await self._api.request("DELETE", f"/v1/vms/{self.sandbox_id}", missing_ok=True)

    async def exec(self, *command: str) -> _FinishedProcess:
        response = await self._api.request(
            "POST", f"/v1/vms/{self.sandbox_id}/exec", json={"argv": list(command)}, timeout=None,
        )
        result = response.json()
        return _FinishedProcess(result["exit_code"], result["stdout"], result["stderr"])

    async def write_host_file(self, data: bytes, vm_path: str) -> None:
        """One upload through the API, where the base class streams base64 over exec."""
        await self._api.request("PUT", f"/v1/vms/{self.sandbox_id}/files", params={"path": vm_path}, content=data)

exec must return an object with async stdout.read(), stderr.read() and wait(), and _FinishedProcess holds a result the service returns once the command has exited. exec_script runs sudo bash -c through exec and raises Script failed (exit N) on a nonzero exit.

Output
create_sandbox -> GpuVmSandbox gpu_vm vm vm-166505b8 {18765: 'http://127.0.0.1:58000'} https://vnc.vm-service.test/vm-166505b8
exec_with_output: (0, 'Linux\n', '')
exec_script: 'hello from vm-166505b8-d473b699\n'
write_host_file stored on the VM service: b'DISPLAY=:0\n'
exec_script failure: Script failed (exit 3):

VM and Container Mode

A sandbox's mode tells the caller what it got: vm is a bare machine, and container means the image already runs on the requested port. Like modal_vm and e2b, gpu_vm returns a bare VM from create_sandbox, ignoring image_name and env, and an agent deploy loads and runs the image itself.

create_container is inherited: it calls create_vm, runs docker pull and docker run in the VM, and sets mode to container. Neither that path nor an agent or gateway on gpu_vm was exercised, since both need Docker inside the VM and the stand-in has none.

Configure It

Core resolves the env: and secret: references in the provider's config table, then calls from_config(**config), whose default is cls(**config). gpu_vm overrides it to name the missing keys.

src/agentenv_openciv3/gpu_sandbox.py
    @classmethod
    def from_config(cls, **config: Any) -> Self:
        if missing := [key for key in ("base_url", "api_token") if not config.get(key)]:
            raise ConfigError(f"[sandbox.providers.gpu_vm.config] needs {' and '.join(missing)}")
        return cls(**config)

The table has no impl, because the entry point supplies the class and a table without one only configures it. secret:GPU_VM_TOKEN resolves through the secret store, so the token never sits in the file.

.agentenv/config.toml
[sandbox.providers.gpu_vm]
config = { base_url = "env:GPU_VM_URL?https://vms.example.test", api_token = "secret:GPU_VM_TOKEN", boot_timeout_seconds = 600 }
Output
built: GpuVmSandboxProvider api token resolved from secret: True
from_config without api_token: ConfigError [sandbox.providers.gpu_vm.config] needs api_token

Register It

The entry point's name is the registry key, and it must equal the type of the sandboxes the provider returns, which for GpuVmSandbox is gpu_vm. A deploy records that type next to the sandbox id, and core reconnects later by building the provider of that name.

pyproject.toml
[project.entry-points."agent_env.sandbox_providers"]
gpu_vm = "agentenv_openciv3.gpu_sandbox:GpuVmSandboxProvider"

agent-env plugin show agentenv-openciv3 then lists a sandbox provider named gpu_vm as active, and plugin check confirms every contribution is in effect:

Terminal
agent-env plugin check
Output
ok: 2 plugin package(s), 9 contribution(s), all in effect

The Type Guard

Core wraps create_sandbox, create_vm and create_container on every provider that isn't built in. When a sandbox comes back with a type other than the provider's name, it terminates the sandbox and raises SandboxProviderTypeError, a ConfigError. Registering the same class a second time, as [sandbox.providers.gpu_vm_eu] with an impl, and creating a VM gives:

Output
SandboxProviderTypeError (a ConfigError: True): [sandbox.providers.gpu_vm_eu] produced a sandbox with .type='gpu_vm'; the config name must equal the produced Sandbox.type. Rename the key to 'gpu_vm' or set the type to 'gpu_vm_eu'.
fake service log tail: ['POST /v1/vms', 'GET /v1/vms/vm-b8a5562b', 'DELETE /v1/vms/vm-b8a5562b']

Reconnect and Tear Down

A teardown_sandboxes step reads each sandbox_id and sandbox_type from the env's record, calls build_sandbox_provider(sandbox_type).get_sandbox(id).terminate(), and appends the ids to metadata["torn_down_sandbox_ids"]. In the task run from Environment Plugins, deploy_env recorded sandbox_type: gpu_vm and teardown recorded ["vm-11aae0b9"].

The record has no vnc_url, so a caller that wants to watch the client calls get_sandbox the same way and reads it from the sandbox. terminate treats a 404 as already gone, so terminating twice is harmless:

Output
reconnect by (id, type): GpuVmSandbox True True https://vnc.vm-service.test/vm-166505b8
terminated twice (404 on the second is fine)
get_sandbox after terminate: VM service GET /v1/vms/vm-166505b8: HTTP 404: {"error":"no such vm"}

Select It

[sandbox] default and agent_default select gpu_vm for envs and agents that deploy on the default slots. env deploy --sandbox gpu_vm, task run --env-sandbox gpu_vm and a deploy_env step's sandbox_type hand it to one env's deploy as sandbox_type, and task run --agent-sandbox does the same for agents.

A custom env can choose for itself, and openciv3_client takes the sandbox_type it was given, or else the one its VM image names, never the default slot:

src/agentenv_openciv3/client_env.py
        spec = sandbox_type or self.vm_image.sandbox_type
        if not spec:
            raise ValueError(f"VM image '{self.vm_image.id}' names no sandbox_type; set it, or deploy with sandbox_type")
        sandbox_provider = build_sandbox_provider(spec)

That task run got gpu_vm from the image alone, with no default slot set. env deploy without --sandbox still prints Sandbox backend: config default, so read the record's sandbox_type for where it ran; and name gpu_vm on its own, since a fallback chain has no create_vm.

Last updated on

Ask a question · Report an issue

On this page