# Sandbox Provider Plugins (https://www.agentenvframework.com/docs/plugins/sandbox-provider-plugins)

> Run envs and agents on compute of your own, here GPU VMs with a display for the OpenCiv3 client

A sandbox provider is where envs and agents run: it creates the machine, runs commands on it and
tears it down. The OpenCiv3 plugin ships none; this page writes `gpu_vm`, which creates GPU VMs
with a display from an HTTP VM service.

## Why the Live Client Needs a GPU VM

The client that `agent-env openciv3 setup --client` adds renders on the CPU, at 0.8 to 1.3 s a turn, so
the live view skips turns. The `openciv3_client` env from [Environment
Plugins](https://www.agentenvframework.com/docs/plugins/environment-plugins.md) boots a GPU VM image, and no built-in (`local`,
`modal`, `modal_vm`, `e2b`) boots a named VM image with a GPU: `local` and `modal_vm` ignore
`create_vm(image=...)`, `e2b` raises on it, and `modal` has no `create_vm`.

So the env never deploys on `[sandbox] default`. On `local`, its [environment
provider](https://www.agentenvframework.com/docs/plugins/environment-provider-plugins.md) would run `systemctl` and write under `/etc`
on your own machine.

## The Provider

`SandboxProvider` has one abstract method, `create_sandbox`. `create_vm` and `get_sandbox` raise
`NotImplementedError` until you override them, and `gpu_vm` overrides both: the client's
environment provider calls `create_vm`, and teardown calls `get_sandbox`.

```python title="src/agentenv_openciv3/gpu_sandbox.py"
class GpuVmSandboxProvider(SandboxProvider):
    async def create_vm(self, *, image: Optional[str] = None, boot_mode: Optional[str] = None, cpu: float = 1.0,
                        memory: int = 8192, disk_size_gb: float = 10, timeout: int = 3600 * 2,
                        exposed_ports: Optional[list[int]] = None, setup_for_gateway: bool = True,
                        attribution: Optional[Attribution] = None, priority: Optional[int] = None,
                        network_policy: Optional[NetworkPolicy] = None) -> GpuVmSandbox:
        refuse_unenforceable_policy(self, network_policy)
        response = await self._api.request("POST", "/v1/vms", json={
            "image": image, "cpu": cpu, "memory_mb": memory, "disk_gb": disk_size_gb, "ttl_seconds": timeout,
            "ports": list(exposed_ports or []), "labels": dict(attribution or {}), "priority": priority,
        })
        vm = await self._until_running(response.json())
        sandbox = GpuVmSandbox(self._api, vm, self.effective_network_policy(network_policy))
        try:
            if setup_for_gateway:
                await sandbox.setup_vm_for_gateway(exposed_ports)
        except BaseException:
            await sandbox.terminate()
            raise
        return sandbox

    async def create_sandbox(self, *, image_name: str, port: int, env: dict[str, str], cpu: float = 1.0,
                             memory: int = 8192, disk_size_gb: float = 10, timeout: int = 3600 * 2,
                             attribution: Optional[Attribution] = None, priority: Optional[int] = None,
                             network_policy: Optional[NetworkPolicy] = None) -> GpuVmSandbox:
        return await self.create_vm(cpu=cpu, memory=memory, disk_size_gb=disk_size_gb, timeout=timeout,
                                    exposed_ports=[port], attribution=attribution, priority=priority,
                                    network_policy=network_policy)

    async def get_sandbox(self, sandbox_id: str) -> GpuVmSandbox:
        response = await self._api.request("GET", f"/v1/vms/{sandbox_id}")
        return GpuVmSandbox(self._api, response.json())
```

`_until_running` polls the VM every two seconds, and deletes it and raises once it reports `failed`
or outlives `boot_timeout_seconds`. The provider can't enforce an egress allowlist, so
`refuse_unenforceable_policy` fails before anything is created when a caller asks for one.

The provider defines the VM service's API; every request carries `Authorization: Bearer <api_token>`. Each request opens its own HTTP client, because core builds a new provider on every
use and never calls `close()`.

| Request                        | Body or params                                                           | Response                                                                               |
| ------------------------------ | ------------------------------------------------------------------------ | -------------------------------------------------------------------------------------- |
| `POST /v1/vms`                 | `{image, cpu, memory_mb, disk_gb, ttl_seconds, ports, labels, priority}` | 201 `{id, status, image, ports: {"<port>": url}, vnc_url, labels}`                     |
| `GET /v1/vms/{id}`             |                                                                          | the same shape; `status` is `booting`, `running` or `failed`, with an optional `error` |
| `POST /v1/vms/{id}/exec`       | `{"argv": [...]}`                                                        | `{exit_code, stdout, stderr}`                                                          |
| `PUT /v1/vms/{id}/files?path=` | raw bytes                                                                | 204                                                                                    |
| `DELETE /v1/vms/{id}`          |                                                                          | 204; a 404 counts as already gone                                                      |

Every VM on this page came from a local stand-in for that service, a real HTTP server on loopback
that boots nothing. It has no GPU, and nothing served a display at `vnc_url`.

## The Sandbox

`Sandbox` has one abstract method, `terminate`, and `exec` raises `NotImplementedError` until you
override it. `VmSandbox` builds `exec_script` and `write_host_file` on top of `exec`.

```python title="src/agentenv_openciv3/gpu_sandbox.py"
class GpuVmSandbox(VmSandbox):
    type = "gpu_vm"

    def __init__(self, api: VmServiceApi, vm: dict[str, Any], network_policy: NetworkPolicy | None = None) -> None:
        self._api = api
        self.sandbox_id = vm["id"]
        self.mode = SANDBOX_MODE_VM
        self.tunnel_urls = {int(port): url for port, url in vm["ports"].items()}
        self.vnc_url = vm.get("vnc_url")
        self.network_policy = network_policy

    async def terminate(self) -> None:
        await self._api.request("DELETE", f"/v1/vms/{self.sandbox_id}", missing_ok=True)

    async def exec(self, *command: str) -> _FinishedProcess:
        response = await self._api.request(
            "POST", f"/v1/vms/{self.sandbox_id}/exec", json={"argv": list(command)}, timeout=None,
        )
        result = response.json()
        return _FinishedProcess(result["exit_code"], result["stdout"], result["stderr"])

    async def write_host_file(self, data: bytes, vm_path: str) -> None:
        """One upload through the API, where the base class streams base64 over exec."""
        await self._api.request("PUT", f"/v1/vms/{self.sandbox_id}/files", params={"path": vm_path}, content=data)
```

`exec` must return an object with async `stdout.read()`, `stderr.read()` and `wait()`, and
`_FinishedProcess` holds a result the service returns once the command has exited. `exec_script`
runs `sudo bash -c` through `exec` and raises `Script failed (exit N)` on a nonzero exit.

```text title="Output"
create_sandbox -> GpuVmSandbox gpu_vm vm vm-166505b8 {18765: 'http://127.0.0.1:58000'} https://vnc.vm-service.test/vm-166505b8
exec_with_output: (0, 'Linux\n', '')
exec_script: 'hello from vm-166505b8-d473b699\n'
write_host_file stored on the VM service: b'DISPLAY=:0\n'
exec_script failure: Script failed (exit 3):
```

## VM and Container Mode

A sandbox's `mode` tells the caller what it got: `vm` is a bare machine, and `container` means the
image already runs on the requested port. Like `modal_vm` and `e2b`, `gpu_vm` returns a bare VM
from `create_sandbox`, ignoring `image_name` and `env`, and an agent deploy loads and runs the image
itself.

`create_container` is inherited: it calls `create_vm`, runs `docker pull` and `docker run` in the
VM, and sets `mode` to `container`. Neither that path nor an agent or gateway on `gpu_vm` was
exercised, since both need Docker inside the VM and the stand-in has none.

## Configure It

Core resolves the `env:` and `secret:` references in the provider's config table, then calls
`from_config(**config)`, whose default is `cls(**config)`. `gpu_vm` overrides it to name the missing
keys.

```python title="src/agentenv_openciv3/gpu_sandbox.py"
    @classmethod
    def from_config(cls, **config: Any) -> Self:
        if missing := [key for key in ("base_url", "api_token") if not config.get(key)]:
            raise ConfigError(f"[sandbox.providers.gpu_vm.config] needs {' and '.join(missing)}")
        return cls(**config)
```

The table has no `impl`, because the entry point supplies the class and a table without one only
configures it. `secret:GPU_VM_TOKEN` resolves through the [secret
store](https://www.agentenvframework.com/docs/registry/secret-store.md#references), so the token never sits in the file.

```toml title=".agentenv/config.toml"
[sandbox.providers.gpu_vm]
config = { base_url = "env:GPU_VM_URL?https://vms.example.test", api_token = "secret:GPU_VM_TOKEN", boot_timeout_seconds = 600 }
```

```text title="Output"
built: GpuVmSandboxProvider api token resolved from secret: True
from_config without api_token: ConfigError [sandbox.providers.gpu_vm.config] needs api_token
```

## Register It

The entry point's name is the registry key, and it must equal the `type` of the sandboxes the
provider returns, which for `GpuVmSandbox` is `gpu_vm`. A deploy records that type next to the
sandbox id, and core reconnects later by building the provider of that name.

```toml title="pyproject.toml"
[project.entry-points."agent_env.sandbox_providers"]
gpu_vm = "agentenv_openciv3.gpu_sandbox:GpuVmSandboxProvider"
```

`agent-env plugin show agentenv-openciv3` then lists a `sandbox provider` named `gpu_vm` as
`active`, and `plugin check` confirms every contribution is in effect:

```bash title="Terminal"
agent-env plugin check
```

```text title="Output"
ok: 2 plugin package(s), 9 contribution(s), all in effect
```

## The Type Guard

Core wraps `create_sandbox`, `create_vm` and `create_container` on every provider that isn't built
in. When a sandbox comes back with a `type` other than the provider's name, it terminates the
sandbox and raises `SandboxProviderTypeError`, a `ConfigError`. Registering the same class a
second time, as `[sandbox.providers.gpu_vm_eu]` with an `impl`, and creating a VM gives:

```text title="Output"
SandboxProviderTypeError (a ConfigError: True): [sandbox.providers.gpu_vm_eu] produced a sandbox with .type='gpu_vm'; the config name must equal the produced Sandbox.type. Rename the key to 'gpu_vm' or set the type to 'gpu_vm_eu'.
fake service log tail: ['POST /v1/vms', 'GET /v1/vms/vm-b8a5562b', 'DELETE /v1/vms/vm-b8a5562b']
```

## Reconnect and Tear Down

A `teardown_sandboxes` step reads each `sandbox_id` and `sandbox_type` from the env's record, calls
`build_sandbox_provider(sandbox_type).get_sandbox(id).terminate()`, and appends the ids to
`metadata["torn_down_sandbox_ids"]`. In the task run from [Environment
Plugins](https://www.agentenvframework.com/docs/plugins/environment-plugins.md), `deploy_env` recorded `sandbox_type: gpu_vm` and teardown
recorded `["vm-11aae0b9"]`.

The record has no `vnc_url`, so a caller that wants to watch the client calls `get_sandbox` the
same way and reads it from the sandbox. `terminate` treats a 404 as already gone, so terminating twice is
harmless:

```text title="Output"
reconnect by (id, type): GpuVmSandbox True True https://vnc.vm-service.test/vm-166505b8
terminated twice (404 on the second is fine)
get_sandbox after terminate: VM service GET /v1/vms/vm-166505b8: HTTP 404: {"error":"no such vm"}
```

## Select It

`[sandbox] default` and `agent_default` select `gpu_vm` for envs and agents that deploy on the
default slots. `env deploy --sandbox gpu_vm`, `task run --env-sandbox gpu_vm` and a `deploy_env`
step's `sandbox_type` hand it to one env's `deploy` as `sandbox_type`, and `task run --agent-sandbox`
does the same for agents.

A custom env can choose for itself, and `openciv3_client` takes the `sandbox_type` it was given, or
else the one its VM image names, never the default slot:

```python title="src/agentenv_openciv3/client_env.py"
        spec = sandbox_type or self.vm_image.sandbox_type
        if not spec:
            raise ValueError(f"VM image '{self.vm_image.id}' names no sandbox_type; set it, or deploy with sandbox_type")
        sandbox_provider = build_sandbox_provider(spec)
```

That task run got `gpu_vm` from the image alone, with no default slot set. `env deploy` without
`--sandbox` still prints `Sandbox backend: config default`, so read the record's `sandbox_type` for
where it ran; and name `gpu_vm` on its own, since a fallback chain has no `create_vm`.