Sandbox Provider Plugins
Run envs and agents on compute of your own, here GPU VMs with a display for the OpenCiv3 client
A sandbox provider is where envs and agents run: it creates the machine, runs commands on it and
tears it down. The OpenCiv3 plugin ships none; this page writes gpu_vm, which creates GPU VMs
with a display from an HTTP VM service.
Why the Live Client Needs a GPU VM
The client that agent-env openciv3 setup --client adds renders on the CPU, at 0.8 to 1.3 s a turn, so
the live view skips turns. The openciv3_client env from Environment
Plugins boots a GPU VM image, and no built-in (local,
modal, modal_vm, e2b) boots a named VM image with a GPU: local and modal_vm ignore
create_vm(image=...), e2b raises on it, and modal has no create_vm.
So the env never deploys on [sandbox] default. On local, its environment
provider would run systemctl and write under /etc
on your own machine.
The Provider
SandboxProvider has one abstract method, create_sandbox. create_vm and get_sandbox raise
NotImplementedError until you override them, and gpu_vm overrides both: the client's
environment provider calls create_vm, and teardown calls get_sandbox.
class GpuVmSandboxProvider(SandboxProvider):
async def create_vm(self, *, image: Optional[str] = None, boot_mode: Optional[str] = None, cpu: float = 1.0,
memory: int = 8192, disk_size_gb: float = 10, timeout: int = 3600 * 2,
exposed_ports: Optional[list[int]] = None, setup_for_gateway: bool = True,
attribution: Optional[Attribution] = None, priority: Optional[int] = None,
network_policy: Optional[NetworkPolicy] = None) -> GpuVmSandbox:
refuse_unenforceable_policy(self, network_policy)
response = await self._api.request("POST", "/v1/vms", json={
"image": image, "cpu": cpu, "memory_mb": memory, "disk_gb": disk_size_gb, "ttl_seconds": timeout,
"ports": list(exposed_ports or []), "labels": dict(attribution or {}), "priority": priority,
})
vm = await self._until_running(response.json())
sandbox = GpuVmSandbox(self._api, vm, self.effective_network_policy(network_policy))
try:
if setup_for_gateway:
await sandbox.setup_vm_for_gateway(exposed_ports)
except BaseException:
await sandbox.terminate()
raise
return sandbox
async def create_sandbox(self, *, image_name: str, port: int, env: dict[str, str], cpu: float = 1.0,
memory: int = 8192, disk_size_gb: float = 10, timeout: int = 3600 * 2,
attribution: Optional[Attribution] = None, priority: Optional[int] = None,
network_policy: Optional[NetworkPolicy] = None) -> GpuVmSandbox:
return await self.create_vm(cpu=cpu, memory=memory, disk_size_gb=disk_size_gb, timeout=timeout,
exposed_ports=[port], attribution=attribution, priority=priority,
network_policy=network_policy)
async def get_sandbox(self, sandbox_id: str) -> GpuVmSandbox:
response = await self._api.request("GET", f"/v1/vms/{sandbox_id}")
return GpuVmSandbox(self._api, response.json())_until_running polls the VM every two seconds, and deletes it and raises once it reports failed
or outlives boot_timeout_seconds. The provider can't enforce an egress allowlist, so
refuse_unenforceable_policy fails before anything is created when a caller asks for one.
The provider defines the VM service's API; every request carries Authorization: Bearer <api_token>. Each request opens its own HTTP client, because core builds a new provider on every
use and never calls close().
| Request | Body or params | Response |
|---|---|---|
POST /v1/vms | {image, cpu, memory_mb, disk_gb, ttl_seconds, ports, labels, priority} | 201 {id, status, image, ports: {"<port>": url}, vnc_url, labels} |
GET /v1/vms/{id} | the same shape; status is booting, running or failed, with an optional error | |
POST /v1/vms/{id}/exec | {"argv": [...]} | {exit_code, stdout, stderr} |
PUT /v1/vms/{id}/files?path= | raw bytes | 204 |
DELETE /v1/vms/{id} | 204; a 404 counts as already gone |
Every VM on this page came from a local stand-in for that service, a real HTTP server on loopback
that boots nothing. It has no GPU, and nothing served a display at vnc_url.
The Sandbox
Sandbox has one abstract method, terminate, and exec raises NotImplementedError until you
override it. VmSandbox builds exec_script and write_host_file on top of exec.
class GpuVmSandbox(VmSandbox):
type = "gpu_vm"
def __init__(self, api: VmServiceApi, vm: dict[str, Any], network_policy: NetworkPolicy | None = None) -> None:
self._api = api
self.sandbox_id = vm["id"]
self.mode = SANDBOX_MODE_VM
self.tunnel_urls = {int(port): url for port, url in vm["ports"].items()}
self.vnc_url = vm.get("vnc_url")
self.network_policy = network_policy
async def terminate(self) -> None:
await self._api.request("DELETE", f"/v1/vms/{self.sandbox_id}", missing_ok=True)
async def exec(self, *command: str) -> _FinishedProcess:
response = await self._api.request(
"POST", f"/v1/vms/{self.sandbox_id}/exec", json={"argv": list(command)}, timeout=None,
)
result = response.json()
return _FinishedProcess(result["exit_code"], result["stdout"], result["stderr"])
async def write_host_file(self, data: bytes, vm_path: str) -> None:
"""One upload through the API, where the base class streams base64 over exec."""
await self._api.request("PUT", f"/v1/vms/{self.sandbox_id}/files", params={"path": vm_path}, content=data)exec must return an object with async stdout.read(), stderr.read() and wait(), and
_FinishedProcess holds a result the service returns once the command has exited. exec_script
runs sudo bash -c through exec and raises Script failed (exit N) on a nonzero exit.
create_sandbox -> GpuVmSandbox gpu_vm vm vm-166505b8 {18765: 'http://127.0.0.1:58000'} https://vnc.vm-service.test/vm-166505b8
exec_with_output: (0, 'Linux\n', '')
exec_script: 'hello from vm-166505b8-d473b699\n'
write_host_file stored on the VM service: b'DISPLAY=:0\n'
exec_script failure: Script failed (exit 3):VM and Container Mode
A sandbox's mode tells the caller what it got: vm is a bare machine, and container means the
image already runs on the requested port. Like modal_vm and e2b, gpu_vm returns a bare VM
from create_sandbox, ignoring image_name and env, and an agent deploy loads and runs the image
itself.
create_container is inherited: it calls create_vm, runs docker pull and docker run in the
VM, and sets mode to container. Neither that path nor an agent or gateway on gpu_vm was
exercised, since both need Docker inside the VM and the stand-in has none.
Configure It
Core resolves the env: and secret: references in the provider's config table, then calls
from_config(**config), whose default is cls(**config). gpu_vm overrides it to name the missing
keys.
@classmethod
def from_config(cls, **config: Any) -> Self:
if missing := [key for key in ("base_url", "api_token") if not config.get(key)]:
raise ConfigError(f"[sandbox.providers.gpu_vm.config] needs {' and '.join(missing)}")
return cls(**config)The table has no impl, because the entry point supplies the class and a table without one only
configures it. secret:GPU_VM_TOKEN resolves through the secret
store, so the token never sits in the file.
[sandbox.providers.gpu_vm]
config = { base_url = "env:GPU_VM_URL?https://vms.example.test", api_token = "secret:GPU_VM_TOKEN", boot_timeout_seconds = 600 }built: GpuVmSandboxProvider api token resolved from secret: True
from_config without api_token: ConfigError [sandbox.providers.gpu_vm.config] needs api_tokenRegister It
The entry point's name is the registry key, and it must equal the type of the sandboxes the
provider returns, which for GpuVmSandbox is gpu_vm. A deploy records that type next to the
sandbox id, and core reconnects later by building the provider of that name.
[project.entry-points."agent_env.sandbox_providers"]
gpu_vm = "agentenv_openciv3.gpu_sandbox:GpuVmSandboxProvider"agent-env plugin show agentenv-openciv3 then lists a sandbox provider named gpu_vm as
active, and plugin check confirms every contribution is in effect:
agent-env plugin checkok: 2 plugin package(s), 9 contribution(s), all in effectThe Type Guard
Core wraps create_sandbox, create_vm and create_container on every provider that isn't built
in. When a sandbox comes back with a type other than the provider's name, it terminates the
sandbox and raises SandboxProviderTypeError, a ConfigError. Registering the same class a
second time, as [sandbox.providers.gpu_vm_eu] with an impl, and creating a VM gives:
SandboxProviderTypeError (a ConfigError: True): [sandbox.providers.gpu_vm_eu] produced a sandbox with .type='gpu_vm'; the config name must equal the produced Sandbox.type. Rename the key to 'gpu_vm' or set the type to 'gpu_vm_eu'.
fake service log tail: ['POST /v1/vms', 'GET /v1/vms/vm-b8a5562b', 'DELETE /v1/vms/vm-b8a5562b']Reconnect and Tear Down
A teardown_sandboxes step reads each sandbox_id and sandbox_type from the env's record, calls
build_sandbox_provider(sandbox_type).get_sandbox(id).terminate(), and appends the ids to
metadata["torn_down_sandbox_ids"]. In the task run from Environment
Plugins, deploy_env recorded sandbox_type: gpu_vm and teardown
recorded ["vm-11aae0b9"].
The record has no vnc_url, so a caller that wants to watch the client calls get_sandbox the
same way and reads it from the sandbox. terminate treats a 404 as already gone, so terminating twice is
harmless:
reconnect by (id, type): GpuVmSandbox True True https://vnc.vm-service.test/vm-166505b8
terminated twice (404 on the second is fine)
get_sandbox after terminate: VM service GET /v1/vms/vm-166505b8: HTTP 404: {"error":"no such vm"}Select It
[sandbox] default and agent_default select gpu_vm for envs and agents that deploy on the
default slots. env deploy --sandbox gpu_vm, task run --env-sandbox gpu_vm and a deploy_env
step's sandbox_type hand it to one env's deploy as sandbox_type, and task run --agent-sandbox
does the same for agents.
A custom env can choose for itself, and openciv3_client takes the sandbox_type it was given, or
else the one its VM image names, never the default slot:
spec = sandbox_type or self.vm_image.sandbox_type
if not spec:
raise ValueError(f"VM image '{self.vm_image.id}' names no sandbox_type; set it, or deploy with sandbox_type")
sandbox_provider = build_sandbox_provider(spec)That task run got gpu_vm from the image alone, with no default slot set. env deploy without
--sandbox still prints Sandbox backend: config default, so read the record's sandbox_type for
where it ran; and name gpu_vm on its own, since a fallback chain has no create_vm.
Last updated on