# isobox > isobox is an open-source HTTP sandbox that runs untrusted / AI-generated code in gVisor-isolated, resource-capped, ephemeral containers. It is security-first and built for AI agents: POST code to `/execute` and get back stdout/stderr/exitCode, optionally streamed over SSE. Base URL (live): `https://isobox.hsingh.app` — self-host default: `http://127.0.0.1:8090`. ## Endpoints - `GET /` — AI chat + code-interpreter UI (HTML). - `GET /healthz` — liveness. `200 application/json` body `{"status":"ok","version":""}` (`version` is the build-stamped isoboxd version; `dev` if unstamped). When kernel sessions are enabled it also carries `"kernels":{"used":N,"max":M}` (resident-kernel slot occupancy). - `GET /readyz` — readiness. `200 {"ready":true,"backend":"gvisor"}` or `503 {"ready":false,"reason":"..."}`. Ready only if the protective cgroup slice is memory-capped AND the backend is healthy. - `GET /runtimes` — list runtimes. `200 [{"language","version","aliases":[...],"compiled":bool,"backend":"gvisor"}]`. - `POST /execute` — run code. Synchronous JSON, or live SSE if `Accept: text/event-stream`. - `GET /llms.txt` — this document. `text/plain`. - `GET /openapi.yaml` — OpenAPI spec. `application/yaml`. ## Auth Optional. If the server has an API key configured, send it as `X-API-Key: ` OR `Authorization: Bearer `. If no key is configured (e.g. the public demo), the endpoint is open and only rate limiting applies. A bad/missing key when one is configured returns `401 {"error":"unauthorized"}`. ## POST /execute — Request `Content-Type: application/json`. Request body hard cap: 256 KiB. ```json { "language": "python", "version": "3.14-agent", "code": "print(1)", "files": [{"name": "main.py", "content": "...", "encoding": "utf8"}], "stdin": "", "args": [], "limits": { "memoryBytes": 268435456, "cpus": 1.0, "pids": 128, "outputBytes": 65536, "wallTimeMs": 10000 }, "network": true } ``` Fields: - `language` (string, REQUIRED) — name or alias; see Languages / `GET /runtimes`. - `version` (string, optional) — pins a specific version; otherwise the default for that language. - `code` (string) — single-file convenience source. Provide `code` OR `files`. At least one is REQUIRED. - `files` (array, optional) — explicit multi-file source. Each item: `name` (string), `content` (string), `encoding` (one of `utf8` (default) | `base64` | `hex`). - `stdin` (string, optional) — fed to the program's stdin. - `args` (array of strings, optional) — program argv; passed safely as positional parameters. - `limits` (object, optional) — all sub-fields optional and individually CLAMPED to hard ceilings. When a field is omitted the per-language default applies (defaults vary by runtime — e.g. python defaults to 384 MiB / 1.5 CPU / 15000 ms wall time; most other runtimes to 256 MiB / 1.0 CPU / 10000 ms; compiled go/rust to 512 MiB / 2.0 CPU / 20000 ms — `GET /runtimes` plus the bounds below describe the envelope): - `memoryBytes` (int) — min 8388608 (8 MiB), max 536870912 (512 MiB). - `cpus` (float) — min 0.1, max 2.0. - `pids` (int) — default 128 (all runtimes). Min 1, max 256. - `outputBytes` (int) — per-stream cap (stdout and stderr are each capped independently at this value, so total output can be up to 2× this). Default 65536 (64 KiB). Min 1024 (1 KiB), max 262144 (256 KiB). - `wallTimeMs` (int) — min 100, max 20000. - `network` (bool, optional) — default `true` (omitted = ON): FILTERED egress (public internet only). Pass `false` explicitly to run with no network. See Network mode. Values outside a limit's range are clamped to the nearest bound, not rejected. ## POST /execute — Response (200, application/json) ```json { "language": "python", "version": "3.14-agent", "backend": "gvisor", "run": { "stdout": "...", "stderr": "...", "exitCode": 0, "timedOut": false, "oomKilled": false, "truncated": false, "durationMs": 214, "network": false }, "warning": "network requested but currently unavailable (egress firewall not verified); ran with network OFF" } ``` Fields: - `language`, `version`, `backend` (string) — resolved language, resolved version, execution backend (`gvisor`). - `run.stdout`, `run.stderr` (string) — captured output (each stream capped at `outputBytes`). - `run.exitCode` (int) — process exit code. A wall-time or memory kill yields `137`. - `run.timedOut` (bool) — `true` if killed by the wall-time cap. - `run.oomKilled` (bool) — `true` if killed by the memory cap. Disambiguates timeout vs OOM (both surface exitCode 137). - `run.truncated` (bool) — `true` if output hit the `outputBytes` cap. - `run.durationMs` (int) — wall-clock execution time in milliseconds. - `run.network` (bool) — whether filtered egress was ACTUALLY applied. May be `false` even when `network:true` was requested (fail-closed). - `warning` (string, optional) — present ONLY when network was requested (or defaulted on) but egress was denied. Exact text: `network requested but currently unavailable (egress firewall not verified); ran with network OFF`. ## Streaming (SSE) Send header `Accept: text/event-stream` (alternatively `?stream=1`). The response is a stream of Server-Sent Events. Event names and payloads: ``` event: stdout data: {"chunk":"he"} event: stderr data: {"chunk":"oops\n"} event: done data: {"language":"python","version":"...","backend":"gvisor","run":{...},"warning":"..."} event: error data: {"error":"..."} ``` - `stdout` / `stderr` — incremental output chunks; `data` is `{"chunk": ""}`. - `done` — final structured result; `data` is the same object as the non-streaming 200 response (`language`, `version`, `backend`, `run`, and `warning` when applicable). - `error` — emitted only on internal failure; `data` is `{"error": ""}`. The HTTP status for an SSE response is always `200`; per-run outcomes (non-zero exit, timeout, OOM) are reported inside the `done` event's `run` object, not via HTTP status. ## Errors (non-2xx, application/json) - `400 {"error":"invalid_json","detail":"..."}` — malformed body or body exceeds 256 KiB. - `400 {"error":"language_required"}` — `language` missing/empty. - `400 {"error":"unknown_language","detail":"\"x\" is not a known runtime; see GET /runtimes"}`. - `400 {"error":"no_source","detail":"provide `code` or `files`"}` — neither `code` nor `files` supplied. - `401 {"error":"unauthorized"}` — only when an API key is configured and the request key is wrong/missing. - `429 {"error":"rate_limited"}` — per-IP token bucket (default 30/min, burst 10). Includes a `Retry-After` header. - `429 {"error":"capacity","detail":"no free execution slot"}` — global concurrency cap reached (default 4). Includes a `Retry-After` header. ## Languages Use `GET /runtimes` for the live list. Current runtimes (language, default version, aliases): - `python` — `3.14-agent` — aliases `py`, `py3`, `python3`. Batteries-included: `requests`, `httpx`, `beautifulsoup4`, `lxml`, `pandas`, `numpy`, `matplotlib`, `yt-dlp`. - `python-slim` — `3.14.5` — alias `pyslim`. - `javascript` — `node-26` — aliases `js`, `node`, `nodejs`. - `typescript` — `node-26` — alias `ts`. - `ruby` — `3-alpine` — alias `rb`. - `go` — `1.24-warm` — alias `golang`. Compiled, pre-warmed. - `rust` — `1-alpine` — alias `rs`. Compiled. Pass the language name or any alias in the request's `language` field. ## Network mode `"network": true` by default (an omitted `network` field means ON; pass `"network": false` explicitly to run with no network). With network on, the code reaches the PUBLIC internet ONLY. Firewalled off: cloud metadata (`169.254.169.254`), all private / RFC1918 ranges, the host (including any VPN), outbound SMTP, IPv6, and other sandboxes. It fails CLOSED: if the egress firewall is not currently verified, the run executes with network OFF and the response reports `run.network: false` plus the `warning` field. Always check `run.network` to confirm egress was actually applied before assuming a fetch could have succeeded. ## Examples ```bash # simple — synchronous JSON curl -s https://isobox.hsingh.app/execute \ -H 'content-type: application/json' \ -d '{"language":"python","code":"print(40+2)"}' # streaming — live SSE output curl -N https://isobox.hsingh.app/execute \ -H 'accept: text/event-stream' \ -H 'content-type: application/json' \ -d '{"language":"python","code":"import time\nfor i in range(3):\n print(i,flush=True);time.sleep(.3)"}' # network fetch / scrape — filtered egress is on by default (check run.network in the response) curl -s https://isobox.hsingh.app/execute \ -H 'content-type: application/json' \ -d '{"language":"python","code":"import requests;print(requests.get(\"https://api.github.com\").status_code)"}' # offline run — opt out of egress explicitly curl -s https://isobox.hsingh.app/execute \ -H 'content-type: application/json' \ -d '{"language":"python","network":false,"code":"print(40+2)"}' # limits override — clamped to hard ceilings curl -s https://isobox.hsingh.app/execute \ -H 'content-type: application/json' \ -d '{"language":"python","code":"print(\"hi\")","limits":{"memoryBytes":134217728,"cpus":0.5,"pids":64,"outputBytes":131072,"wallTimeMs":5000}}' ``` ## Stateful sessions (v1) For multi-step / agent workflows that need to share state. A *filesystem session* is a persistent `/workspace` directory shared across steps; each step still runs in a fresh hardened sandbox (so a session at rest costs no RAM). Files written to `/workspace` (or the session's working directory) in one step are visible to the next step and to the file API. Every `/v1/sessions/{id}/*` call requires the capability token returned at create, as header `X-Session-Token: `. - `POST /v1/sessions` — create. Body `{"runtime":"python","ttlSec":86400}` (runtime required; ttlSec optional). Returns `201 {"id","token","runtime","type","createdAt"}`. Save the `token` — it is shown only once and is required on every later call. - `POST /v1/sessions/{id}/exec` — run one step. Body is the same as `/execute` MINUS `language` (the session's runtime is used). The step runs with `/workspace` mounted read-write and as its working directory. Sync JSON, or SSE with `Accept: text/event-stream` (same event shape as `/execute`). `network` works the same as `/execute` (on by default; `false` opts out). - `GET /v1/sessions/{id}/fs?path=` — if `path` is a dir, returns `{"path","dir":true,"entries":[{"name","size","dir","modTime"}]}`; if a file, returns the raw bytes (`application/octet-stream`, with a `Content-Disposition: attachment` header so a browser downloads it by name). Add `&format=zip` on a dir to stream it as a zip archive (`application/zip`, `Content-Disposition: attachment`, capped at 256 MiB); on a file it `400`s. - `PUT /v1/sessions/{id}/fs?path=` — upload a file (raw body, up to 16 MiB, streamed to disk not buffered). Returns `{"path","bytes"}`. Counts against the per-session disk quota. - `DELETE /v1/sessions/{id}/fs?path=` — delete a file/subtree. `204`. - `DELETE /v1/sessions/{id}` — end the session and delete its workspace. `204`. Limits & errors: per-session disk quota (default 512 MiB, enforced against actual usage); a session/global storage ceiling and max-session count may return `429 {"error":"too_many_sessions"}` or `507 {"error":"storage_full"}` at create. Idle sessions are reaped after their TTL (default 24h). `404 session_not_found`, `401 invalid_session_token`, `413 quota_exceeded`, `400 invalid_path` as applicable. Paths are confined to the workspace (traversal is collapsed inside it). ```bash # create a session S=$(curl -s -X POST https://isobox.hsingh.app/v1/sessions -d '{"runtime":"python"}') ID=$(echo "$S" | jq -r .id); TOK=$(echo "$S" | jq -r .token) # step 1 writes state, step 2 (a separate sandbox) reads it curl -s -X POST https://isobox.hsingh.app/v1/sessions/$ID/exec -H "X-Session-Token: $TOK" \ -d '{"code":"open(\"/workspace/x.txt\",\"w\").write(\"42\")"}' curl -s -X POST https://isobox.hsingh.app/v1/sessions/$ID/exec -H "X-Session-Token: $TOK" \ -d '{"code":"print(open(\"/workspace/x.txt\").read())"}' # -> 42 # download a file the code produced, then end the session curl -s "https://isobox.hsingh.app/v1/sessions/$ID/fs?path=x.txt" -H "X-Session-Token: $TOK" curl -s -X DELETE https://isobox.hsingh.app/v1/sessions/$ID -H "X-Session-Token: $TOK" ``` ## Persistent memory (v1) Two durable tiers an agent can store and recall across sessions. In open/demo mode (no API key) everything is a single shared "public" tenant; configure `ISOBOX_API_KEYS` to get one isolated tenant per key (tenant is derived server-side from the key, never from the request). ### Structured KV — `/v1/memory` - `PUT /v1/memory/{namespace}/{key}` — store the raw request body. Optional `X-TTL-Seconds: ` for expiry. `204`. Per-tenant quota 10 MiB / 10k keys. - `GET /v1/memory/{namespace}/{key}` — returns the value (with `ETag`, `X-Expires-At`), or `404` if absent/expired. - `DELETE /v1/memory/{namespace}/{key}` — `204`. - `GET /v1/memory/{namespace}?prefix=&limit=&cursor=` — list keys. ### Filesystem volumes — `/v1/volumes` (long-term `/memory`) A named directory that survives sessions and re-attaches to new ones. - `POST /v1/volumes` — body `{"name":"..."}` → `{"id","name","createdAt",...}`. - `GET /v1/volumes` (list, tenant-scoped) · `GET /v1/volumes/{id}` · `DELETE /v1/volumes/{id}`. - Attach to a session by passing `"volumeId":""` at `POST /v1/sessions`; it is mounted **read-write at `/memory`** for every step. Only one session may hold a volume RW at a time (others get `409`). Per-volume quota 512 MiB. ```bash # remember across sessions via a volume VID=$(curl -s -X POST https://isobox.hsingh.app/v1/volumes -d '{"name":"mem"}' | jq -r .id) S=$(curl -s -X POST https://isobox.hsingh.app/v1/sessions -d "{\"runtime\":\"python\",\"volumeId\":\"$VID\"}") ID=$(echo "$S"|jq -r .id); TOK=$(echo "$S"|jq -r .token) curl -s -X POST https://isobox.hsingh.app/v1/sessions/$ID/exec -H "X-Session-Token: $TOK" \ -d '{"code":"open(\"/memory/note.txt\",\"w\").write(\"recall me later\")"}' # a LATER session attaching the same volume sees /memory/note.txt # structured KV curl -s -X PUT https://isobox.hsingh.app/v1/memory/facts/pi --data-binary '3.14159' curl -s https://isobox.hsingh.app/v1/memory/facts/pi # -> 3.14159 ``` ## Live-kernel sessions (persistent variables) Add `"type":"kernel"` at `POST /v1/sessions` (default is `"filesystem"`). A kernel session keeps a long-lived Python interpreter, so **variables, imports, and definitions persist across `exec` steps** (Code-Interpreter style) and steps are near-instant (a warm process, not a fresh container per step). Same hardening as every sandbox (filtered egress on by default — pass `"network": false` at create to run offline; read-only rootfs, non-root, capped). The network setting is fixed at create for the kernel's whole lifetime (per-step `network` is ignored). Same exec / fs / DELETE endpoints; `/workspace` and an attached `/memory` volume still work. Runtime `pip install` works inside a kernel (containers set `HOME` and `PYTHONUSERBASE` under `/workspace/.local`, whose `bin/` is on `PATH`); packages installed at runtime persist for the session's lifetime only. Kernels are a scarce resource (a small fixed pool); creating one past the cap returns `429`. Idle kernels are reaped (~30 min). Note: a step that BLOCKS past its wall-time (e.g. `time.sleep(1e9)`) terminates the kernel session — keep long waits bounded. **Exit codes.** A kernel `exec` step reports a real `run.exitCode`: `0` on success, `1` when the step raised an exception, and `n` when it ends via `sys.exit(n)` — which `run_shell` cells use to propagate bash's real exit status (so a missing command reports `127`). A wall-time kill still yields `137`. **Rich-display artifacts.** A kernel `exec` step best-effort captures rich outputs and returns them as an `artifacts` array — top-level in the buffered JSON result and in the SSE `done` event, omitted when empty. Each entry is `{"type":"image/png","dataB64":""}` (every open matplotlib figure, auto- closed after capture) or `{"type":"text/html","data":""}` (the step's trailing bare expression's `_repr_html_`/`_repr_png_`, e.g. a pandas DataFrame). Caps: at most 8 per step, each payload ≤ 2 MiB (base64 length for PNG, HTML byte length for HTML); beyond that the artifact is dropped and the step's `run.truncated` is set. Capture is best-effort and never breaks execution. ```bash S=$(curl -s -X POST https://isobox.hsingh.app/v1/sessions -d '{"runtime":"python","type":"kernel"}') ID=$(echo "$S"|jq -r .id); TOK=$(echo "$S"|jq -r .token) curl -s -X POST https://isobox.hsingh.app/v1/sessions/$ID/exec -H "X-Session-Token: $TOK" -d '{"code":"import numpy as np; arr = np.arange(10)"}' curl -s -X POST https://isobox.hsingh.app/v1/sessions/$ID/exec -H "X-Session-Token: $TOK" -d '{"code":"print(arr.sum())"}' # -> 45 (arr persisted) ``` ## Server-side AI agent (v1) Optional. When the operator sets `ISOBOX_AI_BASE` (an OpenAI-compatible upstream), isobox runs the chat's tool-calling loop **server-side** so the browser never holds the model key. Off by default; when unset, `GET /v1/agent` reports `enabled:false` and the browser-direct chat mode is the fallback. - `GET /v1/agent` — availability probe (no auth, no upstream call). Returns `{"enabled":false}` when off, or `{"enabled":true,"model":""}` when on. - `POST /v1/agent` — run the loop against your existing KERNEL session, streamed as SSE. Create a kernel session first (`POST /v1/sessions` with `"type":"kernel"`) and pass its `id`+`token`. Guarded by auth (when keys are configured), a 1 MiB request-body cap, and a stricter per-IP limit (`ISOBOX_AI_RATE_PER_MIN`, default 6/min). The session is checked with the same tenant+token rules as `/exec`. Request body: ```json { "messages": [{"role":"user","content":"fetch example.com and count the links"}], "session": {"id":"","token":""}, "maxIters": 8 } ``` `maxIters` is optional and clamped to `ISOBOX_AI_MAX_ITERS` (default 8). A malformed body returns `400 {"error":"invalid_json"}` before the stream starts. SSE events (each a UTF-8 frame terminated by a blank line): - `token` — `{"text":"..."}` — an assistant text delta. - `tool_call` — `{"id","name","arguments"}` — `name` is `run_python`, `run_shell`, or `web_search`; `arguments` is a raw JSON string (`{"code":...}`, `{"command":...}`, or `{"query":...}`). - `tool_result` — `{"id","stdout","stderr","exitCode","truncated","durationMs","artifacts"?}` — each result (stdout+stderr) is truncated to 16 KiB before being fed back to the model; `truncated` says whether anything was cut. The optional `artifacts` array (same shape as a kernel exec step: image/png + text/html) is forwarded to the UI only, never back to the model. - `done` — `{"finishReason":"stop|max_iterations|error","iterations":}`. - `error` — `{"code","message"}` — codes include `session_required`, `session_invalid`, `upstream_error`, and `kernel_dead` (the kernel session died mid-run; the loop STOPS instead of looping so the UI can start a fresh session). Two tools run in your kernel session: `run_python{code}` (persistent Python — variables/imports survive across calls) and `run_shell{command}` (bash in the same sandbox and working directory: curl, wget, jq, git, ripgrep, unzip, file available). A `run_shell` command's `exitCode` is bash's real exit status (e.g. `127` for a missing command); a `run_python` step reports `0` on success or `1` if the code raised. `run_shell` steps get a 180s wall and `run_python` steps 60s; a shell command that exceeds its wall is killed IN PLACE with exit code `124` and `[shell timed out after Ns]` on stderr — the kernel SURVIVES (a slow step no longer takes the whole kernel container down). A third tool, `web_search{query,max_results}`, runs SERVER-SIDE (DuckDuckGo) and returns titles, URLs, and snippets, so the model can find URLs even when the session has no egress; it is advertised only when `ISOBOX_AI_SEARCH` is on (the default; set it off to disable). ## Note: memory requires a key (this deployment) On the public demo, `/v1/memory` and `/v1/volumes` require an API key (`X-API-Key: ` or `Authorization: Bearer `) so memory is per-tenant isolated; `/execute` and `/v1/sessions` remain open. Self-hosters: keys are optional unless `ISOBOX_MEMORY_REQUIRE_KEY=1` (then memory needs a key) or `ISOBOX_REQUIRE_AUTH=1` (then everything does).