Sandbox Execution

A sandbox is a real, disposable computer you provision on demand. It gives an AI agent, a test run, or a developer a shell, a Python kernel, a Node.js runtime, a headless browser, and a browser-based IDE in one isolated workspace, then tears down cleanly when the work is done. It is built to run untrusted, model-generated code: the sandbox has no access to your production data stores or platform secrets, and its outbound network is denied by default except for a fixed list of package and source registries.

Sandboxes run on usage credits and are available on every plan, including Free. A sandbox's active time counts as Real-Time Session Minutes, at the same rate whatever its size. See pricing and the usage credits guide.

What Is Inside a Sandbox

After this section you will know what you get in a single sandbox and why you do not need to stitch separate services together. Most agent workflows otherwise juggle one service for compute, another for a browser, and another for an editor. A Backbuild sandbox bundles all of them in one workspace, so the same session that runs your code can also drive a browser and open an IDE.

RuntimeWhat it is for
ShellAn interactive shell for running commands, driven through the terminal API.
PythonA Python kernel for executing Python code and notebooks.
Node.jsA JavaScript runtime for executing Node code.
IDEA browser-based code editor, reached through the sandbox proxy.
Browser + displayA headless browser and a virtual display, viewable through the proxy, for automated browsing and rendering.

These runtimes coexist in the same workspace. There is no per-language base image to choose at provision time: you provision one sandbox and use whichever runtimes the task needs.

Your First Sandbox

After this section you will be able to provision a sandbox, run code in it, and tear it down. The lifecycle is a small set of session-scoped operations. Each sandbox is bound to a session you name, so every call refers to the same workspace by its session id.

  1. Provision a sandbox for a session, choosing a size and, optionally, a git repository to clone.
  2. Run commands and code through the terminal, Python, and Node.js runtimes, or open the IDE, browser, and display through the proxy.
  3. Collect the artifacts the work produced.
  4. Tear down the sandbox, which saves the workspace and releases the compute.
# Provision a sandbox for a session
curl -X POST https://api.backbuild.ai/v1/sandbox \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "sessionId": "019d...",
    "size": "standard-2"
  }'
# Run a command in the sandbox shell
curl -X POST https://api.backbuild.ai/v1/sandbox/terminal/019d... \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "action": "exec",
    "command": "python --version"
  }'

Provisioning a sandbox needs the project.update permission, the same permission that lets you change a project. See Roles & Permissions.

Sandbox Sizes

After this section you will be able to pick the right instance size for a workload. Three named sizes are offered. If you omit the size, you get standard-2. Larger sizes have more CPU, memory, and disk.

SizeCPUMemoryDisk
standard-2 (default)1 vCPU6 GiB12 GB
standard-32 vCPU8 GiB16 GB
standard-44 vCPU12 GiB20 GB

Older size names are still accepted and mapped to a current size, so existing callers keep working: basic, standard-1 and small map to standard-2, standard to standard-3, and large to standard-4. The sandbox image does not fit the smaller disks of basic and standard-1, so standard-2 is the smallest size. Prefer the named sizes above for new work.

Session Persistence: Teardown and Resume

After this section you will understand exactly what survives a teardown and how to pick up where you left off. A sandbox is tied to its session, and the session record is kept after teardown so it can be resumed. This is what lets an agent stop consuming compute between bursts of work without losing state.

  • Teardown flushes the workspace to your organization's storage, commits and pushes the changes in a cloned repository (see Cloning a Git Repository), and stops the container, releasing its compute.
  • Resume restarts the sandbox for the same session and reconstitutes the workspace from where teardown left it.
  • A fresh start is a new session id. Provision with a new session id whenever you want a clean workspace.
  • A stopped sandbox stays stopped until you resume it. A sandbox also stops on its own after a period of inactivity, saving its workspace and pushing a cloned repository's changes the same way. While it is stopped, requests to its terminal, IDE, and files answer 409 with code SANDBOX_STOPPED; resume it to bring the workspace back.
One session, provisioned and resumed 1. Provision a session id, a size, and an optional clone 2. Running compute active, the workspace live 3. Teardown, or idle too long saves the workspace, pushes a clone's changes, and frees the compute 4. Stopped the workspace kept, no compute running 5. Resume restores the workspace, running again A new session id starts a fresh workspace
Teardown releases the compute but keeps the session, so resuming brings the workspace back. A new session id always starts clean.

What survives a teardown? The workspace files are flushed to your organization's storage, and in a sandbox that cloned a repository the changes are committed and pushed to the cloned branch. Resuming the same session restores that workspace. The running compute is released, so you stop paying for it while torn down.

How do I start completely fresh? Provision a new session id. Each session id is its own workspace.

How many sandboxes can run at once? Your plan sets how many of your organization's sandboxes can run at the same time; free plans allow one. Provisioning or resuming beyond that answers 429 with code SANDBOX_LIMIT_REACHED. Tear one down, or let an idle one stop, and try again. Stopped sandboxes do not count.

Who can open a sandbox? A sandbox's terminal, IDE, files, and artifacts open only for the member who started it, because it works with that member's own connected credentials. Other members of the organization can see its status and tear it down.

Cloning a Git Repository

After this section you will be able to start a sandbox with your code already checked out. Pass a repository and, optionally, a branch when you provision, and the sandbox clones it into the workspace before you start.

curl -X POST https://api.backbuild.ai/v1/sandbox \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "sessionId": "019d...",
    "size": "standard-2",
    "gitRepo": "owner/repo",
    "gitBranch": "main"
  }'

The clone uses your own connected GitHub credentials, scoped to your account, so a sandbox only ever reaches repositories you can already access. If you ask for a repository but have not connected GitHub, provisioning fails rather than proceeding without credentials. Repository and branch names are validated to a strict character set, so only well-formed values are accepted. Agents that do not need a repository simply omit gitRepo.

Two things follow from a clone using your own credential, and both matter when the sandbox runs code you did not write:

  • The credential is in the sandbox. Code running in a sandbox that cloned a repository can read your GitHub credential and use it on GitHub, which the outbound allowlist reaches, with the access you granted when you connected GitHub.
  • Changes are pushed for you. When the sandbox is torn down, or stops after sitting idle, Backbuild commits every change in the workspace and pushes it to the branch you cloned, using your credential.

So clone a branch made for the sandbox's work rather than one you protect, keep branch protection on the branches that matter, and review what arrives before you merge it. For code you do not trust with your GitHub access, provision the sandbox without a repository.

Artifacts

After this section you will be able to retrieve the files a run produced. The work a sandbox does, generated files, build outputs, reports, is available as artifacts you can list and download for the session.

Method & pathDescription
GET /v1/sandbox/artifacts/:sessionIdList the artifacts produced in the session.
GET /v1/sandbox/artifacts/:sessionId/:nameDownload a single named artifact.

Artifact downloads are served defensively: they are delivered as attachments with a safe content type so that a downloaded file is saved, not executed or rendered in place.

Reaching the IDE, Notebook, and Display

The browser-based IDE, the Python notebook, and the virtual display are reached through a transparent proxy scoped to the session, so you can open them in a browser or embed them in your own product. The proxy path is ALL /v1/sandbox/proxy/:sessionId/*, with convenience routes for the editor, the notebook, and the display surface.

Safety and Isolation

After this section you will be able to answer, for a security review, whether it is safe to run untrusted code here. Every time an agent executes model-generated code, it runs software you did not write, so the sandbox is built to contain it by default rather than to trust it. The controls below are the user-facing security posture of a sandbox.

  • Per-session and per-organization isolation: a sandbox belongs to one session in one organization. It cannot reach another sandbox, and a request that reaches across the organization boundary is treated as not found.
  • Deny-by-default egress: outbound network access is restricted to a fixed allowlist of package and source registries (such as GitHub, npm, PyPI, container registries, and operating-system package mirrors). Everything else is denied, so a sandbox cannot call arbitrary or attacker-chosen hosts. Name lookups (DNS) are answered by the platform's resolvers and are not limited by the allowlist.
  • Egress quota: outbound web traffic (HTTP and HTTPS) is metered and capped per hour.
  • Output secret scrubbing: command and code output is scrubbed for secret material before it is returned, so credentials that appear in output are not surfaced back verbatim.
  • Resource limits: CPU, memory, and disk are capped by the selected instance size.
  • No production data or platform secrets: a sandbox has no access to your production data stores or platform secrets. What it does hold is what you give it: the code and data you put in it and, when you ask for a clone, your own GitHub credential.
  • Safe artifact delivery: downloads are forced to save as attachments with a safe content type rather than render in place.
  • Audit logging: sandbox operations are recorded with the identity of the user who performed them.
Where a sandbox can connect The sandbox the code you run, including untrusted code Allowed GitHub npm PyPI container registries OS package mirrors Denied everything else, including any arbitrary or attacker-chosen host No path to your production data stores or platform secrets What it does hold with a clone, your GitHub credential, which code in the sandbox can use on GitHub
A sandbox reaches only a fixed set of package and source registries, and every other outbound host is denied. It has no path to your production data stores or platform secrets, but a sandbox that cloned a repository holds your GitHub credential.

Is it safe to run untrusted, model-generated code? It is what a sandbox is built for. Each session is isolated, outbound network is denied by default except for a fixed registry allowlist and name lookups, outbound web traffic is quota-capped, output is scrubbed for secrets, and there is no access to your production data stores or platform secrets. The code can still use whatever you put in the sandbox, so for code you do not trust with your GitHub access, provision without a repository.

Can a sandbox reach my internal services or another tenant? No. Outbound traffic is limited to a fixed list of package and source registries, and cross-organization access is treated as not found.

Access and Credits

Provisioning, resuming, tearing down, and driving a sandbox require the project.update permission. Sandboxes run on usage credits and are available on every plan, including Free; a sandbox's active time counts as Real-Time Session Minutes, at the same rate whatever its size. The usage credits guide explains how credits are metered and the pricing page lists each plan's included usage.

Use Cases

AI agent code execution

When an agent needs to write and run code, it provisions a sandbox for its session, runs the generated code through the shell and runtimes, validates the output, and tears the sandbox down. The untrusted code runs inside the isolation boundary the whole time, so a bad prompt cannot reach production resources or any host outside the registry allowlist.

Automated testing

Run a test suite in a clean, isolated workspace. Each run can start from a fresh session, which eliminates state left over from a previous run, and clone the repository under test on provision.

Sandboxed development

Give a person an interactive environment, IDE, shell, and runtimes, to experiment with code or configuration without touching a production workspace, then flush it to storage and resume it later.

API Reference

Method & pathDescription
POST /v1/sandboxProvision a sandbox for a session (size, optional git clone).
GET /v1/sandbox/status/:sessionIdQuery the live state of a session's sandbox.
POST /v1/sandbox/terminal/:sessionIdDrive the shell: create, write, read, and exec.
POST /v1/sandbox/resume/:sessionIdResume a torn-down sandbox, restoring its workspace.
POST /v1/sandbox/teardown/:sessionIdFlush the workspace, push git, and stop the sandbox.
GET /v1/sandbox/artifacts/:sessionIdList artifacts; add /:name to download one.
ALL /v1/sandbox/proxy/:sessionId/*Transparent proxy into the IDE, notebook, browser, and display.

See the Tools & Sandbox API for complete request and response documentation.