> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tensorlake.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Non-blocking Sandbox Creation

> Create a sandbox without waiting for it to start, poll for readiness, and limit how long it may wait for capacity

A sandbox create normally waits until the sandbox is running. When the fleet has room that takes a few seconds. When it does not, a new host has to boot first, which can take several minutes on bare-metal hosts, and a call that waits for it can outlive its own timeout.

Instead of waiting on the create call, create the sandbox, poll for it when convenient, and set a limit on how long it may wait for capacity. Use this when you create many sandboxes at once.

## Create without waiting

Pass `wait=False` and the create returns as soon as the sandbox is durable, with the sandbox in the `pending` state. The sandbox is scheduled and started exactly as a normal create; only the waiting moves to you.

<Tabs>
  <Tab title="CLI">
    ```bash theme={null}
    # Prints the sandbox id and returns
    tl sbx create --no-wait

    # Later, from any shell that has the id
    tl sbx wait <sandbox-id> --timeout 1800
    ```
  </Tab>

  <Tab title="Python">
    ```python theme={null}
    from tensorlake.sandbox import Sandbox, SandboxPending

    pending = Sandbox.create(cpus=4, memory_mb=8192, wait=False)
    print(pending.sandbox_id, pending.state, pending.pending_reason)

    # Block until it runs; returns a connected Sandbox
    sandbox = pending.ready(timeout=1800)
    ```
  </Tab>

  <Tab title="TypeScript">
    ```typescript theme={null}
    import { Sandbox, SandboxPending } from "tensorlake";

    const pending = await Sandbox.create({ cpus: 4, memoryMb: 8192, wait: false });
    console.log(pending.sandboxId, pending.state, pending.pendingReason);

    // Block until it runs; resolves to a connected Sandbox
    const sandbox = await pending.ready({ timeout: 1800 });
    ```
  </Tab>

  <Tab title="HTTP">
    ```bash theme={null}
    curl -X POST https://api.tensorlake.ai/sandboxes \
      -H "Authorization: Bearer $TL_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{"resources": {"cpus": 4, "memory_mb": 8192}, "wait": false}'
    # 202 {"sandbox_id": "...", "state": "pending", "pending_reason": "scheduling"}
    ```
  </Tab>
</Tabs>

The handle has two methods. `ready()` polls the sandbox every two seconds until it is running and returns a connected `Sandbox`. `status()` fetches the current state once. A sandbox known only by id, in another process for example, is picked up with `Sandbox.connect(sandbox_id)`, which also waits while the sandbox is pending.

## Why a sandbox is pending

While a sandbox waits, its `pending_reason` says why. Read it from `status()`, from `GET /sandboxes/{id}`, or from a `SandboxPending` error.

| `pending_reason` | Meaning |
| - | - |
| `scheduling` | Just created; the scheduler has not looked at it yet. |
| `no_executors_available`, `no_resources_available` | No host has room. The fleet is growing to make some. |
| `pool_at_capacity`, `max_capacity_reached` | Your pool or namespace is at its own limit. Capacity is not the issue. |
| `no_executor_for_snapshot_location` | A restore is waiting for a host in the region its snapshot lives in. |
| `waiting_for_container` | Placed on a host; the image is being fetched and the VM booted. |

## Limit the queue time

By default a pending sandbox waits until there is room, however long that takes. If you would rather give up at some point, set `max_pending_secs` on the create. When the bound runs out while the sandbox is still waiting for capacity, the server terminates it with `termination_reason: no_capacity`, and the last `pending_reason` explains what it was waiting for. `0` fails at once if the sandbox cannot be placed right away.

<Tabs>
  <Tab title="CLI">
    ```bash theme={null}
    tl sbx create --no-wait --max-pending-secs 2700
    ```
  </Tab>

  <Tab title="Python">
    ```python theme={null}
    pending = Sandbox.create(cpus=4, memory_mb=8192, wait=False, max_pending_secs=2700)
    ```
  </Tab>

  <Tab title="TypeScript">
    ```typescript theme={null}
    const pending = await Sandbox.create({ cpus: 4, memoryMb: 8192, wait: false, maxPendingSecs: 2700 });
    ```
  </Tab>

  <Tab title="HTTP">
    ```bash theme={null}
    curl -X POST https://api.tensorlake.ai/sandboxes \
      -H "Authorization: Bearer $TL_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{"resources": {"cpus": 4, "memory_mb": 8192}, "wait": false, "max_pending_secs": 2700}'
    ```
  </Tab>
</Tabs>

<Warning>
  Bare-metal hosts take up to 20 minutes to boot. A bound shorter than that expires the sandbox just as the host it triggered comes up, and the next burst pays for a boot again. Use 30 minutes or more for capacity waits.
</Warning>

The bound applies only while the sandbox waits for capacity. Once it is placed and booting, the platform's own startup limits apply. A named sandbox whose resume cannot find capacity within the bound returns to `suspended` rather than terminating, so nothing it holds is lost.

## Timeouts

The blocking `Sandbox.create()` and `pending.ready(timeout=...)` raise `SandboxPending` when their wait runs out. **The sandbox is not deleted.** It keeps its place in the queue and starts when there is room. The error carries `sandbox_id` and the last `pending_reason`, so you can wait again, connect from elsewhere, or terminate it.

<Tabs>
  <Tab title="Python">
    ```python theme={null}
    try:
        sandbox = Sandbox.create(cpus=4, memory_mb=8192, request_timeout=120)
    except SandboxPending as still_queued:
        print(still_queued.sandbox_id, still_queued.pending_reason)
        sandbox = Sandbox.connect(still_queued.sandbox_id)   # keep waiting, or terminate it
    ```
  </Tab>

  <Tab title="TypeScript">
    ```typescript theme={null}
    try {
      const sandbox = await Sandbox.create({ cpus: 4, memoryMb: 8192, requestTimeout: 120 });
    } catch (err) {
      if (err instanceof SandboxPending) {
        console.log(err.sandboxId, err.pendingReason);
        const sandbox = await Sandbox.connect(err.sandboxId);   // keep waiting, or terminate it
      }
    }
    ```
  </Tab>
</Tabs>

If you want the old behaviour, where a timed-out create deletes the sandbox, pass `cancel_on_timeout=True` (`cancelOnTimeout: true`). Deleting a pending sandbox yourself terminates it with `termination_reason: cancelled`.

<Note>
  Before SDK versions with `wait`, a create that timed out deleted its sandbox. A retry loop written for that behaviour ("create failed, create again") now accumulates pending sandboxes until they run or are cancelled. Reuse the id from the error instead, or set `cancel_on_timeout`.
</Note>

## Retries

A create with `wait=False` answers in well under a second, so a lost response is rare, but a retry of an unanswered create can still make a second sandbox. When that matters, give the sandbox a `name`. Names are unique per namespace: a repeated create of the same name is an HTTP 409, and [`get_or_create`](/sandboxes/sdk-reference#get-or-create) resolves it to the existing sandbox. Named sandboxes suspend rather than terminate when their `timeout_secs` elapses, so terminate them explicitly when the work is done.

## Create many sandboxes at once

For a batch, create everything up front with `wait=False`, then poll the list endpoint. One list call sees every sandbox in the namespace, so ten thousand sandboxes cost one request per poll, not ten thousand.

<Tabs>
  <Tab title="Python">
    ```python theme={null}
    import asyncio, json, os
    from tensorlake.sandbox import AsyncSandbox, AsyncSandboxClient, SandboxStatus

    N = 10_000
    SPEC = dict(image="my-org/worker:latest", cpus=4, memory_mb=12_288)
    MAX_PENDING_SECS = 45 * 60      # give up on a sandbox after 45 min in the queue
    IN_FLIGHT = 500                 # sandboxes doing work at once
    STATE = "batch-ids.json"        # ids survive a restart of this script


    async def request_all(client: AsyncSandboxClient) -> dict[str, int]:
        """Create everything without waiting. Names make a rerun resolve to the
        same sandboxes instead of creating more."""
        if os.path.exists(STATE):
            return json.load(open(STATE))
        limit = asyncio.Semaphore(64)          # be polite to the API, not to capacity

        async def one(i: int) -> tuple[str, int]:
            async with limit:
                pending = await client.create(
                    name=f"batch-{i}", max_pending_secs=MAX_PENDING_SECS, wait=False, **SPEC
                )
                return pending.sandbox_id, i

        ids = dict(await asyncio.gather(*(one(i) for i in range(N))))
        json.dump(ids, open(STATE, "w"))
        return ids


    async def work(sandbox: AsyncSandbox, i: int) -> None:
        await sandbox.run("python3", ["job.py", str(i)], timeout=1800)
        await sandbox.terminate()


    async def main() -> None:
        client = AsyncSandboxClient.for_cloud()
        todo = await request_all(client)          # sandbox_id -> job index
        slots = asyncio.Semaphore(IN_FLIGHT)
        running: set[asyncio.Task] = set()

        async def use(sandbox_id: str, i: int) -> None:
            async with slots:
                await work(await AsyncSandbox.connect(sandbox_id), i)

        while todo:
            # One list call sees every sandbox: start the ones that just turned
            # running, drop the ones that failed.
            seen = {s.sandbox_id: s for s in await client.list()}
            for sandbox_id, i in list(todo.items()):
                info = seen.get(sandbox_id)
                if info is None:                              # gone: expired or failed
                    archived = await client.get_archived(sandbox_id)
                    print(f"{sandbox_id}: {archived.termination_reason}")   # e.g. no_capacity
                    del todo[sandbox_id]
                elif info.status == SandboxStatus.RUNNING:
                    running.add(asyncio.create_task(use(sandbox_id, i)))
                    del todo[sandbox_id]
            await asyncio.sleep(5)
        await asyncio.gather(*running)


    asyncio.run(main())
    ```
  </Tab>

  <Tab title="TypeScript">
    ```typescript theme={null}
    import { Sandbox, SandboxClient, SandboxStatus } from "tensorlake";

    const N = 10_000;
    const SPEC = { image: "my-org/worker:latest", cpus: 4, memoryMb: 12_288 };
    const MAX_PENDING_SECS = 45 * 60;
    const client = SandboxClient.forCloud();

    // Create everything without waiting; names keep a rerun from duplicating.
    const todo = new Map<string, number>();
    for (let i = 0; i < N; i++) {
      const pending = await client.create({ ...SPEC, name: `batch-${i}`, maxPendingSecs: MAX_PENDING_SECS, wait: false });
      todo.set(pending.sandboxId, i);
    }

    // One list call per poll sees every sandbox.
    while (todo.size > 0) {
      const seen = new Map((await client.list()).map((s) => [s.sandboxId, s]));
      for (const [sandboxId, i] of [...todo]) {
        const info = seen.get(sandboxId);
        if (info === undefined) {
          const archived = await client.getArchived(sandboxId);
          console.log(sandboxId, archived.terminationReason);       // e.g. no_capacity
          todo.delete(sandboxId);
        } else if (info.status === SandboxStatus.RUNNING) {
          void Sandbox.connect(sandboxId).then((sandbox) => runJob(sandbox, i));
          todo.delete(sandboxId);
        }
      }
      await new Promise((r) => setTimeout(r, 5000));
    }
    ```
  </Tab>
</Tabs>

Sandboxes start oldest first as hosts come up, so a batch that waited for a host is not overtaken by creates that arrived after it. Sandboxes still waiting when their bound runs out fail with `no_capacity` and drop out of the list. Webhook subscribers can replace the poll with the `sandbox.running` and `sandbox.failed` events.

## Learn more

* [Sandbox Lifecycle](/sandboxes/lifecycle) for the full state machine.
* [SDK Reference](/sandboxes/sdk-reference#create) for every `create()` parameter.
* [Create Sandbox](/api-reference/v2/sandboxes/create) for the HTTP contract.
