Beyond Git Worktree

Cold Start Latency in On-Demand Agent Sandboxes

As agent systems hit production scale, millisecond delays add up to real operating costs.

Columnist · · 13 min read
Cover illustration for “Cold Start Latency in On-Demand Agent Sandboxes”
Sandbox Infrastructure · September 29, 2026 · 13 min read · 2,814 words

Cold Start Latency in On-Demand Agent Sandboxes.

Why cold start latency is now a production engineering problem, not a prototype concern

Cold start latency in agent sandboxes has stopped being a rounding error and started being a line item. Agentic coding has moved out of the demo phase and into production systems, with 86% of organizations reporting they've gone past pure experimentation and 91% of enterprises now deploying agents against production code Anthropic 2026 State of AI Agents Report. Adoption at that scale changes the math on what counts as acceptable delay. Teams aren't asking whether the agent can write the patch. They're asking whether the infrastructure underneath it holds up when it's running ten thousand times a day Firecrawl AI Agent Sandbox Modal Best Code Execution Sandboxes for Coding Agents Alan Hou Blog DeltaBox Modal Best Code Execution Sandboxes for Tool-Calling AI Agents.

That volume is why milliseconds start to matter. Multiplied across a fleet of agents, the number becomes an operating cost rather than a curiosity.

Whether agents are synchronous or asynchronous is the other axis to consider. A developer waiting on an interactive coding assistant experiences a multi-second cold start as a broken product. A background agent chewing through a long-running refactor barely notices the same delay. Both contexts matter, but the tolerance for latency is not the same animal in each case. 80% of developers now use AI coding agents in workflows, yet trust in AI accuracy has dropped from 40% to 29% year-over-year, indicating that reliability, not just capability, is what teams are measuring now Bunnyshell Agentic Development Guide.

What cold start latency is in the context of an agent sandbox

Cold start latency is the gap between an agent initiating a tool call and the execution environment actually being ready to run code. Every tool invocation that requires a fresh compute environment pays this tax somewhere.

It helps to be precise about what makes this different from the cold starts engineers already know from serverless computing. A Lambda function or a Cloud Functions invocation is built to be stateless: it spins up, does one thing, and disappears. Agent sandboxes can't work that way, because agents need persistent filesystem state and long-running, interactive sessions across multiple steps. That statefulness is the whole point of the sandbox, and it's also the reason the cold start problem is harder here than it is in traditional serverless architecture.

The delay compounds because agents don't make one tool call, they make chains of them. A coding agent doing something as ordinary as reading a file, running a test, and writing a patch touches three separate tool calls in sequence, and if each one lands on a cold sandbox, the agent stacks three cold starts inside a single user interaction. There's a second, quieter version of this same problem: a sandbox that boots without the right runtime pre-installed forces the agent's first action to be a package install, which adds its own delay and introduces dependency resolution as a new way for the whole thing to fail.

The stack of delays that make up a single cold start

Think of a cold start as five layers stacked on top of each other, each one adding its own tax before the agent ever executes a line of code.

Layer one is image pulling and filesystem setup. The container or microVM image has to be fetched and unpacked before anything runs. Heavier images with toolchains baked in take longer to pull, but they avoid forcing a runtime install later. Leaner images start faster on paper, then hand the agent a dependency-install step as its very first action. Layer caching and optimized filesystems are the main tools for managing this tradeoff.

Layer two is process initialization and runtime startup. This is where the guest kernel, in the case of a microVM, or the intercepting runtime, in the case of something like gVisor, has to spin up before anything else happens. On top of that sits the language runtime itself: a Python interpreter or a Node.js process adds its own startup overhead on top of whatever the isolation layer already cost. The choice of isolation technology has the most direct effect on the floor at this layer, a point covered in more detail shortly.

Layer three is credential minting and secrets injection. Sandboxes built for production ought to mint fresh, short-lived credentials for each session rather than injecting a long-lived secret that sits there indefinitely. That's the safer design, but it isn't free: the round trip to issue a credential adds to the time before the sandbox is genuinely ready to act. Teams that skip this step and inject static credentials instead are buying speed with governance, and that trade gets more expensive the more write access an agent actually has to real systems.

Layer four is network setup and egress policy enforcement. This layer is easy to miss in a benchmark, because most benchmarks stop measuring once the container reports itself as ready, not once the agent can actually make its first outbound call.

Layer five, relevant only for stateful agents, is checkpoint and rollback state initialization. Agents doing test-time tree search or reinforcement learning don't just need a running sandbox, they need a known, reproducible starting state, and duplicating an entire sandbox's state every time a checkpoint is taken gets expensive fast. This layer gets its own full treatment further down, because it turns out to be where some of the most interesting recent engineering work is happening.

None of these layers exist in isolation, and an optimization in one quietly reshapes the costs in another. An optimization at layer one, like baking more state into a heavier image so the agent never has to install a dependency, can quietly inflate layer two by giving the runtime more to initialize. Fixing one layer without watching its effect on the next one isn't optimization, it's just moving the cost somewhere less visible. A single cold start involves configuring the sandbox's network interface, applying egress filtering rules, and establishing any tunnels or proxies required for the agent to reach external APIs.

How isolation technology choice shapes the latency floor

Every technique discussed later in this piece operates above a floor, and that floor is set by which isolation technology the sandbox uses in the first place. Three models dominate production use today, and each one buys a different combination of speed and containment.

Standard containers, the Docker model, share the host kernel Northflank Top AI Sandbox Platforms What’s the best code execution sandbox for AI agents in 2026? | Blog…. That's what makes them fast: Daytona gets sub-90-millisecond cold starts running on this model Northflank Top AI Sandbox Platforms What’s the best code execution sandbox for AI agents in 2026? | Blog…. It's also what makes them the weakest isolation boundary of the three, since a kernel exploit in a shared-kernel environment has a much shorter path to escaping the sandbox entirely.

gVisor sits in the middle. It intercepts application system calls in user space and effectively acts as its own guest kernel, which shrinks the kernel attack surface without paying for a dedicated virtual machine. Modal runs on this model. It costs a bit more than a native container in overhead, but it avoids the heavier initialization tax that a full microVM carries.

MicroVMs, built on technologies like Firecracker, Kata Containers, or Cloud Hypervisor, give each workload its own dedicated kernel Webfuse Agentic Coding in 2026 Northflank Top AI Sandbox Platforms What’s the best code execution sandbox for AI agents in 2026? | Blog…. The cost is a heavier boot: a full lightweight virtual machine has to come up before anything runs.

This isn't an arbitrary menu of options. Untrusted AI-generated code, the kind where a prompt injection or a hallucinated command could try to run something destructive, genuinely benefits from the blast-radius containment a microVM provides. An internal coding agent working with tightly controlled inputs doesn't need that much armor, and gVisor or even a standard container may be entirely sufficient. Lambda itself runs on Firecracker under the hood, but it's tuned for stateless execution, and that is why agent sandboxes have carved out their own category rather than just riding on top of existing serverless infrastructure.

The main techniques for reducing cold start latency above the isolation floor

Once the isolation model is fixed, the remaining latency gets attacked through a handful of engineering techniques, each with its own cost.

Pre-warmed sandbox pools are the most direct approach. A pool of sandboxes sits ready before any request arrives, already booted, with a server running, a repo pulled, and dependencies installed. A user's request claims one from the pool instantly, and a replacement gets provisioned in the background to refill it. Modal's own documentation describes this as a standard pattern for cutting perceived cold start latency in production deployments. The tradeoff is straightforward: idle sandboxes sitting in a pool still cost compute, so the pool size has to be tuned against actual traffic, not left oversized out of caution.

Snapshot-based resume works differently. Instead of keeping sandboxes warm and idle, a fully initialized sandbox gets captured as a point-in-time snapshot, runtime already started, dependencies already installed, and future sessions resume from that snapshot instead of rerunning the whole boot sequence. Memory snapshots are the aggressive end of this spectrum, since preserving process state in RAM means resume skips the initialization code path altogether rather than replaying it. Modal itself notes alpha-stage limitations on memory snapshots, and not every platform offering this has reached production stability with it yet.

Standby and resume is a related but distinct idea. Rather than terminating a sandbox after use, some platforms suspend it in place. Cloudflare's Sandboxes take a related approach with configurable idle sleep, defaulting to a 10-minute timeout, with an option to keep a sandbox alive indefinitely through an explicit keepAlive setting Alan Hou Blog DeltaBox Modal Best Code Execution Sandboxes for Tool-Calling AI Agents. This model trades away the guarantee of a truly fresh environment on every session in exchange for near-instant resume, which fits interactive work where the same user keeps coming back, and fits less well where strict multi-tenant isolation is the priority.

Copy-on-write forking addresses a different problem entirely: parallel execution. Rather than booting N independent sandboxes from scratch for N parallel agent branches, a single initialized parent sandbox forks using copy-on-write semantics, so each branch inherits the parent's state instantly instead of repeating the boot. Daytona supports this pattern, and it matters most for fan-out workloads like test-time tree search, where many execution paths need to diverge from one common starting point.

Image and filesystem optimization rounds out the list. Modal's runtime includes a custom filesystem built specifically so that large images can come online quickly without image size becoming the bottleneck. Pre-baking common language runtimes and dependencies into the base image also removes the first-action package install problem described earlier, at the cost of a larger image to maintain and pull. Modal supports filesystem snapshots, directory snapshots (beta), and memory snapshots (alpha); Together Code Sandbox resumes from snapshot in ~500ms Webfuse Agentic Coding in 2026 Northflank Top AI Sandbox Platforms. Some platforms keep sandboxes in a suspended-but-not-terminated state; Blaxel's "perpetual sandboxes" resume from standby in under 25ms with full state preserved, with no compute charges while idle Blaxel Cold Start Latency in Agent Tool Calls DeltaBox arXiv paper.

What the DeltaBox research shows about checkpoint and rollback latency specifically

Layer five, checkpoint and rollback, deserves its own treatment because it's where the most recent research has actually changed the numbers rather than just shuffling tradeoffs.

Existing mechanisms such as DMTCP and CRIU duplicate the entire sandbox state at each checkpoint, causing latency in the range of hundreds of milliseconds to seconds per checkpoint. When an agent is trying to explore a deeply branching search tree, that per-checkpoint cost becomes the actual bottleneck, worse than the reasoning itself.

The DeltaBox paper, published in May 2026 by Yunpeng Dong, Jingkai He, and coauthors, starts from a simple observation: consecutive checkpoints in agent workloads look almost identical to each other, differing only by small incremental changes step to step. Duplicating the delta does the job, so there's no reason to duplicate the whole state. The system splits into two components. DeltaFS organizes filesystem state into layers, freezes the writable layer, and reduces file updates to copy-on-write operations, so a rollback becomes a layer switch instead of a full restore. DeltaCR handles process memory the same way, taking incremental dumps and rolling back by forking directly from a frozen template process rather than replaying a traditional restore pipeline.

The measured results show exactly that shift. That's not an incremental tuning gain; it's an order-of-magnitude shift in what a checkpoint costs.

The raw speedup unlocks changes that were previously bottlenecked by the time budget going into copying state rather than reasoning. When most of an agent's time budget was being spent copying state rather than reasoning, the search itself was starved. That's a capability change dressed up as a speed improvement.

A companion paper, AgentRewind, approaches the same territory from a different angle: its DART mechanism selects semantically valid restore points for structured tool-calling agents, but it needs explicit control flow and defined recovery boundaries to work DeltaBox arXiv paper. DeltaBox, by contrast, operates at the system level and doesn't require the agent itself to be aware that checkpointing is happening at all DeltaBox arXiv paper. None of this is shipping as a commercial product yet. It's a signal of where the latency floor for stateful agents is headed. The problem this addresses: LLM-powered agents doing test-time tree search or reinforcement learning require high-frequency checkpoint and rollback of the complete sandbox state (files and process memory) so they can explore multiple execution paths from the same starting point. Results from evaluations on SWE-bench and RL micro-benchmarks show checkpoint completes in 14ms, rollback in 5ms, millisecond-level latency versus the prior baseline of hundreds of milliseconds to seconds DeltaBox arXiv paper. Agents can explore 10–100x more execution paths in the same time budget; before DeltaBox, most of the time budget was consumed copying state rather than doing reasoning Alan Hou Blog DeltaBox Modal Best Code Execution Sandboxes for Tool-Calling AI Agents.

How current sandbox platforms perform across the latency layers in practice

It also supports copy-on-write forking for parallel agent runs. The company pivoted from development environments into AI agent infrastructure in February 2025 and launched its current platform in late April of that year Blaxel Cold Start Latency in Agent Tool Calls. The speed comes from that narrower surface it's built on.

Modal runs on gVisor isolation and is engineered specifically for fast cold starts, backed by a custom filesystem and an optimized container runtime. It scales to more than 50,000 concurrent sessions and serves over 10,000 teams Modal Best Code Execution Sandboxes for Coding Agents. Sessions run up to 24 hours, with longer workflows handled through snapshotting rather than raw session extension. Ramp uses Modal Sandboxes in production for background coding agents that generate code changes and push them into commits or pull requests, a concrete example of the asynchronous-agent pattern described earlier in this piece. It doesn't currently offer a bring-your-own-cloud option.

Northflank runs microVM isolation through Kata Containers with Cloud Hypervisor, alongside Firecracker and gVisor, and processes more than 2 million isolated workloads a month. Sessions have no fixed time limit, which removes the workaround problem that platforms with hard session caps introduce for long-running agents. It accepts any OCI container image and offers bring-your-own-cloud deployment across AWS, GCP, Azure, Oracle, CoreWeave, Civo, on-premises, and bare metal.

Taken together, these three platforms make the tradeoff from earlier sections concrete rather than theoretical. Speed, isolation strength, and deployment flexibility do not arrive as a package. Every platform on this list picked which one to lead with, and the right choice for a given team depends on what the agent is actually doing, not on which number looks best on a benchmark page. The fastest published cold start is sub-90ms, using Docker isolation by default (standard containers, not microVMs) What’s the best code execution sandbox for AI agents in 2026? | Blog…. Tradeoffs per Northflank's ranking include limited networking (no first-class tunneling or egress policies) and sandbox-only scope with no broader infrastructure for databases, APIs, or GPUs. Cold start of ~150ms is achieved using Firecracker microVMs (hardware-level isolation, a stronger guarantee than containers) Webfuse Agentic Coding in 2026. The session limit is a 24-hour maximum on the Pro plan, while the Hobby plan is capped at 1 hour; long-running agent workflows that exceed this require workarounds. With AI-first SDK design, E2B claims to be cited by 88% of Fortune 100 companies for frontier agentic workflows, and customer testimonials cite ~1-hour end-to-end integration time. SOC 2 Type II audit has been completed, with HIPAA support available on Enterprise plans.

Sources

  1. AI Agent Sandbox: How to Safely Run Autonomous Agents in 2026
  2. Top AI sandbox platforms in 2026, ranked | Blog — Northflank
  3. Best Code Execution Sandboxes for Tool-Calling AI Agents in 2026 | Modal Blog
  4. Best Code Execution Sandboxes for AI Agents in 2026 | Modal Blog
  5. What’s the best code execution sandbox for AI agents in 2026? | Blog — Northflank
  6. Cold Start Latency in Agent Tool Calls: Fix the Stack | Blaxel Blog
  7. DeltaBox: Scaling Stateful AI Agents with Millisecond-Level ...
  8. DeltaBox: Millisecond-Level Checkpoint/Rollback for AI Agent State Exploration / DeltaBox:为 AI 智能体状态探索提供毫秒级检查点/回滚 | Alan Hou

More in Sandbox Infrastructure