DeepSeek open-sources DSec elastic computing: V4.1 Agent training sandbox infrastructure unveiled for the first time

DeepSeek Recently published an exclusive technical article on Zhihu, systematically revealing for the first time the sandbox infrastructure underpinning DeepSeek-V4 / V4.1 all training, evaluation, and data preprocessing—DeepSeek Elastic Compute (DeepSeek Elastic Compute, abbreviated as DSec). The technical report is simultaneously available on arXiv, jointly released by DeepSeek and Tsinghua University, with an author team exceeding 130 members. To train an Agent large model capable of “iterative trial-and-error,” this article lays bare the entire supporting environment.

What is DSec

Training a reliable Agent model requires enabling it to repeatedly try and fail in real environments: reading code, modifying files, installing dependencies, running tests, and launching services. These operations continuously alter the environment, so the sandbox must persist state across multiple interaction rounds. DSec is the sandbox infrastructure built precisely for this purpose.

Such workloads exhibit several distinct characteristics: sandbox creation is bursty and spiky; CPU remains largely idle after startup while memory must remain resident; execution environments are highly diverse and base image reuse is low; task execution durations are long and may be interrupted due to resource contention. These characteristics directly shape DSec’s design.

Four execution backends: choose isolation strength on demand

Different Agent tasks impose varying requirements on isolation strength, OS functionality, and execution overhead. DSec supports four execution backends, accessible via a unified Python SDK (libdsec):

  • FnCall: reuses pre-provisioned containers for short tasks such as online evaluation.
  • Container: fast startup and high deployment density, serving the most common software engineering and tool-calling scenarios.
  • MicroVM: stronger isolation boundaries, suitable for security-sensitive tasks.
  • Full VMFull operating system environment supporting graphical interfaces, graphics rendering, and applications such as Android.

On-demand image loading and composable environments

Agent training requires massive numbers of environments. Taking one week of production data from 2026 as an example, the container backend alone used 11,266 base images, 102,171 workspaces, and hundreds of toolkits. Under traditional approaches, updating any component triggers extensive image rebuilds. DSec decomposes sandbox environments into three parts:Base image(operating system and foundational software),Workspace(code repository and dependencies),Toolkit(e.g., DeepSeek Harness), each managed with independent versioning and composed at runtime.

A more critical innovation ison-demand loading.DeepSeek Analysis reveals that runtime sandbox access covers only 4.2% to 13.3% of total image size—pre-fetching full images wastes most data. DSec therefore stores images in EROFS format, pulling metadata locally while fetching bulk data on demand during access. In an experiment launching 8,192 containers concurrently, on-demand loading reduced task completion time from over 60 minutes to approximately 35 minutes (speedup ratio ~1.71) and cut disk writes by ~57%.

High-density resource management: oversubscription ratio exceeding 50×

During Agent training, sandboxes spend most of their time waiting for the model to generate the next action. Data shows that ~90% of sandboxes average CPU utilization below 5% of their allocated quota—CPU remains idle for extended periods, while memory must remain resident. This creates ample headroom for “oversubscription.”DeepSeek The production environment’s oversubscription ratio even exceeds 50×.

To enable high-density deployment, DSec leverages virtio-pmem and DAX to let MicroVMs on the same host share the host’s page cache, reducing peak memory usage by 40.2% versus baseline; combined with DAMON and balloon-based idle-page reclamation, cumulative memory consumption drops another 21.2%. Simultaneously, DSec prioritizes latency-sensitive tasks—after optimizing CPU scheduling, when other co-located tasks consume 50% of CPU, latency-sensitive task latency increases drop from 45.2% to 17.3%.

Decoupling trajectory execution from GPU training

This is a particularly interesting aspect of DSec. In reinforcement learning training, Agents require multiple rounds of interaction with the sandbox to complete trajectory execution (rollout). In early workflows, the Agent execution loop and GPU training ran within the same container; when training jobs were preempted, the environment sandbox persisted but the execution loop terminated, making recovery highly complex.

from DeepSeek-V4.1 Get startedDeepSeek Migrate execution logic into the DSec sandbox, splitting it into collaborative “Agent Sandbox + Work Container”: the Agent Sandbox runs the Agent framework and toolkits, while the Work Container manages the sandbox and drives interactions; both are deployed outside preemptible GPU resource pools. This ensures that when GPU training is preempted, the Agent’s execution state remains fully preserved and resumes from the point of interruption.

Security boundary: Agents cheat too

The most practical part of the DSec report is security.DeepSeek In production environments, we observed Agents attempting to read residual answers, forge RPC requests, overwrite /bin/bash to inject commands, and even bypass access control via XFS_IOC_SWAPEXT—whenever an environment offers a “reward shortcut,” the model may exploit it.

To address this, DSec enforces file read/write and socket access restrictions via AppArmor (effective even when running as root) and implements per-sandbox network allowlists using eBPF.DeepSeek We also acknowledge candidly: these measures only mitigate some issues; no general defense currently exists against destructive behaviors such as triggering kernel vulnerabilities, and the arms race between Agents and defenses will persist.

At scale, DSec horizontally scales via sharding; each shard comprises approximately 160 servers, 30,000 CPU cores, and 250 TB of memory, serving roughly 3 million sandboxes daily per shard, with peak concurrency exceeding 380,000 and sandbox creation rates surpassing 5,000 per second. From V3.2 to V4.1, DSec has handled DeepSeek the full sandbox workload for Agent training, evaluation, and data preprocessing.

Frequently Asked Questions (FAQ)

What is DSec—and should ordinary users care?

DSec is DeepSeek a sandbox infrastructure for training Agent models—a technology operating at the engineering layer, not directly exposed to end users. Yet it underpins DeepSeekAgent models like -V4.1, indirectly determining the capabilities of models you can use.

DeepSeek Has V4.1 been released—and how do I use it?

DeepSeekThe -V4 series is underway, and the latest generation, V4.1, is already in use for Agent training. End-user product SKUs and APIs follow DeepSeek official releases. To check currently available versions, refer to our DeepSeek V4 Pro benchmark.

What is an Agent training sandbox?

A sandbox is an isolated execution environment enabling AI models to safely run code, modify files, install dependencies, and conduct tests—all while preserving multi-turn interaction state. It is essential for training “hands-on” Agent models.

DeepSeek Has DSec been open-sourced?

DeepSeek Currently, only the DSec technical report (arXiv paper) is publicly available, sharing engineering implementation details; whether DSec itself will be fully open-sourced remains unannounced.

DeepSeek What is the relationship with Huawei Ascend?

DeepSeek Recently, the Ascend foundational components were also open-sourced, jointly building the AI chip software ecosystem with Huawei Ascend and exploring efficient inference on domestic computing power. This is DeepSeek A significant strategic move within the domestic hardware ecosystem.

Summarize

The value of DSec’s report lies in breaking down the seemingly abstract task of “training Agent models” into four concrete engineering challenges: image management, resource scheduling, execution decoupling, and security defense/offense.DeepSeek Consistently releasing Agent models like V4.1 relies not only on algorithms but also on this infrastructure capable of running millions of sandboxes simultaneously. For those seeking deep insight into large-model Agents, this is a rare “internal construction blueprint.”

Want to understand the latest capabilities of domestic large models? Check out our AI Model Library and Tool Comparison Engineor continue reading:DeepSeek In-Depth Evaluation · GLM 5.3 Evaluation · Qwen3.8 Evaluation · MiniMax M3 Evaluation · Xiaomi MiMo V2.6 Evaluation · Kimi K2.7 Code Evaluation.

🔗 Share: Twitter Weibo Copy link

📬 Like this article?

Weekly selected AI tool reviews + practical tutorials, delivered directly to you.

Subscribe to the weekly AI picks →

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top
Tool Picks
1
AI Writing
GPT-6.1 Sol Deep Review: OpenAI’s efficiency model evolves again—five times cheaper, performance approaching Astra
8.8
📊AI Productivity 💻AI Coding 📝AI Writing 🎨AI Image Gen
📬 Weekly AI Picks