All essays

At One Million Tokens, Post-Training Becomes State Management

Kimi K3 shows why million-token agentic reinforcement learning depends on persistent state, resumable environments, and deployment-aware training.

  • Post-Training
  • Reinforcement Learning
  • AI Systems
  • Kimi

A one-million-token context window sounds like a model feature. During reinforcement learning, it becomes a distributed-systems problem.

Long-horizon agents do not merely generate a long answer. They alternate between reasoning, tool calls, observations, retries, and changes to an external environment. A rollout may survive several training iterations. Its language-model cache, sandbox filesystem, application state, and reward evidence must still describe the same trajectory when computation resumes.

Kimi K3 is interesting because its technical report treats this persistence as part of post-training rather than background infrastructure. My main takeaway is that, at this scale, state management becomes part of the learning algorithm.

What the paper builds

K3 is a 2.78-trillion-parameter mixture-of-experts model with 104.2 billion active parameters and a context window of up to one million tokens. Its architecture scales information flow in three directions. Hybrid attention mixes three Kimi Delta Attention layers with one global Gated MLA layer. Attention Residuals allow layers to retrieve representations from earlier depth blocks. Stable LatentMoE expands the routed pool to 896 experts while activating 16 per token.

The report says that these architecture, data, and training changes collectively improve scaling efficiency by about 2.5 times over K2. “Collectively” is important: the paper does not establish that one component alone produced that gain.

Context length is extended progressively from 8K to 64K during pretraining and from 256K to 1M during cooldown. The post-training pipeline then creates nine specialist policies: three domains—general tasks, general agents, and coding agents—crossed with low, high, and max reasoning effort. Multi-Teacher On-Policy Distillation consolidates them into one model.

The reinforcement-learning environments are deliberately heterogeneous. A configurable harness can vary tool schemas, system prompts, skills, memory, subagents, and context-management strategies rather than teaching one fixed interface. Other environments cover professional workflows, GPU kernel optimization, visual reasoning with Python tools, personal-assistant work across mock applications, autonomous execution with hidden verifiers, and web development.

The paper reports 93.5 on GPQA Diamond, 88.3 on Terminal-Bench 2.1, 91.2 on BrowseComp, 95.0 F1 on DeepSearchQA, and 81.2 on FrontierSWE. These are reported results under the paper’s stated harnesses and mostly maximum reasoning effort; they are not all independent comparisons under identical conditions.

Why partial rollout changes the training problem

Synchronous RL normally waits for a batch of trajectories to finish before updating the policy. Long-running agents create severe stragglers. K3 instead pauses generation after a chosen fraction has completed, trains on available work, and resumes unfinished trajectories in the next iteration.

That improves utilization but creates stale data: a paused trajectory may have been generated by an older policy. The optimization therefore uses per-token regularization to limit how far updates move from the behavior policy. Algorithm and scheduler are coupled. Without resumable state, partial rollout is impossible; without off-policy tolerance, it is unstable.

The physical environment must resume too. AgentENV uses Firecracker microVMs and incremental checkpoints. The report gives best-case checkpoint and resume latencies of 133 ms and 49 ms. A paused sandbox consumes no CPU or memory, which matters because waiting for model inference can occupy as much as 98% of its lifetime. Across K3 training and evaluation, the authors report creating 51,219,741 sandboxes from 1,505,678 images.

Model state is equally demanding. A missed prefix cache at one million tokens can make resumption prohibitively expensive. K3 writes evicted reusable prefixes to CPU memory, restores them before reuse, dynamically throttles requests as contexts grow, and manages recurrent KDA state together with MLA’s token-level KV cache. What looks like a context-window feature is actually a lifecycle for multiple forms of state.

Deployment enters post-training

K3 also applies quantization-aware training throughout SFT and RL, using MXFP4 for expert weights and MXFP8 activations. Rollout and training use the same quantization scheme. This avoids teaching a high-precision policy and only later discovering how quantization changes its behavior.

I think this is a consequential design choice. Post-training is often treated as the last model-quality step, followed by a separate deployment optimization phase. K3 makes serving constraints part of the policy’s experience. When an agent may consume millions of tokens and hundreds of tool calls, cost and latency are not downstream details; they influence which behaviors are usable.

The remaining gaps

The report openly says K3 still trails the strongest proprietary systems overall. On HLE-Full it reports 43.5 without tools and 56.0 with tools, versus 53.3 and 63.0 for Claude Fable 5. On CritPt it scores 23.4 versus 32.3 for GPT-5.6 Sol. Research-level reasoning remains weaker than the most impressive agentic-search results suggest.

Comparison is also difficult. Evaluations mix different agent harnesses, fallbacks, cyber safeguards, and inference budgets. Many results use maximum reasoning effort and temperature 1.0; some benchmarks or judges are internal. The report does not clearly disclose total pretraining tokens, total training FLOPs, or the complete cost of producing the model.

Open weights should not be confused with easy deployment. A 2.78T-parameter model remains a major infrastructure commitment even when only 104.2B parameters activate per token.

There is also a serious dual-use dimension. The report describes previously unknown vulnerabilities and partial success on exploit-development tasks, while an independent suite still found a large gap on end-to-end exploitation. Both facts matter: the capability is meaningful, and it remains uneven. Any account of the model should discuss access controls and verification, not only benchmark progress.

My takeaway

K3 suggests a broader definition of post-training. It includes the policy objective, but also the mechanisms that preserve a trajectory across time: cache retention, sandbox snapshots, hidden verifiers, resource admission, reasoning budgets, and deployment precision.

At short horizons, these can look like implementation details. At one million tokens, they determine which experiences the model can learn from at all.

Primary source