Mooncake

Mooncake for Miles: From Fragmented Rollout Data to Efficient Bulk I/O
Mooncake for Miles: From Fragmented Rollout Data to Efficient Bulk I/O

Mooncake introduces high-performance structured-object transfer for heterogeneous and fragmented data, bringing this capability to Miles as a new rollout data-transfer backend. The integration delivers 10–14× faster remote GET and 1.2–1.6× faster PUT compared with the existing Ray path.

Aug 20, 2026

When Prefix Cache Meets KDA: How Mooncake Enabled Day-0 Support for Kimi K3
When Prefix Cache Meets KDA: How Mooncake Enabled Day-0 Support for Kimi K3

On July 27, Moonshot AI officially open-sourced Kimi K3. Having been part of the journey across multiple Kimi generations, Mooncake, together with SGLang, vLLM, and TokenSpeed, delivered Day-0 support to enable efficient distributed inference for Kimi K3.

Aug 3, 2026

Scaling KV Cache Beyond Memory with Mooncake SSD Offloading
Scaling KV Cache Beyond Memory with Mooncake SSD Offloading

Mooncake extends KV cache beyond expensive memory by pooling local NVMe SSDs into a distributed, persistent cache tier that preserves long-context reuse and reduces TTFT for agentic workloads.

Jul 15, 2026

How Much KV Cache Budget Do We Need for LLM Serving?
How Much KV Cache Budget Do We Need for LLM Serving?

Estimate the KV Cache capacity for LLM inference workloads by analyzing hit rate and prefill speedup under different cache budgets, with the help of KV Cache Hit Rate Simulator.

Jun 26, 2026

OpenClaw + Mooncake: A Stability Upgrade for Real-World Multi-Session Inference
OpenClaw + Mooncake: A Stability Upgrade for Real-World Multi-Session Inference

By integrating Mooncake into OpenClaw's real inference path, we not only improved fast-path latency, but also sharply reduced TTFT tail latency in multi-session, long-context workloads, turning a system that was usually fast but occasionally slow into one that feels consistently smooth.

Mar 19, 2026

Mooncake Joins the PyTorch Ecosystem
Mooncake Joins the PyTorch Ecosystem

Mooncake is now part of the PyTorch Ecosystem, complementing PyTorch-native LLM serving with high-performance disaggregated data transfer and storage.

Feb 12, 2026