Explore our latest articles on LLM inference, optimization techniques, and system architecture.
KTransformers v0.7.0 focuses on fine-tuning, with native FP8 Expert weights for LoRA, an AVX512 CPU path for compatible AMD/x86 platforms, full fine-tuning, checkpoint lifecycle support, and a complete Cookbook.
Mooncake introduces high-performance structured-object transfer for heterogeneous and fragmented data, bringing this capability to Miles as a new rollout data-transfer backend. The integration delivers 10–14× faster remote GET and 1.2–1.6× faster PUT compared with the existing Ray path.
On July 27, Moonshot AI officially open-sourced Kimi K3. Having been part of the journey across multiple Kimi generations, Mooncake, together with SGLang, vLLM, and TokenSpeed, delivered Day-0 support to enable efficient distributed inference for Kimi K3.
AgentENV is an open-source execution environment for large-scale Agentic RL that combines strong microVM isolation with fast snapshots, copy-on-write forks, and efficient resource reuse.
Mooncake extends KV cache beyond expensive memory by pooling local NVMe SSDs into a distributed, persistent cache tier that preserves long-context reuse and reduces TTFT for agentic workloads.
Estimate the KV Cache capacity for LLM inference workloads by analyzing hit rate and prefill speedup under different cache budgets, with the help of KV Cache Hit Rate Simulator.
KT-FT v0.6.1 connects MoE SFT and local SGLang serving into an end-to-end loop; split LoRA serving bridges KT expert and SGLang non-expert adapters for Qwen3.5 MoE.
By integrating Mooncake into OpenClaw's real inference path, we not only improved fast-path latency, but also sharply reduced TTFT tail latency in multi-session, long-context workloads, turning a system that was usually fast but occasionally slow into one that feels consistently smooth.
Mooncake is now part of the PyTorch Ecosystem, complementing PyTorch-native LLM serving with high-performance disaggregated data transfer and storage.
A low-cost, low-memory end-to-end fine-tuning and inference workflow for large MoE models with KTransformers, LLaMA-Factory, and SGLang.