AI: post transformers
AI: post transformers

MagicDec: Breaking Latency-Throughput Tradeoffs via KV-Compressed Speculative Decoding

28 February 2026 17:46 mcgrof

Listen to episode

About this episode

We review an April 3, 2025 research collaboration between CMU, Moffett AI and Together AI which introduces MagicDec, a new framework designed to accelerate the serving of long-context large language models through speculative decoding.


Previously, conventional wisdom discouraged using speculative decoding (SD) for large batches, as the verification step was believed to be too compute-heavy and inefficient. MagicDec proves that this limitation only applies to short sequences. The paper demonstrates that once sequences pass a critical length, the massive memory cost of loading the KV cache becomes the true bottleneck, shifting inference from being compute-bound to memory-bound.


The authors address the memory bottleneck by applying KV selection algorithms to compress the draft model's KV cache during speculative decoding. They evaluated different KV selection algorithms, both static (SnapKV, StreamingLLM) and dynamic (PQKache). They observed that PQCache leads to high token acceptance rates but it incurs substantial, batch-size-dependent search costs. On tasks like common word extraction and question answering, SnapKV dominated PQCache because it achieved similar acceptance rates without the heavy search overhead. For complex tasks like "needle in a haystack," PQCache initially performed better because its acceptance rate was near 100%. However, as batch sizes increased, PQCache's search costs became too expensive, and SnapKV once again outperformed it.


By effectively managing the memory pressure through KV compression, the system can maintain a high token acceptance rate, minimize costly verification steps, and achieve significant speedups for large batches. The authors test sequence (prefill) lengths ranging from 1k up to 100k tokens. In their theoretical memory footprint analyses, they project context lengths up to 128k tokens. For batch sizes, the core end-to-end speedup experiments focus on large batch sizes ranging from 32 to 256. Additionally, some ablation studies test batch sizes up to 512, and theoretical trade-off analyses chart batch sizes up to 1024. To validate their framework across different hardware capabilities, the researchers used configurations of 4 to 8 GPUs. Specifically, their experiments were run on clusters of 8xA100, 8xH100, 4xH100, and 8xL40 GPUs.


The paper provides the industry with a framework to break the latency-throughput tradeoff when serving long-context Large Language Models (LLMs) at scale. This enables the highly efficient scaling of long-context applications—such as retrieval-augmented generation (RAG), extensive document analysis, code generation, and complex agent workflows—across large batches of concurrent users.


Source:


2024

MAGICDEC: BREAKING THE LATENCY-THROUGHPUT TRADEOFF FOR LONG CONTEXT GENERATION WITH SPECULATIVE DECODING

Carnegie Mellon University, Moffett AI, Together A

Want to find AI jobs?

Join thousands of AI professionals finding their next opportunity

We respect your inbox. Unsubscribe at any time.

© 2026 AI: post transformers. All rights reserved.

Common Questions

Frequently asked questions

Quick answers about how DevFound's AI matching, resumes, and referrals work.

DevFound's AI Copilot ingests your profile, goals, and live job data to deliver curated matches in seconds. Every match includes a resume variant, suggested referrals, and interview prep so you can act immediately. The more feedback you provide, the sharper the Copilot becomes.

AI-led job searches shrink the hours spent sifting through boards and formatting resumes. DevFound pairs automation with your personal outreach, so you reserve energy for interviews and negotiation. Traditional networking still matters, but AI gives you a lift before you even send a message.

Modern AI roles expect comfort with production-grade code, data fluency, and practical ML tooling. The strongest candidates pair deep technical chops with storytelling—translating model impact to product, GTM, and exec partners. Continuous learning keeps you ahead as stacks evolve.

DevFound rewards active seekers. Keep your profile fresh, respond to match quality prompts, and enable alerts so you never miss a role. The AI prioritizes companies and teams that align with your feedback, accelerating both introductions and interview invites.

High-density tech hubs continue to host the deepest AI talent pools, yet distributed teams are catching up fast. Use DevFound filters to hone in on onsite, hybrid, or fully remote roles and watch openings expand across time zones.

DevFound aggregates thousands of remote AI openings and flags the nuances—core hours, async culture, and visa needs—up front. The Copilot also recommends how to position your distributed work experience so hiring managers know you can thrive on a remote team.