The Autoregressive Generation Bottleneck
Standard Large Language Model inference is notoriously memory-bandwidth bound. To generate a single token, the GPU or Unified Memory bus must read every single parameter of the model from VRAM into the compute cores. For a 14B parameter model in Q4_K_M precision (~9.2GB), generating 50 tokens requires reading 460 Gigabytes of memory sequentially.
No matter how many teraflops of compute your Apple Silicon M3/M4 or NVIDIA RTX GPU possesses, your token streaming rate is strictly capped by your memory bus bandwidth. Enter Speculative Decoding: an algorithmic breakthrough that shatters this memory-bandwidth limit without sacrificing a single bit of mathematical output quality.
How Speculative Decoding Works
Speculative decoding pairs a heavy, high-intelligence Target Model (e.g., Qwen-2.5-Coder-14B) with a tiny, lightning-fast Draft Model from the exact same model family (e.g., Qwen-2.5-Coder-0.5B).
┌─────────────────────────────────────────────────────────────────┐
│ SPECULATIVE DECODING VERIFICATION PIPELINE │
└─────────────────────────────────────────────────────────────────┘
[ Draft Model (0.5B Qwen-Coder) ] ── Fast Speculation (120 tok/s)
│ Generates K candidate tokens: [ "def", " ", "parse", "_", "json" ]
▼
[ Target Model (7B/14B Qwen-Coder) ] ── Single Parallel Forward Pass
│ Evaluates all K tokens concurrently against target distribution
▼
[ Verification & Acceptance Gate ]
├── Tokens 1..4 Accepted ✓ (Draft accurately predicted output)
└── Token 5 Rejected ✗ -> Target model emits corrective token
│
▼
Result: Effective speedup jumps from 24 tok/s -> 58 tok/s!
Because code syntax is filled with predictable patterns (like import numpy as np, standard boilerplate, and repetitive variable declarations), the tiny draft model predicts these sequences with 85%+ accuracy. The large model validates all $K$ tokens in a single parallel forward pass, effectively generating 3 to 5 tokens for the memory cost of one!
Setting Up Speculative Decoding in llama.cpp
You can run speculative decoding directly in modern versions of llama.cpp using the -md (model draft) flag:
# Run Qwen 2.5 Coder 14B assisted by Qwen 2.5 Coder 0.5B
$ ./llama-cli -m ./models/qwen2.5-coder-14b-instruct-q4_k_m.gguf -md ./models/qwen2.5-coder-0.5b-instruct-q4_k_m.gguf -ngl 99 -ngld 99 --draft-max 8 --draft-min 3 -p "Write an asynchronous Swift actor for LRU caching:"
Empirical Speedups Across Coding Tasks
We benchmarked code generation throughput on an Apple MacBook Pro M3 Max (36GB Unified Memory):
| Task Type | Standard Baseline (14B) | Speculative (14B + 0.5B) | Acceptance Rate | Net Speedup |
|---|---|---|---|---|
| Python Boilerplate & Classes | 26.2 tok/s | 68.4 tok/s | 84.2% | 2.61x |
| Algorithmic Logic (LeetCode Hard) | 25.8 tok/s | 47.1 tok/s | 61.5% | 1.82x |
| JSON Schema & API Mappings | 26.5 tok/s | 78.2 tok/s | 91.0% | 2.95x |
The Memory Footprint Trade-Off
The only downside of speculative decoding is that both models must reside in memory simultaneously. However, because the 0.5B draft model in Q4_K_M consumes less than 420 MB of VRAM, the trade-off is an absolute no-brainer for any workstation with 16GB or more RAM.