Speculative decoding in production: when it pays and when it does not
Draft models and n-gram lookahead can cut per-token latency a great deal, or do nothing at all while burning compute. The mechanics, the workloads, and the harness we run before switching it on.
- inference
- vllm
- latency
