Tagged "speculative-decoding"
- vLLM v0.28.0 Features Major Kimi-K3 Optimization and Decode Context Parallel Support
- Liquid AI Releases LFM2.5-DSpark Draft Models with 3.18x Faster Decoding
- HackerNoon Compares 7 Best Self-Hosted Inference Servers for Open-Source Models
- 7 Best Self-Hosted Inference Servers for Open-Source Models Compared (2026)
- vLLM v0.27.0 – Kimi K3 Support and 561 Commits from 242 Contributors
- Ollama v0.32.6: Faster Apple GPU Inference with Speculative Decoding
- Theoretical Bottlenecks for Scaling LLM Inference to Achieve Higher Token per Second
- NVIDIA DFlash Block Diffusion Accelerates Autoregressive LLM Inference
- Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
- Prefill Once, Fan Out: KV Snapshot Sharing for Multi-Agent LLM Pipelines
- Meet EAGLE 3.1: The Speculative Decoding Algorithm That Fixes Attention Drift in LLM Inference
- DFlash Speculative Decoding Delivers 8.5x Speed Improvement for LLM Inference
- Google Releases Gemma 4 Multi-Token Prediction Drafters To Accelerate AI Inference
- Google Accelerates Gemma 4 Inference Speed 3x With Multi-Token Prediction Drafters
- Supercharging LLM Inference on Google TPUs: Achieving 3X Speedups With Diffusion-Style Speculative Decoding
- Prefill Is Compute-Bound, Decode Is Memory-Bound: Optimizing GPU Utilization for LLM Inference
- DFlash Doubles Token Generation Speed of Qwen3.5 27B on Mac M5 Max
- Speculative Decoding Achieves 29% Speed Boost for Gemma-4 31B
- DFlash Speculative Decoding Achieves 3.3x Speedup on Apple Silicon
- I Replaced My Local LLM With a Model Half Its Size and Got Better Results — and It Wasn't About the Parameters
- Speculative Decoding Made My Local LLM Actually Usable
- Qwen 3.5 27B Achieves 1.1M Tokens/Second on B200 GPUs with Optimized vLLM Config
- P-EAGLE: Faster LLM Inference with Parallel Speculative Decoding in vLLM
- Qwen 3.5-27B Demonstrates Exceptional Performance with Thoughtful Prompt Engineering