- Bookmark stories with reactions via GitHub
- Comment on any post — no account needed to read
- Write your own posts or guides
Ask Our Expert — answered from our published articles, with a link to every source. More about this →
Recent Posts
-
LFM2.5-VL-DSpark Brings Accelerated Vision-Language Models to Local Inference
Hugging Face announces LFM2.5-VL-DSpark, an optimized vision-language model designed for local deployment with improved inference speed. The model combines efficient architecture with quantization-friendly design for edge execution.
-
Llama.cpp Fork Delivers 2-4x Speedup for Multi-GPU MoE Model Inference
A specialized llama.cpp fork optimizes mixture-of-experts models for multi-GPU setups, achieving 2-4x performance improvements for models exceeding single-GPU VRAM limits. This enables practical local deployment of large MoE architectures.
-
Oh My Pi Adds Custom Model Support via vLLM, Llama.cpp, and SGLang
A new guide demonstrates running custom quantized models on Raspberry Pi using multiple inference engines including vLLM, Llama.cpp, and SGLang. This enables practical multi-engine inference workflows on edge devices with detailed configuration examples.
-
vLLM Introduces Watermarking Capabilities for Local Model Serving
vLLM's latest update adds watermarking support for locally-served language models, enabling detection of model-generated content and enhancing control over generated outputs. This feature matters for security and accountability in local deployment scenarios.
-
Llama.cpp Optimizes Kernel Execution with RMS_NORM and SCALE Fusion
The latest llama.cpp release fuses RMS_NORM and SCALE operations into a single kernel, eliminating 96 extra kernel launches per batch on large models like Qwen3.8-27B. This optimization reduces computational overhead without sacrificing accuracy.
-
Ollama v0.40.0 Makes MLX the Default Runner for Apple Silicon
Ollama's latest release shifts to MLX as the default inference engine for Apple Silicon devices, enabling better performance for supported model architectures. This change simplifies local LLM deployment on Mac hardware.
-
Practical Guide: Running Local LLMs on Your Mac - What Fits, What's Free
A comprehensive guide exploring which local LLMs run efficiently on Mac hardware, including free options and performance tradeoffs between commercial and open-source models. Covers model selection, quantization options, and realistic expectations.