Tagged "consumer-gpu"
-
JetBrains Releases Junie Local: On-Device Coding Agent for macOS
-
Qwen 3.6 Now Easier to Run Locally on Mac with JetBrains Integration
-
8 Free Tools to Assess Your PC's Local AI Capabilities
-
llama.cpp Build 10605: Mamba2 GEMM Optimization Improves State-Space Model Performance
-
FreeToken: Edge-Native MoE Serving Engine Runs 753B GLM-5.2 on Single Workstation GPU
-
Strong Domain Adaptation Results with Qwen 3 4B Fine-Tuning
-
llama.cpp Adds CUDA Pool Operations Support
-
vLLM's Disaggregated Serving Cuts GPU Interference, Delivering 2.5x Higher Goodput
-
llama.cpp Build b10581 Adds DSpark Support for Faster Local Inference
-
Liquid AI Releases DSpark Version of Compact LFM2.5 Models with Up to 2.67x Speedup
-
Qwen3.8-27B: Running a Frontier-class Open Model on Your Local GPU
-
llama.cpp b10549: Tensor Parallelism Support for LFM2/LFM2MOE Models
-
llama.cpp b10524 Makes MoE Expert Scatter Deterministic in OpenCL
-
Liquid AI Releases LFM2.5 Q4_0 Checkpoints from Quantization-Aware Distillation
-
Ollama Runs Free AI Models Locally on Mac, Windows and Linux
-
Native vLLM and ROCm 7.15 Support for AMD RDNA2 GPUs on Windows
-
Qwen3.8-27B Matches Claude Opus 4.6 on Coding, Runs on Consumer GPUs
-
What If Local LLM Inference Is Using Consumer Hardware Wrong?
-
GGUF Quantization Deep Dive: Q4_K_M vs IQ4_XS vs IQ4_NL Performance
-
Qwen3.8-27B Surpasses 1 Million Downloads, Overseas Developers Race to Maximize Local Deployment
-
GGUF Quantization Compared: Q4_K_M vs. IQ4_XS vs. IQ4_NL Performance Analysis
-
AMD Adds Day 0 Qwen3.8 Support, Radeon AI PRO R9700 Hits 51.8 Tokens per Second
-
Show HN: I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac
-
Unsloth Releases Qwen 3.8 27B GGUF Quantised Weights
-
Qwen 3.8 27B Successfully Runs on 16GB RAM Using LM Studio
-
Ollama Adds Qwen 3.8 27B with Optimised Apple Silicon Support
-
AMD Optimizes Qwen 3.8 27B for Ryzen AI Max and Radeon GPUs
-
Meta's Muse Glimmer on ExecuTorch Enables Fast On-Device Agentic AI
-
Liquid AI Releases LFM2.5-VL-3B: Compact Vision-Language Model for Edge Inference
-
Running DeepSeek's 284B LLM on a Laptop: Quantisation and GGUF Optimization
-
Building Local LLM Rigs with Used Server GPUs: 32GB VRAM for €220
-
AMD Launches Gorgon Halo and ROCm.AI for Local AI Inference with Workstation Hardware
-
Benchmarking Local LLMs on Consumer Hardware: Real-World Performance Data
-
Ollama Releases NVIDIA Nemotron 3.5 Lightning for Local Agent Deployment
-
Meta's Muse Glimmer Now Available Across All Platforms in Ollama
-
Meta Releases Muse Glimmer: 30B Open-Source LLM for Local Deployment
-
How to Install Ollama on Windows 11 for Local AI Inference
-
Muse Glimmer Now Available on Ollama – Meta's Open Multimodal Agent Model
-
On-Device AI Market Combines AI Operations With Local Processing
-
llama.cpp Improves CUDA Performance with Kernel Fusion
-
DeepSeek V4 Flash Achieves 82.7% on Terminal-Bench 2.1
-
MSI Crosshair A16 HX: Professional Gaming Laptop Built for AI and Gaming
-
Llama.cpp Adds LRU Scheduler for Multi-Model Serving
-
Llama.cpp B10327 Fixes CUDA Quantized Copy Kernel Performance
-
Chrome and Edge Browsers Quietly Deploy Up to 20GB AI Models on Windows 11
-
Self-Hosted LLM Costs 2026: Comprehensive Pricing Comparison
-
Show HN: Benchmark Local LLMs Fit for Your Device Specs
-
llama.cpp Build b10301: CUDA Optimization and Compiler Warning Fixes
-
Google Chrome Reveals Storage Requirements for Integrated Local AI Models
-
vLLM v0.27.0rc1: Latest Release Candidate for High-Performance Inference
-
LFM2.5-2.6B: On-Device Agentic Model With 128K Context and Tool Calling
-
Gainz.fast – Local Inference, Faster
-
Bubo: AI Code-Reviewer That Learns From Review Comments
-
Ask HN: How Are You Operating OSS AI Infrastructure?
-
llama.cpp b10256 – SYCL SDPA Extended to Quantized KV Caches
-
llama.cpp Build b10258: Sampling Architecture Refinements
-
llama.cpp Release b10257 – Vulkan LLVMpipe Fixes
-
K-EXAONE 2.0 Brings 262K Context to Frontier AI
-
HP Looks to On-Device AI to Reinvent Desktop Computing
-
Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
-
The KV Cache Survival Guide: Why Your GPU Runs Out of Memory with Local LLMs
-
NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework
-
The KV Cache Survival Guide: Why Your GPU Runs Out of Memory with Local LLMs
-
Squeezing Silicon Limits: Effective Strategies to Eliminate GPU Idle Time and Maximize GPU Utilization
-
AI Efficiency Layer Cuts Energy Use and Expands Server Capacity on Existing Hardware
-
Q4 vs Q6 vs Q8: The Quantization Decision Framework for Local LLMs
-
Run Ollama Locally on Windows 11: Setup Guide
-
GPU Half-Idle: The Hundred-Billion-Dollar Race to Squeeze 10x Efficiency from Silicon
-
4 Reasons I'm Canceling My ChatGPT Subscription for Local AI
-
Phi-4 Mini vs Gemma 3 vs Llama 3.2: 128K vs 32K Context Window Comparison
-
Ask HN: What are you using for LLM inference in production?
-
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
-
Titan Transients and LLM Scalability
-
Gemma 4's Quantized Models Finally Made Local AI Practical in Homelab
-
K3 Model Achieves 20 Tokens/Second on 80x RTX 5090 Cluster
-
Netflix Details Its In-House LLM Serving Platform with Triton and vLLM
-
Claude Code Cut System Prompt by 80%: Implications for Small Local Models
-
Code Mode Can Help Smaller LLM Models
-
Show HN: TS Compiler Knowledge Graph Reducing AI Tokens About 90%
-
Nvidia Isn't the Only Choice for Local LLMs Anymore, and AMD Test Proves It
-
How To Build Your Own LLM Runtime From Scratch
-
AI Inference is Rewriting the GPU Buying Playbook
-
Microsoft Strikes Multibillion-Dollar Deal with French AI Firm Mistral
-
llama.cpp b10075 Packs Four Local AI Runtime Upgrades
-
Sunday Reboot: Shrinking Models and an On-Device AI Future
-
LLM Wiki Implementation: Community Resource for Local Deployment
-
AI Data Center Power Constraints Are the Real 2026 Bottleneck
-
Agentic Test Processes and LLM Benchmarks: Evaluating Local AI Agents
-
AI Inference Costs: Build vs. Rent
-
How to Run an LLM Locally: 13 Steps, 90 Min
-
I Thought My Local AI Would Replace My Claude Subscription — Then I Tried Automating My PC
-
Mira Murati's Thinking Machines Launches Open-Weight AI Model
-
ConlangCrafter: Constructing Languages with a Multi-Hop LLM Pipeline
-
Rapid Rise of Open Source Models in the U.S.: Nvidia Nemotron Ultra Grows Quickly on Ollama
-
Show HN: GGUFun, Play Snake and a Simple Maze on Ollama Using Hand Crafted GGUFs
-
WSL Transforms Windows Into a Viable Local LLM Development Platform
-
GitHub Copilot With Ollama: Run Local AI Models In VS Code Offline & Free
-
Intel-Scaler-vLLM 0.21.0-b1 Brings Latest Features for vLLM on Intel GPUs
-
Exploiting Sparsity for Long Context Inference: Million Token on Commodity GPUs
-
Ollama Raises $65M Series B Funding, Reaches Nearly 9 Million Users
-
AMD Lemonade Enables Local AI Portability With New Nvidia Support
-
Viability of Local Models for Coding
-
Ollama is the Easiest Way to Start Local LLMs, But These 6 Alternatives Are Also Worth Trying
-
Compressor V2: Three Compression Layers for 50% LLM Agent Cost Cut
-
LongCat-2.0 Released
-
How to Build Your Own Local AI Server in 2026
-
Intent-Addressable Code for AI Coding Agents
-
Ollama is the Open-Source App That Finally Made Free Local AI Useful on My PC
-
I Quantized a Local LLM on My Home Server and Ditched Cloud AI for Smart Home Control Entirely
-
GLM-5.2's Code Reviews Are Only as Good as Your Prompt
-
How to Choose Between Small and Frontier Models
-
Using Local Coding Agents
-
PewDiePie's Open-Source AI Workspace Gains Traction as Practical Local Deployment Platform
-
Liquid AI Ships LFM2.5-230M with Broad Framework Support for On-Device Inference
-
TriAttention Solves KV Cache Memory Bottleneck in Local LLM Inference
-
Hermes MoA Virtual Models: 8% Higher Than Opus 4.8, 11% Higher Than GPT 5.5
-
Developer Replaces Entire Browser Extension Stack With Single Local LLM
-
ORA: Smaller Models. Same Intelligence
-
NVIDIA DFlash Block Diffusion Accelerates Autoregressive LLM Inference
-
Why Small Local AI Models Get More Use Than Claude or Gemini
-
Developers Run Local LLMs on Windows 11
-
Build Your Own Local AI Coding Agent with Gemma 4 and OpenCode
-
On-Device AI Hardware and Software Acceleration Expected Throughout 2025
-
I Built a Bedside AI Assistant That Reads Me the News Without Touching the Cloud
-
Free Tool Helps Match Local AI Models to Your Hardware
-
Gaming PC vs Phone Local LLM Deployment: Only One Remains in Daily Use
-
Developer Replaces Entire Browser Extension Stack with Single Local LLM
-
Tryll Engine Raises $600K to Deploy On-Device AI Characters in Games
-
Qwen and Fable: Open-Weights 35B Mixture-of-Experts Agentic Coding Model
-
Google's DiffusionGemma Brings Novel Text Generation to Local LLMs
-
Ollama Emerges as Leading Open-Source Local AI Platform
-
AMD Brings Data Center-Level AI Performance to PCs
-
Stop Guessing Which Local AI Models Fit Your Hardware — This Free Tool Does It for You
-
Most People Use Ollama or llama.cpp for Local LLMs, but These Are the Tools I Switch to When It Gets Serious
-
Chrome Downloads 4GB AI Model: Implications for Local On-Device AI
-
RTX 5080 and RTX 3090 Setup Achieves 80 Tok/s on Qwen 3.6 27B Q8
-
What is Ollama? Introduction to the AI Model Management Tool
-
Google's DiffusionGemma Achieves 4x Faster Text Generation for Local Deployment
-
vLLM vs Ollama 2026: 793 vs 41 TPS Performance Benchmark
-
DiffusionGemma: The Developer Guide for Local Deployment
-
AMD's Lemonade SDK Adds NVIDIA CUDA Support for Cross-Platform Local AI
-
Developer Switches from LM Studio to llama.cpp, Citing Performance and Simplicity
-
Google Releases Gemma 4 QAT Models with Reduced Memory Requirements for Mobile and Laptop Deployment
-
Ask HN: What is the AI setup for an experienced dev starting on a new project?
-
NVIDIA Unveils First PC Chips at Computex 2026; CEO Jensen Huang Details New Hardware
-
Best Local LLM Setup for RTX 5090: llama.cpp Fork with TurboQuant
-
Google's New Gemma 4 12B AI Model Is Built for Laptops
-
Running Infinite Context Lengths on 8GB GPU Without Out Of Memory
-
LLM Checker Tool Helps Identify Models for Your PC
-
Maybe Coding Agents Don't Need a Bigger Memory. Maybe They Need Continuity
-
Sawtooth – An Async, Multi-Tiered Memory Framework for LLM Agents
-
Reducing GPU Costs for AI Inference: FP8, FP4, and vLLM Optimization Techniques
-
Google Releases Gemma 4 12B: Encoder-Free Multimodal Model for 16GB Laptops
-
Apple's Overhauled Siri Will Reportedly Run on Nvidia's Blackwell Chips
-
WSL 3 Brings Near-Native GPU and NPU Passthrough for Local AI on Windows
-
NVIDIA RTX Spark Superchip Delivers 6,144 CUDA Cores for Consumer Local AI Inference
-
Tether AI Upgrades QVAC SDK With TurboQuant for Data Center-Sized Memory on Everyday Devices
-
NVIDIA and Microsoft Team Up to Bring Secure On-Device AI Agents to Windows PCs
-
JetBrains Releases Mellum2: A 12B MoE Model for Fast, Specialized Tasks
-
Meet Memory OS: A 6-Layer Open-Source Memory Stack Built on Hermes Agent
-
Nvidia Enters Windows Laptop Market, Taking on Intel and AMD
-
NVIDIA Levels Up Local AI Agents Across RTX PCs and DGX Spark
-
NVIDIA Launches N1X/N1 CPU-GPU SoC for PC Market, Targeting Heavy On-Device AI Users
-
How to Run LLM Locally Without Falling for the Hype
-
GPUs and RAM Are in Short Supply, but the Real Bottleneck for AI Is Electricians
-
Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
-
The Windows Device Manager, on Linux
-
Tweaking Local Language Model Settings with Ollama
-
Mistral AI Launches Mistral Vibe
-
The Anatomy of an LLM
-
Meet EAGLE 3.1: The Speculative Decoding Algorithm That Fixes Attention Drift in LLM Inference
-
Users Report Superior Performance Switching from LM Studio to llama.cpp
-
Gemma 4: A New Budget-Focused Model in Posit AI
-
AMD Unveils Ryzen AI Halo Developer Platform for On-Device AI Workloads
-
New 8B Local LLM Design Marks Biggest Shift Since DeepSeek R1
-
M5 Max MacBook Runs Local Large Language Models Efficiently
-
110 Tokens/Second on RTX 4070 Super with Qwen 3.6 35B
-
Benchmarking a Portable AI Workstation: Lenovo ThinkPad P16 Gen 3, Part 2
-
Intel llm-scaler-vllm 1.4 Released With Updated Components and Arc Pro B70 Support
-
AMD's New Ryzen AI Max Pro 400 with 192GB LPDDR5X Memory
-
Adobe Photoshop Update Brings On-Device AI Processing
-
Nvidia Raises Video Encoder Limit to 12 on Consumer GPUs
-
I Stopped Trying to Replace My Cloud LLMs, and Local Models Finally Made Sense
-
Open Source Local Audio Stem Separation Tool Released
-
llama.cpp Adds Multi-Token Prediction, Doubles Qwen 3.6B Throughput for Local Inference
-
Local LLMs Enable Intelligent Smart Camera Control Without Cloud Dependency
-
AMD's Lemonade SDK Advances macOS Support for Local AI Inference with ROCm 7.13
-
A Lo-Fi Rebellion Against A.I
-
MegaTrain: Full Precision Training of 100B+ Parameter LLMs on a Single GPU
-
AI/ML Benchmark Tool for Local LLM Inference and XGBoost Training
-
Show HN: Find the best local LLM for your hardware, ranked by benchmarks
-
Open-Source Local LLM Emerges as Viable Cloud AI Competitor
-
llama.cpp Delivers Sharp Performance Gains for AMD RDNA3 Users
-
Kog AI – Building a Real-Time Inference Stack on AMD Instinct GPUs
-
Local LLM Persistent Context Prevents Repetitive Mistakes
-
I Stopped Paying for ChatGPT and Switched to a Local LLM That Runs on My Laptop
-
BT Explainer: Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
$200 NVIDIA V100 Server GPU Mod Beats RTX 3060 in Local LLM Test
-
MDL: Endless Visual Novel Engine Powered by AI
-
One LM Studio Setting Change Makes Local LLMs Competitive With Cloud Models
-
Cotypist – AI Autocomplete for Mac
-
DistillFast: AI Cost Optimization Tool for Model Efficiency
-
Small On-Device AI Model Beats Claude Sonnet 4.5 and GPT-5
-
Lemonade Gives AMD Startups a Wider Path to Local Inference
-
Microsoft VibeVoice C++ Port Enables Local Voice AI on CPU and GPU Without Python
-
Improving Code Quality with Local Claude and Codex Models
-
Google Accelerates Gemma 4 Inference Speed 3x With Multi-Token Prediction Drafters
-
5 Things I Wish Someone Had Told Me Before I Tried Self-Hosting a Local LLM
-
llama.cpp Now Supports Multi-Token Prediction in Beta
-
Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
Supercharging LLM Inference on Google TPUs: Achieving 3X Speedups With Diffusion-Style Speculative Decoding
-
Gemma 4 Just Replaced My Whole Local LLM Stack
-
Running a Serious AI Model on a Consumer GPU Just Got Easier and That Matters More Than the Benchmark
-
Local AI Just Got Easier on Windows and the Implications Go Beyond the Benchmark
-
PFlash Claims 10x Prefill Speedup Over llama.cpp
-
Local LLMs Work Best When You're Not Loyal to Just One
-
AMD Posts HDMI 2.1 FRL Patches for Amdgpu Linux Driver
-
New Open-Source Tool Automatically Matches Local LLMs to Your PC Hardware
-
Xmemory: Benchmarking Structured AI Memory Against RAG and Hybrid RAG
-
Linux Setup for Local LLMs Takes Minutes Compared to Windows Hours
-
Running Capable Local LLMs Without Expensive GPU Hardware
-
IBM Introduces Granite 4.1 Family of Models for Local Deployment
-
Google's Gemma 4 Brings Powerful AI Capabilities to Phones and Laptops
-
Show HN: Arkloop – Open-Source, Local-First Agent Client
-
How Much "Brain Damage" Can an LLM Tolerate?
-
Building a Local AI Stack: Five Docker Containers to Replace ChatGPT Subscriptions
-
Hipfire: A Rust-Native AMD Inference Engine That Outperforms llama.cpp
-
Google's Gemma 4: Powerful AI Models Optimized for Your Phone and Laptop
-
Economic Implications of AI Adoption: Why Local Deployment Matters for Cost Control
-
Local AI Isn't Just Ollama—Here's the Ecosystem That Actually Makes It Useful
-
Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
Unsloth's Custom Kernels Make LLM Fine-Tuning Viable on Consumer GPUs
-
Pluggable's TBT5-AI: First Thunderbolt Dock Explicitly Targeting Local LLM Workstations
-
Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
Elastic KV Cache Memory Breakthrough Enables Efficient Bursty LLM Serving and GPU Sharing
-
Fixing Hallucination in LLM Prediction With Only One 48GB GPU
-
Google's Gemma 4 Brings Powerful On-Device AI to Phones and Laptops
-
GPU Passthrough to LXCs in Proxmox Outperforms VMs and Simplifies Local AI Infrastructure
-
I Replaced My Local LLM With a Model Half Its Size and Got Better Results
-
Llama 4 Scout on MLX: The Complete Apple Silicon Guide (2026)
-
Intel OpenVINO 2026.1 Integrates llama.cpp with Wildcat Lake and Arc Pro B70
-
Intel LLM-Scaler vLLM 0.14.0 Released With Official Arc Pro B70 Support
-
10GB VRAM Local LLM: The Complete Setup Guide (2026)
-
Show HN: We built an OCR server that can process 270 dense images/s on a 5090
-
Externalization in LLM Agents: Unified Review of Memory and Harness Engineering
-
Llama.cpp's Auto Fit Feature Quietly Reshapes Local AI Inference on Consumer Hardware
-
Google's Gemma 4 Finally Makes Local LLM Deployment Compelling for Practitioners
-
The Open-Source AI Ecosystem Keeps Treating llama.cpp Like a Second-Class Citizen
-
ZeusHammer: Built an AI Agent That Thinks Locally
-
llama.cpp Merges Speculative Checkpointing for Major Inference Speed Boost
-
Intel Extends AI PC Reach With New Core Ultra Series 3 Launch
-
Running DeepSeek R1 Locally: Your Complete Setup Guide
-
PCMind: Local AI Analysis of Docs, Audio, Video and Images
-
Gemma 4 Just Replaced My Whole Local LLM Stack
-
Show HN: I Can't Write Python. It Works Anyway – Local LLM Automation
-
Unweight: Lossless MLP Weight Compression for LLM Inference
-
We Built a Local Model Arena in 30 Minutes — Infrastructure Mattered More Than the App
-
Laimark – 8B LLM That Self-Improves on Consumer GPUs
-
Sorting 1M u64 KV-Pairs in 20ms on i9-13980HX Using Branchless Rust Implementation
-
Intel's $949 GPU Has 32GB of VRAM for Local AI, but the Software Is Why Nvidia Keeps Winning
-
Community Computer: Collaborative Autoresearch on a Peer-to-Peer Network
-
Google's Gemma 4: The Most Practical Local LLM Despite Not Being The Smartest
-
Prefill Is Compute-Bound, Decode Is Memory-Bound: Optimizing GPU Utilization for LLM Inference
-
Noi Enables Running ChatGPT and Claude Side-by-Side on Your Desktop
-
GPU Passthrough to LXCs in Proxmox Simplifies Local Inference Infrastructure
-
Google's Gemma 4 Brings Game-Changing Performance to Local Laptop Inference
-
SigMap – Shrink AI Coding Context 97% with Auto-Scaling Token Budget
-
Dynamic Expert Cache in llama.cpp Achieves 27% Faster Inference on Large MoE Models
-
Sovereign AI: Why the Next GPT Will Be Born in Our Living Rooms
-
Qwen 3.5 Small – On-Device Multimodal Models Released
-
MiniMax M2.7 Achieves SOTA Performance Under 64GB on Mac with TQ Quantization
-
Speculative Decoding Achieves 29% Speed Boost for Gemma-4 31B
-
Qwen3 Audio and Vision Support Now Available in llama.cpp
-
Audio Processing Support Lands in llama.cpp with Gemma-4
-
MiniMax-M2.7 Delivers Exceptional Performance on Consumer Hardware
-
Unsloth Completes Comprehensive MiniMax M2.7 GGUF Quantization Suite
-
On-Device AI: Achieving Powerful AI Capabilities Without Internet Connectivity
-
MiniMax M2.7 Released: New Model Available for Local Deployment
-
Google Gemma 4 Delivers Exceptional Speed and Accuracy for Local Inference
-
A Deep Dive into Tinygrad AI Compiler
-
MiniMax M2.7 Advances Scalable Agentic Workflows on NVIDIA Platforms for Complex AI Applications
-
DFlash Speculative Decoding Achieves 3.3x Speedup on Apple Silicon
-
Intel Arc Pro B70 32GB Achieves 12 Tokens/Sec on Qwen 3.5-27B
-
Gemma 4 31B vs Qwen 3.5 27B: Comprehensive Long Context Benchmark
-
ASUS ExpertBook P1 Integrates On-Device AI for Enterprise Collaboration
-
AIYO Wisper: Local Voice-to-Text for macOS Using WhisperKit
-
5 Open-Source Projects Running Transformers on CPUs to GPUs in Pure Java
-
Warp Decode vs. vLLM's Triton Kernel: Performance Crossover Analysis
-
Qwen 3.5 122B Achieves 198 Tokens/sec on Dual RTX PRO 6000 Blackwell GPUs
-
Energy Consumption: The Final Frontier for AI and Local Inference
-
VoxCPM2: New Open-Source TTS Model with Voice Cloning and Design
-
I Replaced My Local LLM With a Model Half Its Size and Got Better Results — and It Wasn't About the Parameters
-
Intel Releases OpenVINO 2026.1 With Backend For Llama.cpp, New Hardware Support
-
Gemma 4 Support Stabilized in Llama.cpp
-
Gemma 4 GGUF Models Updated with Critical Quantization Fixes
-
EXAONE 4.5 33B Model Released with Multiple Quantization Formats
-
Speculative Decoding Made My Local LLM Actually Usable
-
Google's Gemma 4 Brings Powerful On-Device AI to Android and iOS
-
Running AI Natively on Windows 11 Using an eGPU
-
Quansloth Using Google's Turboquant Breaks the VRAM Wall for Local LLMs
-
Your Next Assistant is Your PC: How On-Device AI is Transforming Work, One Workflow at a Time
-
Gemma 4 26B Achieves Impressive Local Performance With Proper Configuration
-
AMD Announces Day 0 Support for Google Gemma 4 Across Processors and GPUs
-
TurboQuant-Optimized llama.cpp Fork Delivers GFX906 GPU Acceleration
-
Verbatim 140W GAN: One of the First Chargers With USB PD 3.2 AVS (SPR) Support
-
TurboQuant in Llama.cpp Achieves 6X Smaller KV Cache
-
Show HN: Lightweight LLM Tracing Tool with CLI
-
HunyuanOCR 1B: High-Quality OCR Now Viable on Budget Consumer Hardware
-
Google AI Edge Gallery Tops App Store Charts with On-Device Gemma 4
-
Real-time Multimodal AI on Apple Silicon: Gemma E2B Demo Shows Practical Edge Deployment
-
Gemma 4 31B Achieves Exceptional Performance on Local Hardware
-
Show HN: Turn Photos Into Wordle Puzzles with AI That Runs 100% in Your Browser
-
Quantization Strategy Comparison: Balancing Quality and Speed on Consumer Laptops
-
Context Window Optimization: Extending Gemma 4 Context Length Through Efficient Projection Quantization
-
GPU Memory for LLM Inference (Part 1)
-
Gemma 4 31B Achieves Third Place on FoodTruck Bench, Beating Larger Models
-
Gemma 4 26B MoE Emerges as Optimal All-Around Local Model for Consumer Hardware
-
Qwen 3.5 397B Reduced to 35% Parameters With Usable Quality on 96GB GPU
-
DGX Spark Hardware Limitations: Missing NVFP4 Support Undermines Local AI Value Proposition
-
Samsung Launches Galaxy Book6 Series with NVIDIA RTX 5070 and On-Device AI
-
NVIDIA and Google Optimize Gemma 4 AI Models for Local RTX Deployment
-
GPUs vs. TPUs: Decoding the Powerhouses of AI
-
Google Launches Gemma 4 For Advanced On-Device AI
-
Gemma 4 31B Outperforms GLM 5.1 in Real-World Testing
-
Gemma 4 KV Cache Memory Issues Fixed in llama.cpp
-
AMD Rolls Out Gemma 4 Model Support Across Full Range of GPUs & CPUs
-
SkillCompass – Diagnose and Improve AI Agent Skills Across 6 Dimensions
-
NVIDIA Accelerates Gemma 4 for Local Agentic AI on RTX GPUs
-
VRAM Optimization Technique Cuts Gemma 4 Memory Usage by 3x
-
Google Gemma 4 Released with GGUF Quantizations
-
Google Launches Gemma 4 Open Models for Local On-Device AI
-
Gemma 4 Makes Local AI Agents Practical
-
AMD Provides Day 0 Support for Gemma 4 on Ryzen AI Processors and GPUs
-
OpenUMA – Apple-Style Unified Memory for x86 AI Inference
-
TinyGPU Adds Mac Support for External Nvidia GPU Acceleration
-
Intel's $949 GPU Has 32GB of VRAM for Local AI, but Software is Why Nvidia Keeps Winning
-
Show HN: Extra-Platforms, Python Library to Detect OS, Arch, Shell, CI, AI
-
Bonsai 1-Bit Models Deliver Exceptional Local Inference Performance
-
TurboQuant Enables Qwen 3.5-27B on 16GB Consumer GPUs
-
Apple Silicon Macs Run Local AI Faster with Ollama's New MLX Support
-
Qwen 3.5-27B Demonstrates Superior Performance vs Gemini 3.1 Pro and GPT-5.3
-
Intel's Arc GPU Offers 32GB VRAM for Local AI, But Software Ecosystem Lags Behind
-
ByteShape Releases Qwen 3.5 9B Quantisations with Hardware-Matched Tuning Guide
-
Is Anyone Working on an AI Operating System?
-
ROCm Integration in Ubuntu 26.04 Advances Linux GPU Inference
-
Samsung launches Galaxy Book6 series in India with Nvidia RTX 5070 graphics and on-device AI
-
Intel's $949 GPU has 32GB of VRAM for local AI, but the software is why Nvidia keeps winning
-
Select the Right Hardware for Your Local LLM Deployment with This Online Guide
-
Samsung Launches Galaxy Book6 Series in India with NVIDIA RTX 5070 Graphics and On-Device AI
-
Dell Technologies Unveils 10 AI PC Models for Business, from Ultralight Laptops to Ultracompact Desktops
-
DeepSeek V3 Complete Guide: Deploy and Optimize Local AI in 2026
-
Google's TurboQuant Shows Memory Constraints Remain Critical for Local LLM Inference
-
Samsung Galaxy Book6 Brings Consumer-Grade On-Device AI Hardware to Market
-
IBM Granite 4.0 3B Vision: Compact Enterprise-Grade Document AI
-
DaVinci-MagiHuman: Open-Source AI Model for Realistic Video Generation
-
TurboQuant: Understanding the Quantization Breakthrough
-
Scion: Running Concurrent LLM Agents with Isolated Identities and Workspaces
-
Mixed KV Cache Quantization: Performance Risks and Pitfalls
-
Samsung Galaxy Book6 Series Brings Intel Core Ultra Chips for On-Device LLM Inference
-
TurboQuant KV Cache Compression Achieves 22.8% Faster Decoding at 32K Context
-
Qwen3 512k Context via TurboQuant on Mac mini
-
GPU Passthrough to LXCs in Proxmox Simplifies Local LLM Deployment
-
Coding Implementation to Run Qwen3.5 Reasoning Models Distilled With Claude-Style Thinking Using GGUF and 4-Bit Quantization
-
Hold on to Your Hardware: Implications for Local LLM Deployment
-
TurboQuant Benchmarked in Llama.cpp: Google's Extreme Compression Research Tested in Practice
-
RotorQuant: 10-19x Faster Quantisation Alternative Using Clifford Algebra
-
Pluggable's TBT5-AI: First Thunderbolt Dock Explicitly Targeting Local LLM Workstations
-
Show HN: Beforeyouship – Pre-Build Tool to Estimate LLM Cost
-
Liquid AI's LFM2-24B Achieves 50 Tokens/Second in Web Browser via WebGPU
-
Intel Launches Arc Pro B70/B65 with 32GB VRAM for Local AI Inference
-
Google's TurboQuant: The Unsexy AI Breakthrough Worth Watching
-
NVIDIA Releases GPT-OSS-Puzzle-88B, a Deployment-Optimized Model
-
OmniCoder v2 Released: Improved Code Generation for Local Deployment
-
Researcher Successfully Runs Local LLMs on Legacy "Dead" GPU With Surprising Results
-
Google TurboQuant: Extreme Compression for Local LLM Deployment
-
Llama.cpp Benchmark: RTX 5090 vs Enterprise Systems Compared
-
Running a Private AI Brain on Windows PC as Alternative to Cloud Services
-
Powerful AI Search Engine Built on Single GeForce RTX 5090
-
Ditching Paid AI Services: Building Self-Hosted LLM Solutions as ChatGPT, Claude, and Gemini Alternatives
-
Qwen 3.5 122B Uncensored (Aggressive) Released with New K_P Quantisations
-
Nvidia Nemotron Cascade 2 30B Emerges as Powerful Alternative to Qwen Models
-
Why You Should Use Both ChatGPT and Local LLMs: A Practical Hybrid Approach
-
Careless Whisper – Personal Local Speech to Text
-
AI Playground for Developers Built in Vite and Python
-
Rust Project Perspectives on AI
-
Developer Builds Fully Local Multi-Agent System Using vLLM and Parallel Inference
-
Llama 8B Matches 70B Performance on Multi-Hop QA Using Structured Prompting
-
ik_llama.cpp Fork Delivers 26x Faster Prompt Processing on Qwen 3.5 27B
-
Multi-Token Prediction support coming to MLX-LM for Qwen 3.5
-
DeepSeek R1 RTX 4090 vs Apple M3 Max: Benchmark & Performance Guide
-
Build a $1,500 AI Server with DeepSeek-R1 on RTX 4090
-
Qwen 3.5 397B emerges as top-performing local coding model
-
Apple M5 Max 128GB real-world performance benchmarks for local inference
-
Local AI Coding Assistant: Free Cursor Alternative with VS Code, Ollama & Continue
-
Qwen 3.5 Emerges as Top Performer for Local Deployment with Extensive Quantization Options
-
Repurpose Old GPUs as Dedicated AI Inference Accelerators
-
NVIDIA Nemotron Cascade 2 30B Delivers 120B-Class Performance in Compact Form Factor
-
Llamafile 0.10 Released with GPU Support and Rebuilt Core
-
Community Converges on Optimal KV Cache Quantization Strategies for Qwen 3.5 Models
-
NVIDIA Nemotron 3 Nano 4B Enables On-Device Inference Directly in Web Browsers via WebGPU
-
Meet Sarvam Edge: India's AI Model That Runs on Phones and Laptops With No Internet
-
Tether's QVAC Introduces Cross-Platform Bitnet LoRA Framework for On-Device AI Training
-
Unsloth Studio: Open-Source Web UI for Training and Running LLMs Locally
-
Snapdragon 8 Elite Gen 5 Hands the Galaxy S26 the AI Upgrade We've Been Waiting For
-
MiniMax-M2.7: New Compact Model Announced for Local Deployment
-
I Switched to a Local LLM for These 5 Tasks and the Cloud Version Hasn't Been Worth It Since
-
Mamba 3: State Space Model Architecture Optimized for Inference
-
Custom GPU Multiplexer Achieves 0.3ms Model Switching on Legacy Hardware
-
Run LLMs Locally with Llama.cpp
-
I Ran Local LLMs on a 'Dead' GPU, and the Results Surprised Me
-
Qwen 3.5 4B Outperforms Nvidia Nemotron 3 4B in Local Benchmarks
-
Mistral Small 4 119B Released with NVFP4 Quantisation Support
-
Mistral Releases Small 4 Open-Source Model Under Apache 2.0
-
Kimi Introduces Attention Residuals: 1.25x Compute Performance at <2% Overhead
-
OpenClaw Isn't the Only Raspberry Pi AI Tool—Here Are 4 Others You Can Try This Week
-
OmniCoder-9B: Efficient Coding Model for 8GB GPUs
-
This External GPU Enclosure Tries to Break Cloud Dependence for Local AI Inference
-
Dictare – Open-source Voice Layer for AI Coding Agents (100% Local)
-
AMD Declares 'AI on the PC Has Crossed an Important Line' – Agent Computers as Next Breakthrough
-
Nvidia's Nemotron 3 Super: Understanding the Significance for Local LLM Deployment
-
Running Qwen3.5-27B Across Multiple GPUs Over LAN Achieves Practical Speed for Local Inference
-
Startup Transforms Mac Mini Into Full-Powered AI Inference System With External GPU
-
Two Local Models Prove Competitive Enough to Replace ChatGPT, Gemini, and Copilot
-
India's Mobile-First AI Strategy Could Accelerate Local Inference Adoption in Emerging Markets
-
Hybrid AI Desktop Layer Combining DOM-Automation and API-Integrations
-
Open-Source GreenBoost Driver Augments NVIDIA GPU VRAM With System RAM and NVMe Storage
-
AMD Launches Agent System Optimized for Local AI Inference With Ryzen and Radeon
-
Achieving 2000 Tokens Per Second with QWEN 3.5 27B on RTX-5090
-
Best Local LLM Models 2026: Developer Comparison
-
P-EAGLE: Faster LLM Inference with Parallel Speculative Decoding in vLLM
-
Local Manga Translator: Production LLM Pipeline with YOLO, OCR, and Inpainting
-
3-Path Agent Memory: 8 KB Recurrent State vs. 156 MB KV Cache at 10K Tokens
-
Linux 7.0 AMDGPU Fixing Idle Power Issue For RDNA4 GPUs After Compute Workloads
-
How to Install OpenClaw with Ollama (Step-by-Step Tutorial)
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Qwodel – An Open-Source Unified Pipeline for LLM Quantization
-
Nvidia Pushes Jetson as Edge Hub for Open AI Models
-
Apple M5 Max 128GB Benchmark Results for Local LLM Inference
-
The $1,500 Local AI Setup: DeepSeek-R1 on Consumer Hardware
-
Show HN: VmExit – An Experiment in AI-Native Computing
-
Quantization Explained: Q4_K_M vs AWQ vs FP16 for Local LLMs
-
Nvidia Releases Nemotron 3 Super: 120B MoE Model for Local Deployment
-
Cutile.jl Brings Nvidia CUDA Tile-Based Programming to Julia
-
Local AI Coding Assistant: Complete VS Code + Ollama + Continue Setup
-
Experiment: 0.8B Model Self-Improvement on MacBook Air Yields Surprising Results
-
Texas Instruments Launches NPU-Powered MCUs for Low-Power Edge AI
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Qwen 3.5-35B Uncensored GGUF Models Now Available
-
Llama.cpp Celebrates Major Milestone: From Leak to Industry Standard
-
Qwen 3.5 Ultra-Compact Models Enable On-Device AI from Watches to Gaming
-
HP OMEN MAX 16 Review: Is Local AI on a Laptop Viable in 2026?
-
Fine-Tuned Qwen SLMs (0.6–8B) Demonstrate Competitive Performance Against Frontier LLMs on Specialized Tasks
-
Strix Halo (Ryzen AI Max+ 395) Achieves Strong Local Inference Performance with ROCm 7.2
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Qwen 3.5 Family Benchmark Comparison Shows Strong Performance Across Smaller Models
-
Qwen 3.5 Derestricted Model Available for Local Deployment
-
Nemotron 9B Powers Large-Scale Local Inference: Patent Classification and Real-Time Applications
-
Gyro-Claw – Secure Execution Runtime for AI Agents
-
Engram – Open-Source Persistent Memory for AI Agents
-
When Running Ollama on Your PC for Local AI, One Thing Matters More Than Most
-
Qwen 3.5 27B Achieves Strong Local Inference Performance
-
Mistral AI Prepares Workflows Integration for Le Chat
-
Apple Launches MacBook Neo with A18 Pro Chip for Affordable Local AI Inference
-
Llama.cpp Prompt Processing Optimization: Ubatch Size Configuration Guide
-
ETH Zurich Research Challenges Context-Length Assumptions in LLM Agents
-
Windows 11 Notepad Gets On-Device AI Text Generation Without Subscription
-
Alibaba Releases Qwen 3.5 AI Model with On-Device AI Support
-
Final Qwen3.5 Unsloth GGUF Update with Improved Size/Quality Tradeoffs
-
Alibaba Releases Qwen 3.5 AI Model with On-Device AI Support
-
The Emerging Role of SRAM-Centric Chips in AI Inference
-
Kakao Launches Kanana AI for On-Device Schedule and Recommendation Management
-
Qwen 3.5-35B-A3B Achieves 37.8% on SWE-bench Verified Hard
-
On-Device AI Laptop Lineups Become Standard Across Major Manufacturers
-
AMD Launches Copilot+ Desktop Chips to Compete in On-Device AI Market
-
Qwen 3.5-27B Q4 Quantization Comparison and Analysis
-
VibeWhisper – macOS Voice-to-Text with 100% Local Processing Option
-
Qwen 3.5 0.8B Running in Browser with WebGPU via Transformers.js
-
Intel Arc Pro B70 Workstation GPU Confirmed via vLLM AI Release Notes
-
AMD Ryzen AI 400 Series Desktop Processors Launch with Integrated 60 TOPS NPU
-
Local LLM Performance Improvements: A Year of Progress Since DeepSeek R1 Moment
-
Jan Releases Code-Tuned 4B Model for Efficient Local Code Generation and Development Tasks
-
HP ZBook Ultra 14 G1a Workstation Reclaims Local AI Workflows for Professionals
-
Browser Use vs. Claude Computer Use: Comparing Agent Automation Frameworks
-
Apple Neural Engine Reverse-Engineered for Local Model Training on Mac Mini M4
-
Qwen 3.5-35B-A3B Emerges as Efficient Daily Driver, Replacing 120B Models
-
Nummi – AI Companion with Memory and Daily Guidance
-
Apple Intelligence, Galaxy AI, Gemini: Why Your AI-Powered Phone Is Worth Repairing
-
4 Free Tools to Run Powerful AI on Your PC Without a Subscription
-
Unsloth Dynamic 2.0 GGUFs
-
Qwen 3.5-27B Demonstrates Exceptional Performance with Thoughtful Prompt Engineering
-
On-Device AI in Mobile Apps: What Should Run on the Phone vs the Cloud (A 2026 Decision Guide)
-
The ML.energy Leaderboard
-
LLmFit: Terminal Tool for Right-Sizing LLM Models to Your Hardware
-
LLmFit: One-Command Hardware-Aware Model Selection Across 497 Models and 133 Providers
-
Accuracy vs. Speed in Local LLMs: Finding Your Sweet Spot
-
Qwen3.5-35B Unsloth Dynamic GGUFs Achieve SOTA Across Nearly All Quantisation Levels
-
Qwen3.5-35B RTX 5080 Experiments Confirm KV q8_0 as Free Lunch, Q4_K_M Remains Optimal
-
Krasis: Hybrid CPU/GPU MoE Runtime Achieves 3,324 Tokens/Second Prefill on RTX 5080
-
5 Useful Docker Containers for Agentic Developers
-
Show HN: Caret – Tab to Complete at Any App on Your Mac
-
Qwen3.5 122B Achieves 25 tok/s on 72GB VRAM Setup
-
Qwen 3.5 MoE Delivers 100K Context Window at 40+ TPS on RTX 5060 Ti
-
Qwen 3.5 Underperforms on Hard Coding Tasks—APEX Benchmark Analysis
-
Researchers Develop Persistent Memory System for Local LLMs—No RAG Required
-
DeepSeek Releases DualPath: Addressing Storage Bandwidth Bottlenecks in Agentic Inference
-
DeepSeek Paper – DualPath: Breaking the Bandwidth Bottleneck in LLM Inference
-
The Complete Developer's Guide to Running LLMs Locally: From Ollama to Production
-
Qwen3.5-35B-A3B Emerges as Game-Changer for Agentic Coding Tasks
-
Qwen3.5-27B Identified as Sweet Spot for Mid-Range Local Deployment
-
PyTorch Foundation Announces New Members as Agentic AI Demand Grows
-
Show HN: Pluckr – LLM-Powered HTML Scraper That Caches Selectors and Auto-Heals
-
Mirai Announces $10M to Advance On-Device AI Performance for Consumer Devices
-
Show HN: 100% LLM Accuracy–No Fine-Tuning, JSON Only
-
How AI is Redefining Price and Performance in Modern Laptops
-
Advanced Quantization Techniques Show Surprising Performance Gains Over Standard Methods
-
Enterprise Infrastructure Guide: Running Local LLMs for 70-150 Developers
-
South Korea to Launch $687 Million Project to Develop On-Device AI Semiconductors
-
Qwen3's Voice Embeddings Enable Local Voice Cloning and Mathematical Voice Manipulation
-
A Tool to Tell You What LLMs Can Run on Your Machine
-
Open-Source llama.cpp Finds Long-Term Home at Hugging Face
-
GPT-OSS 20B Demonstrates Practical Agentic Capabilities Running Fully Locally
-
GLM-5 Becomes Top Open-Weights Model on Extended NYT Connections Benchmark
-
Elastic Introduces Best-in-Class Embedding Models for High Performance Semantic Search
-
Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference
-
Custom Portable Workstation Optimized for Local AI Inference Builds
-
Open-Source Framework Achieves Gemini 3 Deep Think Level Performance Through Local Model Scaffolding
-
Nvidia Could Launch Its First Laptops With Its Own Processors
-
Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference
-
Yet Another Fix Coming for Older AMD GPUs on Linux – Thanks to Valve Developer
-
At India AI Impact Summit, Intel Showcases AI PCs and Cost-Efficient Frugal AI
-
O-TITANS: Orthogonal LoRA Framework for Gemma 3 with Google TITANS Memory Architecture
-
Ouro 2.6B Thinking Model GGUFs Released with Q8_0 and Q4_K_M Quantization
-
Strix Halo Performance Benchmarks: Minimax M2.5, Step 3.5 Flash, Qwen3 Coder
-
At India AI Impact Summit, Intel Showcases Its AI PCs and Cost-Efficient Frugal AI
-
GGML.AI Acquired by Hugging Face
-
Qwen3 Coder Next Remains Effective at Aggressive Quantization Levels
-
[Release] Ouro-2.6B-Thinking: ByteDance's Recurrent Model Now Runnable Locally
-
PaddleOCR-VL Now Integrated into llama.cpp for Multilingual OCR
-
NVIDIA Releases Dynamo v0.9.0: Infrastructure Overhaul With FlashIndexer and Multi-Modal Support
-
Mirai Secures $10M to Optimize On-Device AI Amid Cloud Cost Surge
-
Qwen3 Coder Next 8FP Demonstrates Exceptional Long-Context Performance on 128GB System
-
Free ASIC-Accelerated Llama 3.1 8B Inference at 16,000 Tokens/Second
-
AI Integration in Sublime Text: Practical Local LLM Editor Enhancement
-
Enhanced Quantization Visualization Methods for Understanding LLM Compression Trade-offs
-
LayerScale Launches Inference Engine Faster Than vLLM, SGLang, and TRT-LLM
-
Hardware Economics Shift: DDR5 RDIMM Pricing Now Comparable to GPUs for Local Inference
-
Cohere Releases Tiny Aya: Efficient 3.3B Multilingual Model for 70+ Languages
-
ASUS Zenbook 14 Launches in India with AI-Capable Hardware, Starting at Rs 1,15,990
-
Ask HN: What is the best bang for buck budget AI coding?
-
Qwen3-Next 80B MoE Achieves 39 Tokens/Second on RTX 5070/5060 Ti Dual-GPU Setup
-
Qwen 3.5-397B-A17B Now Available for Local Inference with Aggressive Quantisation
-
High Bandwidth Flash Memory Could Alleviate VRAM Constraints in Local LLM Inference
-
GPU-Accelerated DataFrame Library for Local Inference Workloads
-
Alibaba Unveils Major AI Model Upgrade Ahead of DeepSeek Release
-
GPT-OSS 20B Now Runs 100% Locally in Browser via WebGPU
-
GNOME's AI Assistant Newelle Adds llama.cpp Support and Command Execution
-
NVIDIA's Dynamic Memory Sparsification Cuts LLM Inference Costs by 8x
-
MiniMax-M2.5 230B MoE Model Released with GGUF Support for Local Deployment
-
LLaDA2.1 Introduces Token Editing for Massive Speed Gains in Local Inference
-
Context Management Identified as Real Bottleneck in AI-Assisted Coding
-
Ring-1T-2.5 Released with SOTA Deep Thinking Performance
-
The Future of AI Slop Is Constraints - Implications for Local Models
-
Running Mistral-7B on Intel NPU Achieves 12.6 Tokens/Second
-
I Tried a Claude Code Rival That's Local, Open Source, and Completely Free
-
Samsung's REAM: Alternative Model Compression Technique
-
GLM-5 Released: 744B Parameter MoE Model Targeting Complex Tasks
-
NAS System Achieves 18 tok/s with 80B LLM Using Only Integrated Graphics
-
Nanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Acts
-
Carmack Proposes Using Long Fiber Lines as L2 Cache for Streaming AI Data
-
Community Member Builds 144GB VRAM Local LLM Powerhouse