Tagged "hardware-optimization"
-
Ollama Adds Qwen 3.8 27B with Apple Silicon Optimizations
-
Meta's Muse Glimmer Now Available Across All Platforms in Ollama
-
AI Efficiency Layer Cuts Energy Use and Expands Server Capacity on Existing Hardware
-
GPU Half-Idle: The Hundred-Billion-Dollar Race to Squeeze 10x Efficiency from Silicon
-
Nvidia Accelerates Chip Engineering with AI Agents
-
AI Inference is Rewriting the GPU Buying Playbook
-
llama.cpp's 4.26× Intel Gain Has a Narrow Catch
-
Developer Ditches Ollama for llama.cpp's WebUI: A Practical Comparison
-
2026 On-Device AI Market Intensifies: Apple, Google, and Samsung Compete for Local AI Dominance
-
Snapdragon C Specs Revealed: 6nm Process, On-Device AI Engine for Budget Laptops
-
The Anatomy of an LLM
-
Qualcomm's AI-Device Strategy Reflects Growing Market Momentum in On-Device Intelligence
-
Why AI Hardware Is a Chip Layer Problem
-
Nvidia Raises Video Encoder Limit to 12 on Consumer GPUs
-
Open Source Local Audio Stem Separation Tool Released
-
On-Device AI to Be in 80% of Wearables by 2032
-
ROCm 7.2.3 Delivers Performance Improvements Over 7.0.0 on AMD Radeon AI PRO
-
Show HN: Find the best local LLM for your hardware, ranked by benchmarks
-
llama.cpp Delivers Sharp Performance Gains for AMD RDNA3 Users
-
Mlx-serve: Run LLMs Natively on Your Mac
-
Running Capable Local LLMs Without Expensive GPU Hardware
-
IBM Introduces Granite 4.1 Family of Models for Local Deployment
-
Blueprint: AI Hardware Design
-
Google's Gemma 4 Brings Powerful On-Device AI to Phones and Laptops
-
I Replaced My Local LLM With a Model Half Its Size and Got Better Results
-
AI Agent Designs a RISC-V CPU Core from Scratch
-
Intel OpenVINO 2026.1 Integrates llama.cpp with Wildcat Lake and Arc Pro B70
-
Controlling the Secondary Fan on Minisforum AI Pro HX 370
-
Laimark – 8B LLM That Self-Improves on Consumer GPUs
-
Show HN: SkillCompass – Open-Source Quality Evaluator for Your AI Skills
-
The Best Local AI Model for Home Assistant Isn't Always the Biggest One
-
Unsloth Completes Comprehensive MiniMax M2.7 GGUF Quantization Suite
-
MiniMax M2.7 Advances Scalable Agentic Workflows on NVIDIA Platforms for Complex AI Applications
-
Parakeet Streaming ASR on Apple Silicon via CoreML
-
Building Offline AI Companions on Severely Constrained Hardware (8GB RAM)
-
Gemini-CLI, Llama.cpp, and Qwen3.5 Running on NVIDIA Jetson TK1
-
Quansloth Using Google's Turboquant Breaks the VRAM Wall for Local LLMs
-
Your Next Assistant is Your PC: How On-Device AI is Transforming Work, One Workflow at a Time
-
Quantization Strategy Comparison: Balancing Quality and Speed on Consumer Laptops
-
Qualcomm Snapdragon Innovations Enable Advanced On-Device AI for Wearables
-
VRAM Optimization Technique Cuts Gemma 4 Memory Usage by 3x
-
Google Launches Gemma 4 Open Models for Local On-Device AI
-
Bonsai 1-Bit Models Deliver Exceptional Local Inference Performance
-
Apple Silicon Macs Run Local AI Faster with Ollama's New MLX Support
-
ByteShape Releases Qwen 3.5 9B Quantisations with Hardware-Matched Tuning Guide
-
Dell Technologies Unveils 10 AI PC Models for Business, from Ultralight Laptops to Ultracompact Desktops
-
TurboQuant KV Cache Compression Achieves 22.8% Faster Decoding at 32K Context
-
Nota AI and SiMa.ai Partner on Physical AI Technology for Local Deployment
-
Researcher Successfully Runs Local LLMs on Legacy "Dead" GPU With Surprising Results
-
Qualcomm and Samsung's 30-Year AI Alliance Enters a New Phase as On-Device AI Chip Race Heats Up
-
Multi-Token Prediction support coming to MLX-LM for Qwen 3.5
-
Multiverse Computing Targets On-Device AI With Compressed Models and New API Portal
-
Hugging Face Releases One-Liner for Automatic Hardware Detection and Model Selection
-
I Ran Local LLMs on a 'Dead' GPU, and the Results Surprised Me
-
OmniCoder-9B: Efficient Coding Model for 8GB GPUs
-
Nvidia's Nemotron 3 Super: Understanding the Significance for Local LLM Deployment
-
Startup Transforms Mac Mini Into Full-Powered AI Inference System With External GPU
-
Open-Source GreenBoost Driver Augments NVIDIA GPU VRAM With System RAM and NVMe Storage
-
I made Karpathy's Autoresearch work on CPU
-
Lemonade v10 Brings Linux NPU Support and Multi-Modal Capabilities
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Nvidia Pushes Jetson as Edge Hub for Open AI Models
-
Quantization Explained: Q4_K_M vs AWQ vs FP16 for Local LLMs
-
Cutile.jl Brings Nvidia CUDA Tile-Based Programming to Julia
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Qwen 3.5-35B Uncensored GGUF Models Now Available
-
HP Refreshes Lineup with AI-Focused Workstations
-
Llama.cpp Prompt Processing Optimization: Ubatch Size Configuration Guide
-
Final Qwen3.5 Unsloth GGUF Update with Improved Size/Quality Tradeoffs
-
MediaTek Advances Omni Model for Efficient Smartphone Inference
-
Apple Unveils MacBook Pro with M5 Pro and M5 Max Featuring On-Device AI
-
On-Device AI Laptop Lineups Become Standard Across Major Manufacturers
-
Running Local AI Models on Mac Studio 128GB: 4B, 20B & 120B Tested
-
Apple Neural Engine Reverse-Engineered for Local Model Training on Mac Mini M4
-
Qwen 3.5-35B-A3B Emerges as Efficient Daily Driver, Replacing 120B Models
-
Bare-Metal LLM Inference: UEFI Application Boots Directly Into LLM Chat
-
On-Device AI in Mobile Apps: What Should Run on the Phone vs the Cloud (A 2026 Decision Guide)
-
Qwen3.5-35B Unsloth Dynamic GGUFs Achieve SOTA Across Nearly All Quantisation Levels
-
Running LLMs on Raspberry Pi and Edge Devices: A Practical Guide
-
DeepSeek Releases DualPath: Addressing Storage Bandwidth Bottlenecks in Agentic Inference
-
Mirai Announces $10M to Advance On-Device AI Performance for Consumer Devices
-
South Korea to Launch $687 Million Project to Develop On-Device AI Semiconductors
-
Open-Source llama.cpp Finds Long-Term Home at Hugging Face
-
O-TITANS: Orthogonal LoRA Framework for Gemma 3 with Google TITANS Memory Architecture
-
Taalas Etches AI Models onto Transistors to Rocket Boost Inference
-
Mihup and Qualcomm Collaborate to Advance Secure On-Device Voice AI for BFSI
-
Enhanced Quantization Visualization Methods for Understanding LLM Compression Trade-offs
-
Hardware Economics Shift: DDR5 RDIMM Pricing Now Comparable to GPUs for Local Inference
-
GPT4All Replaces Ollama On Mac After Quick Trial
-
Same INT8 Model Shows 93% to 71% Accuracy Variance Across Snapdragon Chipsets
-
Qwen3-Next 80B MoE Achieves 39 Tokens/Second on RTX 5070/5060 Ti Dual-GPU Setup
-
Sourdine: Open-Source macOS App for 100% Local AI Transcription
-
Alibaba Unveils Major AI Model Upgrade Ahead of DeepSeek Release
-
Scaling llama.cpp On Neoverse N2: Solving Cross-NUMA Performance Issues
-
MiniMax-M2.5 230B MoE Model Released with GGUF Support for Local Deployment
-
Simile AI Raises $100M Series A for Local AI Infrastructure
-
Samsung's REAM: Alternative Model Compression Technique
-
Arm SME2 Technology Expands CPU Capabilities for On-Device AI
-
NAS System Achieves 18 tok/s with 80B LLM Using Only Integrated Graphics
-
Carmack Proposes Using Long Fiber Lines as L2 Cache for Streaming AI Data