Tagged "mlx"
- Ollama v0.33.0 Release Candidate Adds Claude Desktop Integration and Performance Improvements
- Ollama v0.32.15 Adds Model Metadata Cache to Reduce Per-Request Overhead
- The Qwen MLX Challenge
- Meta's Muse Glimmer Now Available Across All Platforms in Ollama
- Meta's Muse Glimmer Now Available Across All Platforms in Ollama
- Meta's Muse Glimmer Now Available Across All Platforms via Ollama
- MacPaw and Liquid AI: Complete On-Device AI Stack for macOS
- Muse Glimmer Now Available on Ollama – Meta's Open Multimodal Agent Model
- Ollama v0.32.6: Faster Apple GPU Inference with Speculative Decoding
- Oppo Reno16 Pro 5G Pairs On-Device AI With a 6,700mAh Battery for Creators
- Apple's Hardware Is Ready for On-Device AI and PrismML Just Delivered a Real Breakthrough
- Your Smartwatch Now Detects a Heart Irregularity in Milliseconds – Without Ever Touching the Cloud
- Ask HN: What are you using for LLM inference in production?
- Open-Weights AI Models Have Become Good Enough
- Odysseus - PewDiePie's Self-Hosted AI Finally Runs Fast on Mac
- Arm China Unveils "Tianxuan" CPU and Xingchen 300 Platform, Targeting Ubiquitous AIoT with On-Device AI Portfolio
- Sunday Reboot: Shrinking Models and an On-Device AI Future
- This Open-Source Extension Lets You Rewrite Your X Algorithm Using a Local LLM, and It Healed My Timeline
- Show HN: AITerm – a macOS Terminal with an AI Command Loop and a Safety Gate
- Google's LiteRT.js Enables On-Device AI Inference in Web Browsers
- Running Local AI on Mac With Home Assistant Integration
- Apple's MacBook Lineup Overhaul Features M7 Chip for Enhanced Local AI
- Ollama's New MLX Engine Delivers Significant Performance Gains on Mac
- Apple Updates Creator Studio with AI Video Editing, Image Generation, and Logic Pro Enhancements
- Asahi Linux 7.1 Progress Report
- You Can Now Run Max AI Models on Apple Silicon
- Liquid AI Ships LFM2.5-230M with Broad Framework Support for On-Device Inference
- Apple's M7 Chip Delivers 56% Memory Bandwidth Increase for On-Device AI
- Mac Mini Emerges as Top Choice for Local On-Device AI Deployment
- Mac Mini Positioned as Premier On-Device AI Computer for Local LLM Inference
- Apple unveils Core AI for on-device generative models
- Most People Use Ollama or llama.cpp for Local LLMs, but These Are the Tools I Switch to When It Gets Serious
- Show HN: 11 Model Families Ported to Apple's CoreAI On-Device Framework
- Qualcomm Launches Dragonwing MBM Silicon with Advanced On-Device AI Capabilities
- Apple Unveils AFM 3 Core Advanced with 20 Billion Parameters for On-Device AI
- Google AI Edge Gallery Launches on macOS With Offline Gemini Models
- Google Introduces Gemma 4 QAT for Ultra-Low Memory Local Inference
- Apple iPad Air with M4 Chip Drops to $1349; Powerful On-Device LLM Inference Now More Accessible
- Google's New Gemma 4 12B AI Model Is Built for Laptops
- Samsung's Exynos 2800 Brings HBM Memory to Mobile AI, Enabling Faster Local Model Inference
- Why AI Hardware Is a Chip Layer Problem
- M5 Max MacBook Runs Local Large Language Models Efficiently
- Chrome Is Quietly Downloading a 4GB AI Model Without Your Permission
- Samsung's Exynos 2800 Brings Significant On-Device AI Capabilities
- The Time Bomb Went Off: AI's All-You-Can-Eat Era Just Ended in Real Time
- Chrome Silently Downloads 4GB Gemini Nano Model Without User Consent
- Offline Voice-to-Text and AI Keyboard App for Local Processing
- Apple's M5 MacBook Air Advances On-Device AI with Redesigned Hardware
- Avocado Studio: Open-Source AI Content Editor for Next.js Sites
- Running AI Models Locally on M4 Processors with 24GB Memory
- Cotypist – AI Autocomplete for Mac
- Lython: Experimental Python Compiler Toolchain Based on LLVM
- Mlx-serve: Run LLMs Natively on Your Mac
- Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
- Google's Gemma 4 Brings Powerful AI Capabilities to Phones and Laptops
- NVIDIA Nemotron 3 Nano Omni Powers Multimodal Agent Reasoning in a Single Efficient Open Model
- I Replaced My Local LLM With a Model Half Its Size and Got Better Results
- Llama 4 Scout on MLX: The Complete Apple Silicon Guide (2026)
- DFlash Doubles Token Generation Speed of Qwen3.5 27B on Mac M5 Max
- Sovereign AI: Why the Next GPT Will Be Born in Our Living Rooms
- oMLX Framework Implements DFlash Attention for Optimized Inference
- DFlash Speculative Decoding Achieves 3.3x Speedup on Apple Silicon
- Comprehensive Benchmark: 37 LLMs Tested on MacBook Air M5 With Open-Source Tool
- Qwen 3.6 Free Model Available via OpenRouter
- Ollama Gets Blazing Fast on Macs with Full MLX Support and 2× Speedups
- Mixed Precision Quantization on MLX with TurboQuant Implementation
- Kokoro TTS Achieves 20× Realtime Speed on CPU-Only On-Device Inference
- Apple Silicon Macs Run Local AI Faster with Ollama's New MLX Support
- Is Anyone Working on an AI Operating System?
- Ollama Adopts Apple's MLX Framework for Faster Local AI on Mac
- Select the Right Hardware for Your Local LLM Deployment with This Online Guide
- M5 Max Delivers 1.7x Faster Inference Than M3 Max on Qwen 3.5 Models
- TurboQuant KV Cache Compression Achieves 22.8% Faster Decoding at 32K Context
- mlx-Code: Run Claude Code Locally with MLX-LM
- Apple Plans Slimmed-Down Gemini Models for Local iPhone AI Features
- Google TurboQuant: Extreme Compression for Local LLM Deployment
- Qualcomm and Samsung's 30-Year AI Alliance Enters a New Phase as On-Device AI Chip Race Heats Up
- Multi-Token Prediction support coming to MLX-LM for Qwen 3.5
- Qwen 3.5 Emerges as Top Performer for Local Deployment with Extensive Quantization Options
- Snapdragon 8 Elite Gen 5 Hands the Galaxy S26 the AI Upgrade We've Been Waiting For
- Kimi Introduces Attention Residuals: 1.25x Compute Performance at <2% Overhead
- Dictare – Open-source Voice Layer for AI Coding Agents (100% Local)
- AMD Declares 'AI on the PC Has Crossed an Important Line' – Agent Computers as Next Breakthrough
- LoKI – Local AI Assistant for Linux and WSL
- OpenClaw vs Eigent vs Claude Cowork: Comparing Open-Source AI Collaboration Platforms
- Startup Transforms Mac Mini Into Full-Powered AI Inference System With External GPU
- Local LLMs on Apple Silicon Mac 2026: M1 M2 M3 Guide
- SK Hynix Completes Qualification for LPDDR6 Memory Optimized for AI Inference
- Apple Launches MacBook Neo with A18 Pro Chip for Affordable Local AI Inference
- Real-World Qwen 3.5 9B Agent Performance on M1 Pro Validates Edge Deployment
- Apple Unveils MacBook Pro with M5 Pro and M5 Max Featuring On-Device AI
- Apple Unveils MacBook Pro With M5 Pro and M5 Max for On-Device AI
- Apple M4 iPad Air Targets AI Users with Double M1 Speed Performance
- Qualcomm Launches Snapdragon Wear Elite for On-Device AI on Wearables
- Running Local AI Models on Mac Studio 128GB: 4B, 20B & 120B Tested
- Apple Neural Engine Reverse-Engineered for Local Model Training on Mac Mini M4
- Mirai Announces $10M to Advance On-Device AI Performance for Consumer Devices
- How AI is Redefining Price and Performance in Modern Laptops
- Apple Accelerates U.S. Manufacturing with Mac Mini Production
- Future of Mobile AI: What On-Device Intelligence Means for App Developers
- Qwen3-Code-Next Proves Practical for Local Development: Real-World Coding Tasks on Mac Studio