Tagged "memory-optimization"
-
Llama.cpp Release b10485: GGML Sync with Platform-Specific Optimizations
-
Show HN: I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac
-
How an $8 ESP32 S3 Microcontroller Runs a 28.9M Parameter Local LLM
-
7 Best Self-Hosted Inference Servers for Open-Source Models Compared (2026)
-
Running DeepSeek's 284B LLM on a Laptop: Quantisation and GGUF Optimization
-
vLLM v0.27.0 Released with Major Kernel Improvements and New Model Support
-
vLLM v0.27.0rc2 Release Candidate Available
-
llama.cpp Improves CUDA Performance with Kernel Fusion
-
Chrome's On-Device AI Model Requires 20GB Storage Space
-
Llama.cpp Adds LRU Scheduler for Multi-Model Serving
-
Chrome and Edge Browsers Quietly Deploy Up to 20GB AI Models on Windows 11
-
Optimizing Qwen 3.6 for Local Development: A Developer's Guide
-
Shrinking an AI Model 86% Doesn't Make It 86% Dumber: Compression Breakthroughs
-
llama.cpp Build b10301: CUDA Optimization and Compiler Warning Fixes
-
Google Chrome Reveals Storage Requirements for Integrated Local AI Models
-
vLLM v0.27.0rc1: Latest Release Candidate for High-Performance Inference
-
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems
-
llama.cpp b10298: Multi-Token Multi-Dimension Chunk Serialization Support
-
LFM2.5-2.6B: On-Device Agentic Model With 128K Context and Tool Calling
-
Homebench: Comprehensive Benchmarking Tool for Local LLMs
-
LLM Memory Doesn't Only Get Written Wrong, It Goes Wrong Later
-
llama.cpp b10256 – SYCL SDPA Extended to Quantized KV Caches
-
llama.cpp Build b10258: Sampling Architecture Refinements
-
K-EXAONE 2.0 Brings 262K Context to Frontier AI
-
28.9M-Parameter LLM Runs Locally on ESP32-S3 at 9 Tokens/s
-
Kioxia Is Coming for Samsung and SK Hynix With UFS 5.0 and PCIe 6.0 AI NAND
-
The KV Cache Survival Guide: Why Your GPU Runs Out of Memory with Local LLMs
-
The KV Cache Survival Guide: Why Your GPU Runs Out of Memory with Local LLMs
-
Tim Cook Called Apple's On-Device AI a 'Competitive Weapon' in Final Earnings Call as CEO
-
GPU Half-Idle: The Hundred-Billion-Dollar Race to Squeeze 10x Efficiency from Silicon
-
Rent the Intelligence. Own the Memory
-
Legal and Compliance Considerations for AI Memory Systems in Local Deployments
-
faster-enhancer.c: C Library for Stable Real-Time On-Device Denoising
-
Running Local LLMs on Raspberry Pi: Exploring Edge Inference Boundaries
-
Deploying 1-Bit Bonsai-27B with PrismML and llama.cpp for Local Inference
-
Ruff v0.16.0: 413 Default Rules for Code Quality in AI Development
-
Claude Code Cut System Prompt by 80%: Implications for Small Local Models
-
Show HN: TS Compiler Knowledge Graph Reducing AI Tokens About 90%
-
SK hynix 3D-Stacked DRAM-on-Logic Architecture Could Solve On-Device AI Memory Constraints
-
How To Build Your Own LLM Runtime From Scratch
-
Full Offline Voice Agent Running in 1.2 GB RAM on Android with FunctionGemma
-
Jan: Open, Cross-Platform AI App with Useful Proprietary Models
-
Mira Murati's Thinking Machines Launches Open-Weight AI Model
-
Study: Cerebellum Helps AI Ignore the Ordinary for More Efficient Computing
-
Exploiting Sparsity for Long Context Inference: Million Token on Commodity GPUs
-
Show HN: Trace – Open-source, Self-organizing Memory for LLM Agents
-
Critical GPU Memory Leak Vulnerability Discovered in vLLM
-
Edge AI Transformation Coming to Creative Production Workflows
-
Google Rolls Out Android 17 and Gemma 4 with Advanced On-Device AI
-
Meet EverOS: An Open Source Markdown-First Agent Memory Runtime With Hybrid BM25 + Vector Retrieval
-
Samsung Presents UFS 5.0 Storage Targeted at On-Device AI Performance
-
Show HN: Brain.md – A Persistent Memory Layer for Your Coding Agents
-
TriAttention Solves KV Cache Memory Bottleneck in Local LLM Inference
-
NVIDIA DFlash Block Diffusion Accelerates Autoregressive LLM Inference
-
An Analysis on Why LLMs Perform Badly on Long Loop Tasks
-
Giving AI Human-Like Memory Limits (3–7 Words) Could Improve Language Learning
-
Samsung's UFS 5.0 Addresses Critical Memory Bandwidth Bottleneck in Mobile AI Inference
-
Samsung Unveils UFS 5.0 Storage Optimized for On-Device AI Applications
-
Form Before Data: Addressing the Real Bottleneck in Physical AI Systems
-
Agentic Systems Course: Learn to Build AI Agents with Live AI Coding
-
FlashRT: Execution State for Latency-First AI
-
Google's DiffusionGemma Brings Novel Text Generation to Local LLMs
-
CacheWise Optimizes KVCache Reuse for LLM Coding Agents
-
Ask HN: What Problem Did AI Create at Your Company That Didn't Exist Before?
-
What is Ollama? Introduction to the AI Model Management Tool
-
Open-Source Tool Adds Persistent Memory to Local LLM Deployments
-
Apple Unveils AFM 3 Core Advanced with 20 Billion Parameters for On-Device AI
-
TokenTamer: A Proxy That Reduces LLM Token Usage Through Context Compression
-
Developer Switches from LM Studio to llama.cpp, Citing Performance and Simplicity
-
Google Releases Gemma 4 QAT Models with Reduced Memory Requirements for Mobile and Laptop Deployment
-
Google Introduces Gemma 4 QAT for Ultra-Low Memory Local Inference
-
Best Local LLM Setup for RTX 5090: llama.cpp Fork with TurboQuant
-
AI Memory Systems Show Critical Limitations: 95% Error Rate in Key Benchmarks
-
Google Releases Gemma 4 QAT Models for Local AI Deployment
-
Running Infinite Context Lengths on 8GB GPU Without Out Of Memory
-
Maybe Coding Agents Don't Need a Bigger Memory. Maybe They Need Continuity
-
Sawtooth – An Async, Multi-Tiered Memory Framework for LLM Agents
-
Show HN: LLM Memory Without Context Bleed – 100% Precision vs. <10% Vector Search
-
Longsys Redefines On-Device AI with Groundbreaking Edge Memory Solutions
-
LLM Memory Systems Benchmark: High Recall, Near-Zero Precision for Tested Systems
-
Reducing GPU Costs for AI Inference: FP8, FP4, and vLLM Optimization Techniques
-
Tether AI Upgrades QVAC SDK With TurboQuant for Data Center-Sized Memory on Everyday Devices
-
Meet Memory OS: A 6-Layer Open-Source Memory Stack Built on Hermes Agent
-
Liquid AI Launches Edge-Focused LFM2.5 Model to Power On-Device AI Agents
-
Local LLM Setup: How to Use RAG and an Embedding Model to Stop Wasting Context
-
Samsung's Exynos 2800 Brings HBM Memory to Mobile AI, Enabling Faster Local Model Inference
-
Anker Soundcore Liberty 5 Pro Earbuds Feature Dedicated On-Device AI Chip with Touch Screen
-
Redditor Successfully Runs 1 Trillion Parameter LLM Using Cheap Intel Optane DIMMs
-
New 8B Local LLM Design Marks Biggest Shift Since DeepSeek R1
-
llama.cpp MTP Leak Fix Stabilizes Local AI Agents
-
AMD's New Ryzen AI Max Pro 400 with 192GB LPDDR5X Memory
-
Samsung's Exynos 2800 Brings Significant On-Device AI Capabilities
-
The Time Bomb Went Off: AI's All-You-Can-Eat Era Just Ended in Real Time
-
DwarfStar 4: Native Inference Engine Optimized for DeepSeek V4 Flash
-
SynapseKit: A New Production Framework for Deploying LLMs
-
Arm and Google Collaborate on On-Device AI Optimization Techniques
-
Claude Opus 4.7 System Prompt Leaks Raise Local Deployment Questions
-
Running Local AI LLMs on Mini PCs Without NVIDIA GPUs
-
Local LLM Persistent Context Prevents Repetitive Mistakes
-
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
-
Running a Local LLM on a 12-Year-Old Raspberry Pi
-
AMD's vLLM-ATOM Plugin Supercharges DeepSeek-R1 and Kimi-K2 Inference on MI350/MI400
-
Microsoft Researchers Find AI Models and Agents Can't Handle Long-Running Tasks
-
Ollama Out-of-Bounds Read Vulnerability Allows Remote Process Memory Leak
-
Deploying Frigate & Ollama On A Minisforum MS-A2 Server
-
Lemonade Gives AMD Startups a Wider Path to Local Inference
-
Critical Ollama Memory Leak Vulnerability Exposes 300,000 Servers Globally
-
0ctx – Local-First Project Memory for AI Workflows
-
Show HN: A Local-First Agentic Knowledge Manager
-
Critical Ollama Memory Leak Vulnerability Exposes 300,000 Servers Globally
-
Critical Ollama Memory Leak Vulnerability Exposes 300,000 Servers Globally
-
Agentic AI Community Focus: Building Local Agents in 2026
-
Show HN: Memex, Claude Memory via Local RAG with MCP and Offline Embeddings
-
Gemma 4 Just Replaced My Whole Local LLM Stack
-
Running a Serious AI Model on a Consumer GPU Just Got Easier and That Matters More Than the Benchmark
-
Xmemory: Benchmarking Structured AI Memory Against RAG and Hybrid RAG
-
Building a Raspberry Pi-Based Local LLM Server for Remote Access
-
GraphOS: Visual Runtime and Debugger for AI Agents with Local-First Execution
-
Local AI Isn't Just Ollama—Here's the Ecosystem That Actually Makes It Useful
-
Unsloth's Custom Kernels Make LLM Fine-Tuning Viable on Consumer GPUs
-
Can IBM's RITS Platform and vLLM Reset the Bar for Enterprise AI Access?
-
Elastic KV Cache Memory Breakthrough Enables Efficient Bursty LLM Serving and GPU Sharing
-
Rust Open-Source Headless Browser for AI Agents and Web Scraping
-
Google's Gemma 4 Brings Powerful On-Device AI to Phones and Laptops
-
Show HN: A Karpathy-Style LLM Wiki Your Agents Maintain
-
I Replaced My Local LLM With a Model Half Its Size and Got Better Results
-
10GB VRAM Local LLM: The Complete Setup Guide (2026)
-
Externalization in LLM Agents: Unified Review of Memory and Harness Engineering
-
Llama.cpp's Auto Fit Feature Quietly Reshapes Local AI Inference on Consumer Hardware
-
Bun v1.3.13
-
Unweight: Lossless MLP Weight Compression for LLM Inference
-
Sorting 1M u64 KV-Pairs in 20ms on i9-13980HX Using Branchless Rust Implementation
-
Google's Gemma 4: The Most Practical Local LLM Despite Not Being The Smartest
-
GBrain – System to Make Your AI Agent Better Reflect You
-
SigMap – Shrink AI Coding Context 97% with Auto-Scaling Token Budget
-
Dynamic Expert Cache in llama.cpp Achieves 27% Faster Inference on Large MoE Models
-
MiniMax M2.7 Achieves SOTA Performance Under 64GB on Mac with TQ Quantization
-
Researchers Achieve 1-Bit Quantization of OLMo-3 7B Using Distillation
-
Universal Knowledge Store and Grounding Layer for AI Reasoning Engines
-
A Deep Dive into Tinygrad AI Compiler
-
Self-Hosted LLMs Transform Personal Knowledge Management Systems
-
DMax: New Parallel Decoding Paradigm for Diffusion Language Models
-
LLM Wiki v2: Extended Knowledge Base for LLM Practitioners
-
Building Offline AI Companions on Severely Constrained Hardware (8GB RAM)
-
Intel Releases OpenVINO 2026.1 With Backend For Llama.cpp, New Hardware Support
-
Gemma 4 Support Stabilized in Llama.cpp
-
Gemma 4 GGUF Models Updated with Critical Quantization Fixes
-
MemPalace, the Highest-Scoring AI Memory System Ever Benchmarked
-
Octopoda: Open Source Memory Layer for Fully Offline AI Agents
-
CricketBrain: Neuromorphic Signal Processor in Rust (0.175us/step, 944 bytes)
-
TurboQuant in Llama.cpp Achieves 6X Smaller KV Cache
-
Context Window Optimization: Extending Gemma 4 Context Length Through Efficient Projection Quantization
-
GPU Memory for LLM Inference (Part 1)
-
Vektor – Local-First Associative Memory for AI Agents
-
Gemma 4 26B MoE Emerges as Optimal All-Around Local Model for Consumer Hardware
-
DGX Spark Hardware Limitations: Missing NVFP4 Support Undermines Local AI Value Proposition
-
Mixed Precision Quantization on MLX with TurboQuant Implementation
-
Free AI Video Clipper Using Scene and Speech-Based Segmentation
-
Gemma 4 KV Cache Memory Issues Fixed in llama.cpp
-
VRAM Optimization Technique Cuts Gemma 4 Memory Usage by 3x
-
OpenUMA – Apple-Style Unified Memory for x86 AI Inference
-
TurboQuant Enables Qwen 3.5-27B on 16GB Consumer GPUs
-
SmolLM2-360M Running on Samsung Galaxy Watch 4 with 74% Memory Reduction
-
Show HN: Memsearch – Persistent, Cross-Agent, Cross-Session Memory for AI Agents
-
Llama.cpp Merging TurboQuant Lite (attn-rot) with Major Performance Gains
-
Claw64 – Full Agentic Loop in <4KB on Commodore 64
-
PrismML Announces 1-Bit Bonsai: First Commercially Viable 1-Bit LLMs
-
Ollama Adopts Apple's MLX Framework for Faster Local AI on Mac
-
DeepSeek V3 Complete Guide: Deploy and Optimize Local AI in 2026
-
Google's TurboQuant Shows Memory Constraints Remain Critical for Local LLM Inference
-
Mixed KV Cache Quantization: Performance Risks and Pitfalls
-
TurboQuant KV Cache Compression Achieves 22.8% Faster Decoding at 32K Context
-
Forensic Beats Mem0 with 90.1% on LOCOMO Benchmark
-
Coding Implementation to Run Qwen3.5 Reasoning Models Distilled With Claude-Style Thinking Using GGUF and 4-Bit Quantization
-
TurboQuant Benchmarked in Llama.cpp: Google's Extreme Compression Research Tested in Practice
-
Qwen 3.5 27B Achieves 1.1M Tokens/Second on B200 GPUs with Optimized vLLM Config
-
Google's TurboQuant: The Unsexy AI Breakthrough Worth Watching
-
Running an Open-Weight LLM Locally on an Apple Watch
-
Ultra-Large 400B-Class LLM Runs on iPhone in Test
-
KV Cache Quantization Levels Benchmarked on SWE-bench: Practical Trade-offs for Local Inference
-
FOMOE: Running 397B Parameter Qwen3.5 MoE at 5-9 tok/s on $2,100 Desktop Hardware
-
Ditching Paid AI Services: Building Self-Hosted LLM Solutions as ChatGPT, Claude, and Gemini Alternatives
-
Qwen 3.5 122B Uncensored (Aggressive) Released with New K_P Quantisations
-
Llama 8B Matches 70B Performance on Multi-Hop QA Using Structured Prompting
-
A Little Gap That Will Ensure the Future of AI Agents Being Autonomous
-
DeepSeek R1 RTX 4090 vs Apple M3 Max: Benchmark & Performance Guide
-
Running an AI Agent on a 448KB RAM Microcontroller
-
MacinAI Local brings functional LLM inference to classic Macintosh hardware
-
NVIDIA Nemotron Cascade 2 30B Delivers 120B-Class Performance in Compact Form Factor
-
Community Converges on Optimal KV Cache Quantization Strategies for Qwen 3.5 Models
-
LMCache Dramatically Accelerates LLM Inference on Oracle Data Science Platform
-
Mamba 3: State Space Model Architecture Optimized for Inference
-
Custom GPU Multiplexer Achieves 0.3ms Model Switching on Legacy Hardware
-
Run LLMs Locally with Llama.cpp
-
Mistral Small 4 119B Released with NVFP4 Quantisation Support
-
Researcher Discovers Universal "Danger Zone" in Transformer Model Architecture at 50% Depth
-
The Moment AI Agents Stopped Being a Feature and Started Becoming a System
-
OpenClaw Isn't the Only Raspberry Pi AI Tool—Here Are 4 Others You Can Try This Week
-
OmniCoder-9B: Efficient Coding Model for 8GB GPUs
-
Open-Source GreenBoost Driver Augments NVIDIA GPU VRAM With System RAM and NVMe Storage
-
Best Local LLM Models 2026: Developer Comparison
-
Memory Should Decay: Implementing Temporal Memory Decay in Local LLM Systems
-
3-Path Agent Memory: 8 KB Recurrent State vs. 156 MB KV Cache at 10K Tokens
-
Intel Updates LLM-Scaler-vLLM With Support For More Qwen3/3.5 Models
-
Qwodel – An Open-Source Unified Pipeline for LLM Quantization
-
Apple M5 Max 128GB Benchmark Results for Local LLM Inference
-
Quantization Explained: Q4_K_M vs AWQ vs FP16 for Local LLMs
-
SK Hynix Completes Qualification for LPDDR6 Memory Optimized for AI Inference
-
Experiment: 0.8B Model Self-Improvement on MacBook Air Yields Surprising Results
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
8 Local LLM Settings Most People Never Touch That Fixed My Worst AI Problems
-
Mnemos: Persistent Memory System for Local AI Agents
-
FreeBSD 14.4 Released: Implications for Local LLM Deployment
-
HP OMEN MAX 16 Review: Is Local AI on a Laptop Viable in 2026?
-
SK Hynix Develops 1c LPDDR6 DRAM to Boost On-Device AI Performance in Mobile Devices
-
Qwen 3.5 Derestricted Model Available for Local Deployment
-
Engram – Open-Source Persistent Memory for AI Agents
-
How to Run Your Own Local LLM — 2026 Edition
-
Llama.cpp Prompt Processing Optimization: Ubatch Size Configuration Guide
-
Show HN: Asterode – Multi-Model AI App with Memory and Power Features
-
Mojo: Creating a Programming Language for an AI World with Chris Lattner
-
OPPO and MediaTek Highlight On-Device AI Innovations at MWC 2026
-
Final Qwen3.5 Unsloth GGUF Update with Improved Size/Quality Tradeoffs
-
The Emerging Role of SRAM-Centric Chips in AI Inference
-
Critical: Qwen 3.5 Requires BF16 KV Cache, Not FP16 for Accurate Inference
-
Nummi – AI Companion with Memory and Daily Guidance
-
How to Run High-Performance LLMs Locally on the Arduino UNO Q
-
Qwen3.5-35B Successfully Runs on Raspberry Pi 5 at 3+ Tokens/Second
-
Unsloth Dynamic 2.0 GGUFs
-
LLmFit: Terminal Tool for Right-Sizing LLM Models to Your Hardware
-
Qwen3.5-35B RTX 5080 Experiments Confirm KV q8_0 as Free Lunch, Q4_K_M Remains Optimal
-
Krasis: Hybrid CPU/GPU MoE Runtime Achieves 3,324 Tokens/Second Prefill on RTX 5080
-
Running LLMs on Raspberry Pi and Edge Devices: A Practical Guide
-
Researchers Develop Persistent Memory System for Local LLMs—No RAG Required
-
Show HN: Pluckr – LLM-Powered HTML Scraper That Caches Selectors and Auto-Heals
-
Advanced Quantization Techniques Show Surprising Performance Gains Over Standard Methods
-
What Breaks When AI Agent Frameworks Are Forced Into <1MB RAM and Sub-ms Startup
-
Which Web Frameworks Are Most Token-Efficient for AI Agents?
-
Qwen3's Voice Embeddings Enable Local Voice Cloning and Mathematical Voice Manipulation
-
Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference
-
The Complete Stack for Local Autonomous Agents: From GGML to Orchestration
-
Breaking the Speed Limit: Strategies for 17k Tokens/Sec Local Inference
-
O-TITANS: Orthogonal LoRA Framework for Gemma 3 with Google TITANS Memory Architecture
-
Qwen3 Coder Next 8FP Demonstrates Exceptional Long-Context Performance on 128GB System
-
Enhanced Quantization Visualization Methods for Understanding LLM Compression Trade-offs
-
Running Local LLMs and VLMs on Arduino UNO Q with yzma
-
LayerScale Launches Inference Engine Faster Than vLLM, SGLang, and TRT-LLM
-
Local Vision-Language Models for Document OCR and PII Detection in Privacy-Critical Workflows
-
InitRunner: YAML-Based AI Agent Framework with RAG and Memory
-
Alibaba Unveils Major AI Model Upgrade Ahead of DeepSeek Release
-
Scaling llama.cpp On Neoverse N2: Solving Cross-NUMA Performance Issues
-
NVIDIA's Dynamic Memory Sparsification Cuts LLM Inference Costs by 8x
-
GPT-OSS 120B Uncensored Model Released in Native MXFP4 Precision
-
MiniMax Releases M2.5 Model with SOTA Coding and Agent Capabilities
-
Context Management Identified as Real Bottleneck in AI-Assisted Coding
-
Ring-1T-2.5 Released with SOTA Deep Thinking Performance
-
MiniMax M2.5: 230B Parameter MoE Model Coming to HuggingFace
-
Ming-flash-omni-2.0: 100B MoE Omni-Modal Model Released
-
Running Your Own AI Assistant for €19/Month: Complete Self-Hosting Guide
-
Heaps Do Lie: Debugging a Memory Leak in vLLM
-
Mistral AI Debugs Critical Memory Leak in vLLM Inference Engine
-
Energy-Based Models Compared Against Frontier AI for Sudoku Solving
-
Carmack Proposes Using Long Fiber Lines as L2 Cache for Streaming AI Data
-
Developer Switches from Ollama and LM Studio to llama.cpp for Better Performance