Tagged "benchmarks"
-
Quantization-Aware Healing: 4-Bit Models Outperform Full-Precision Originals
-
8 Free Tools to Assess Your PC's Local AI Capabilities
-
Google COSMO Leak Reveals Gemini Nano and On-Device AI Skills
-
Ollama v0.33.0 Release Candidate Adds Claude Desktop Integration and Performance Improvements
-
Ollama v0.32.15: Time-to-First-Token Cut in Half with Metadata Caching
-
Qwen3.8-27B Matches Claude Opus 4.6 on Coding, Runs on Consumer GPUs
-
What If Local LLM Inference Is Using Consumer Hardware Wrong?
-
GGUF Quantization Deep Dive: Q4_K_M vs IQ4_XS vs IQ4_NL Performance
-
Qwen3.8-27B Surpasses 1 Million Downloads, Overseas Developers Race to Maximize Local Deployment
-
GGUF Quantization Compared: Q4_K_M vs. IQ4_XS vs. IQ4_NL Performance Analysis
-
AMD Adds Day 0 Qwen3.8 Support, Radeon AI PRO R9700 Hits 51.8 Tokens per Second
-
Google Pixel 11 Launches With Faster On-Device Gemini at $899 Starting Price
-
The Qwen MLX Challenge
-
HackerNoon Compares 7 Best Self-Hosted Inference Servers for Open-Source Models
-
Qwen 3.8 27B Successfully Runs on 16GB RAM Using LM Studio
-
Hugging Face State of Open Models: Summer 2026 Observations
-
7 Best Self-Hosted Inference Servers for Open-Source Models Compared (2026)
-
Running DeepSeek's 284B LLM on a Laptop: Quantisation and GGUF Optimization
-
Local Model Performance Benchmarks on MacBook Pro M5 Max: Real-World Inference Metrics
-
How to Run Local LLMs for Free on Slow Laptops: A Practical Guide
-
Benchmarking Local LLMs on Consumer Hardware: Real-World Performance Data
-
Meta Releases Muse Glimmer: 30B Open-Source LLM for Local Deployment
-
DeepSeek V4 Flash Achieves 82.7% on Terminal-Bench 2.1
-
Shrinking an AI Model 86% Doesn't Make It 86% Dumber: Compression Breakthroughs
-
Self-Hosted LLM Costs 2026: Comprehensive Pricing Comparison
-
Show HN: Benchmark Local LLMs Fit for Your Device Specs
-
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems
-
Seeed Studio's reCamera Pro Makes On-Device AI Faster and Easier
-
ASUS Vivobook S16 Arrives with 45 TOPS NPU and OLED Display
-
Homebench: Comprehensive Benchmarking Tool for Local LLMs
-
K-EXAONE 2.0 Brings 262K Context to Frontier AI
-
28.9M-Parameter LLM Runs Locally on ESP32-S3 at 9 Tokens/s
-
AMD's MI355X Undercuts Nvidia's B300 on Cost to Run China's Kimi K3
-
Q4 vs Q6 vs Q8: The Quantization Decision Framework for Local LLMs
-
AMD Ryzen AI PCs Demonstrate 18 Hours Weekly Productivity Gains in Project Management Tasks
-
Phi-4 Mini vs Gemma 3 vs Llama 3.2: 128K vs 32K Context Window Comparison
-
Testing Top Local LLMs Against ChatGPT and Claude Reveals Performance Gaps
-
Building a Dual V100 AI Workstation for Local LLMs
-
Open-Weights AI Models Have Become Good Enough
-
How Much Does a Local LLM Actually Cost to Run? Energy Costs Measured on Apple Silicon
-
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
-
Can a 2.8T Model Run on a Single Node of Nvidia B300 X8?
-
Sol-5.6 and Opus 5 Models Demonstrate Strong One-Shot Game Performance
-
Testing Local LLMs on Real Tasks: Honest Assessment of Practical Utility
-
K3 Model Achieves 20 Tokens/Second on 80x RTX 5090 Cluster
-
AMD Ryzen AI MAX+ 395 Discussed for Local AI Deployment
-
Round-Trip Correctness: New Metric for Generative AI Process Modeling
-
Nvidia Isn't the Only Choice for Local LLMs Anymore, and AMD Test Proves It
-
My Local LLM Struggles with Big Questions—Here's What It's Actually Good At
-
AI Model Release Forecasts from Prediction Markets
-
On-Device AI vs Cloud AI: Which One Should Power Your Next Phone?
-
Deterministic Arena: Testing and Comparing AI Agents Through Code Execution
-
Agentic Test Processes and LLM Benchmarks: Evaluating Local AI Agents
-
AI Inference Costs: Build vs. Rent
-
GPT-5.6 Sol vs. Claude Fable 5 in CNC Red Alert 2 Benchmark
-
'AI Code Is Insane Trash' – David Gerard on Code Generation Quality
-
My Local LLM Struggles With Big Questions—Here's What It's Actually Good At
-
Google Demonstrates New On-Device AI Features for Pixel 10
-
I Thought My Local AI Would Replace My Claude Subscription — Then I Tried Automating My PC
-
AMD Ryzen 7 7700X3D Linux Performance Review
-
Google Gemma 4 Debuts for Pixel 10 With Powerful On-Device AI Features
-
Mira Murati's Thinking Machines Launches Open-Weight AI Model
-
llama.cpp's 4.26× Intel Gain Has a Narrow Catch
-
How Much Does It Actually Cost to Run a Local LLM? (Euros per Million Tokens, Measured)
-
Show HN: Turn Meeting Recordings into Searchable Transcripts. All Local
-
Cost vs. Accuracy in CursorBench 3.1: The Effect of Family and Spend
-
The Triage Is the Product: Running AI Agents Against Ethereum's Protocol Code
-
Tencent Open-Sources Hy3 295B MoE Model Built for STEM Reasoning
-
Making AI Code Review Measurable
-
Viability of Local Models for Coding
-
Ollama Runs 32B Local AI Models on a $599 Mac via Quantization for Free
-
Ask HN: Which AI Model Do You Use for What?
-
RISC-V RVV Vector Benchmarks: SpacemiT K3 SoC Performance for Edge AI
-
Local LLM Performance Gap With Frontier Models Smaller Than Expected
-
3 Local LLM Workflows That Actually Save Me Time
-
How to Choose Between Small and Frontier Models
-
Tiny LLM Benchmark: Jetson Orin Nano Super 8GB
-
Hermes MoA Virtual Models: 8% Higher Than Opus 4.8, 11% Higher Than GPT 5.5
-
The Mac Mini is the Best On-Device AI Computer You Can Buy: Here's Why
-
Claude Opus 4.5 vs. GLM-5.2: Comparative Model Analysis
-
ORA: Smaller Models. Same Intelligence
-
NVIDIA DFlash Block Diffusion Accelerates Autoregressive LLM Inference
-
DeepSWE v1.1 – Updated Execution and Grading for Software Engineering Tasks
-
An Analysis on Why LLMs Perform Badly on Long Loop Tasks
-
GLM-5.2 Challenges Claude Opus in WebGL Game Build
-
Lessons from Building Evals for Financial AI Agents
-
DeepSWE Benchmark Updated with GLM 5.2 and Expanded Model Comparisons
-
The AI Definition of Done: Establishing Quality Standards Beyond Human Review
-
Gaming PC vs Phone Local LLM Deployment: Only One Remains in Daily Use
-
Intel Core Ultra X7 Panther Lake Performance Benchmarked on Linux
-
An End-to-End Machine Learning Pipeline on Time-Series Data
-
Brick: State-of-the-Art LLM Routing
-
South Korea Launches K-On-Device AI Chip Project With 511.1B Won Funding
-
Samsung's Exynos 2600 Doubles On-Device AI Performance in MLPerf Benchmarks
-
Why Tool Calling is More Important Than Model Size for Local LLMs
-
General-Purpose Large Language Models Outperform Specialized Clinical AI
-
Docfai.app Launches With Free Trial for Local Document Processing
-
It Is Beginning: AI Improves Itself
-
RTX 5080 and RTX 3090 Setup Achieves 80 Tok/s on Qwen 3.6 27B Q8
-
vLLM vs Ollama 2026: 793 vs 41 TPS Performance Benchmark
-
Show HN: Tail Panic – a multiplayer game designed for AI agents
-
AMD claims 256-core Zen 6 'Venice' CPU beats Nvidia Vera by 3.3x
-
DeepSeek V4 Performance Analysis: 1.6T Day 0 to Day 43 Scaling Trends
-
Show HN: Veritrooper – find what your AI gets wrong about your own docs
-
AI Memory Systems Show Critical Limitations: 95% Error Rate in Key Benchmarks
-
Running Local AI Models on Old Laptops Without GPU
-
LLM Memory Systems Benchmark: High Recall, Near-Zero Precision for Tested Systems
-
A Cinematic Landing-Page Hero for 80 Cents (GPT Image 2 and Veo 3.1)
-
Liquid AI Launches Edge-Focused LFM2.5 Model to Power On-Device AI Agents
-
Microsoft and Nvidia to Unveil First Windows PCs with Nvidia CPUs and AI Capabilities
-
Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
-
DeepSeek's Flagship V4 Pro Model Drops to 75% Lower Pricing, Increasing Competitive Pressure on Local Inference Economics
-
Show HN: I Built a Debugging Challenge for the AI Coding Age
-
vLLM vs Ollama 2026: Performance Benchmark Reveals 9x Throughput Gap
-
Qualcomm's AI-Device Strategy Reflects Growing Market Momentum in On-Device Intelligence
-
Redditor Successfully Runs 1 Trillion Parameter LLM Using Cheap Intel Optane DIMMs
-
New 8B Local LLM Design Marks Biggest Shift Since DeepSeek R1
-
M5 Max MacBook Runs Local Large Language Models Efficiently
-
A/B Tested Gemini 3.1 Pro vs. Claude Opus 4.6 – Usage Quota and Quality Comparison
-
110 Tokens/Second on RTX 4070 Super with Qwen 3.6 35B
-
Hardware LLM Taalas Reaches >14,000 TPS on Llama 3.1 8B
-
Benchmarking a Portable AI Workstation: Lenovo ThinkPad P16 Gen 3, Part 2
-
Google's Offline AI App Gets Three Major Feature Upgrades
-
Bito's AI Architect Improves Claude Opus Task Success Rate by 35%
-
HP's On-Device AI Needs More If It Is Going to Compete With Copilot
-
Google Limits Gemini Intelligence to New Flagships—Hardware Requirements for Local Deployment
-
AI/ML Benchmark Tool for Local LLM Inference and XGBoost Training
-
ROCm 7.2.3 Delivers Performance Improvements Over 7.0.0 on AMD Radeon AI PRO
-
Show HN: Find the best local LLM for your hardware, ranked by benchmarks
-
LLM temporal and causal reasoning research
-
Claude Opus 4.7 System Prompt Leaks Raise Local Deployment Questions
-
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
-
Researchers Report AI Breaking Every Benchmark for Autonomous Cyber Capability
-
Legacy System Analysis with AI Reveals Modern Architecture Under the Hood
-
Running a Local LLM on a 12-Year-Old Raspberry Pi
-
BT Explainer: Google's Gemma 4 Could Put Powerful AI on Your Phone and Laptop
-
Gemma 4 Replaces Entire Local LLM Stack for Many Practitioners
-
$200 NVIDIA V100 Server GPU Mod Beats RTX 3060 in Local LLM Test
-
One LM Studio Setting Makes Local LLMs Competitive With Cloud Models
-
Small On-Device AI Model Beats Claude Sonnet 4.5 and GPT-5
-
Anthropic Develops Tool to Detect When Claude Recognizes It's Being Tested
-
Local LLM Rewrites Resume Better Than ChatGPT, and It's Not Even Close
-
Building a Local LLM News Brief Taught Me the Real Problem Wasn't the Sources, It Was the Apps
-
Microsoft VibeVoice C++ Port Enables Local Voice AI on CPU and GPU Without Python
-
Improving Code Quality with Local Claude and Codex Models
-
A 49-Line Physics Classifier That Beats kNN on 76% of Benchmarks
-
Supercharging LLM Inference on Google TPUs: Achieving 3X Speedups With Diffusion-Style Speculative Decoding
-
Eval Skills for AI Agents
-
Control AI Risk with Pre-Built Frameworks and Ready-to-Run Evaluations
-
Running a Serious AI Model on a Consumer GPU Just Got Easier and That Matters More Than the Benchmark
-
Local AI Just Got Easier on Windows and the Implications Go Beyond the Benchmark
-
NIST's CAISI Evaluation of DeepSeek V4 Pro Finds It On Par with GPT-5
-
Study: AI Models That Consider User Feelings Are More Likely to Make Errors
-
AI Coding Tools Are Silently Disagreeing with Each Other
-
PFlash Claims 10x Prefill Speedup Over llama.cpp
-
Local LLMs Work Best When You're Not Loyal to Just One
-
Xmemory: Benchmarking Structured AI Memory Against RAG and Hybrid RAG
-
96.8% of MCP Tool Descriptions Don't Warn the Agent About Destructive Behaviour
-
Private LLM vs. ChatGPT: When It Makes Sense for Business
-
Running Capable Local LLMs Without Expensive GPU Hardware
-
IBM Introduces Granite 4.1 Family of Models for Local Deployment
-
Self-Hosted LLMs in Production: Real-World Limits and Practical Lessons
-
How Much "Brain Damage" Can an LLM Tolerate?
-
Estimating Black-Box LLM Parameter Counts via Factual Capacity
-
Intel N150 Mini PC Runs Local LLM for Home Assistant
-
Hipfire: A Rust-Native AMD Inference Engine That Outperforms llama.cpp
-
Linux Crushes Windows on llama.cpp Inference by Double Digits
-
The New Linux Kernel AI Bot Uncovering Bugs Is A Local LLM On Framework Desktop + AMD Ryzen AI Max
-
LLMs Consume 5.4x Less Mobile Energy Than Ad-Supported Web Search
-
Fixing Hallucination in LLM Prediction With Only One 48GB GPU
-
I Replaced My Local LLM With a Model Half Its Size and Got Better Results
-
I Cancelled Codex Two Months Ago. Opus 4.7 Brought Me Back
-
Show HN: We built an OCR server that can process 270 dense images/s on a 5090
-
Gemma 4 Just Replaced My Whole Local LLM Stack
-
Claude vs Local LLM: Real-World Prompt Comparison Reveals Trade-offs
-
Build a More Secure, Always-On Local AI Agent with OpenClaw and NVIDIA NemoClaw
-
115 TOPS in 0.67L: CHUWI AuBox X Packs On-Device AI Power Into a Palm-Sized Mini PC
-
Unweight: Lossless MLP Weight Compression for LLM Inference
-
We Built a Local Model Arena in 30 Minutes — Infrastructure Mattered More Than the App
-
Laimark – 8B LLM That Self-Improves on Consumer GPUs
-
Google's Gemma 4: The Most Practical Local LLM Despite Not Being The Smartest
-
LLM Personalization Breaks Down in High-Stakes Finance
-
Self-Hosted LLMs Transform Personal Knowledge Management Systems
-
MiniMax M2.7 GGUF Investigation Reveals NaN Issues Affecting 21-38% of Hugging Face Conversions
-
DFlash Doubles Token Generation Speed of Qwen3.5 27B on Mac M5 Max
-
OpenClaw at 250K GitHub Stars: Community Explores Practical Limitations Beyond News Digests
-
MiniMax M2.7 Achieves SOTA Performance Under 64GB on Mac with TQ Quantization
-
Running Same Prompts Through Claude and Local LLM Revealed Unexpected Results
-
Show HN: SkillCompass – Open-Source Quality Evaluator for Your AI Skills
-
Speculative Decoding Achieves 29% Speed Boost for Gemma-4 31B
-
MiniMax-M2.7 Delivers Exceptional Performance on Consumer Hardware
-
Self-Hosted LLM Elevates Personal Knowledge Management Systems to New Levels
-
Google Gemma 4 Delivers Exceptional Speed and Accuracy for Local Inference
-
Users Report Significant Performance Improvements After Migrating from Ollama to llama.cpp
-
Google's Gemini Nano 4 Offers Faster, Smarter Local Inference Capabilities
-
Intel Arc Pro B70 32GB Achieves 12 Tokens/Sec on Qwen 3.5-27B
-
Gemma 4 31B vs Qwen 3.5 27B: Comprehensive Long Context Benchmark
-
GLM 5.1 Dominates Agentic Benchmarks, Outperforming Most Models at 1/3 Opus Cost
-
Local Small LLMs Match Enterprise Model Performance on Vulnerability Detection
-
Warp Decode vs. vLLM's Triton Kernel: Performance Crossover Analysis
-
Qwen 3.5 122B Achieves 198 Tokens/sec on Dual RTX PRO 6000 Blackwell GPUs
-
Mano-P: Open-Source On-Device GUI Agent, #1 on OSWorld Benchmark
-
I Replaced My Local LLM With a Model Half Its Size and Got Better Results — and It Wasn't About the Parameters
-
Gemma 4 GGUF Models Updated with Critical Quantization Fixes
-
Comprehensive Benchmark: 37 LLMs Tested on MacBook Air M5 With Open-Source Tool
-
Gemma 4 Achieves Top Multilingual Performance Across European Languages
-
MemPalace, the Highest-Scoring AI Memory System Ever Benchmarked
-
Show HN: Willitrun – Check if Any ML Model Runs on Any Device (Benchmark-Backed)
-
Real-time Multimodal AI on Apple Silicon: Gemma E2B Demo Shows Practical Edge Deployment
-
Gemma 4 31B Achieves Exceptional Performance on Local Hardware
-
Quantization Strategy Comparison: Balancing Quality and Speed on Consumer Laptops
-
Qwen 3.6 Free Model Available via OpenRouter
-
Gemma 4 31B Achieves Third Place on FoodTruck Bench, Beating Larger Models
-
Unpaved: Audit Toolkit for AI Developer Tool Bias in Global South Contexts
-
Qwen 3.5 397B Reduced to 35% Parameters With Usable Quality on 96GB GPU
-
Apple Research Shows Self-Distillation Significantly Improves Local Code Generation
-
GPUs vs. TPUs: Decoding the Powerhouses of AI
-
Gemma 4 31B Outperforms GLM 5.1 in Real-World Testing
-
YC-Bench: GLM-5 Matches Claude Opus 4.6 at 11× Lower Cost
-
Gemma 4 26B A4B Outperforms Qwen 3.5 35B on Apple Silicon
-
Gemma 4 Makes Local AI Agents Practical
-
Gemma 4 Shows Strong Reasoning Performance with Thinking Tokens
-
Qwen 3.6-Plus Released
-
Bonsai 1-Bit Models Deliver Exceptional Local Inference Performance
-
Qwen 3.5-27B Demonstrates Superior Performance vs Gemini 3.1 Pro and GPT-5.3
-
ByteShape Releases Qwen 3.5 9B Quantisations with Hardware-Matched Tuning Guide
-
Llama.cpp Merging TurboQuant Lite (attn-rot) with Major Performance Gains
-
PrismML Announces 1-Bit Bonsai: First Commercially Viable 1-Bit LLMs
-
Does RAG Help AI Coding Tools?
-
M5 Max Delivers 1.7x Faster Inference Than M3 Max on Qwen 3.5 Models
-
Qwen3 512k Context via TurboQuant on Mac mini
-
Forensic Beats Mem0 with 90.1% on LOCOMO Benchmark
-
Reverse-Engineering the Apollo 11 Code with AI
-
Comparison of Two Frameworks: 40% Token Efficiency Improvement
-
TurboQuant Benchmarked in Llama.cpp: Google's Extreme Compression Research Tested in Practice
-
Qwen 3.5 27B Achieves 1.1M Tokens/Second on B200 GPUs with Optimized vLLM Config
-
Homelab Consolidation: Replacing 3 Models with Single 122B MoE Model on AMD Ryzen AI MAX+
-
Real-World Benchmark: DeepSeek-V3 Matches Claude Sonnet on Routine Coding Tasks
-
Llama.cpp Benchmark: RTX 5090 vs Enterprise Systems Compared
-
South Korea Science Ministry Seeks Five On-Device AI Pilot Projects for Public Services
-
KV Cache Quantization Levels Benchmarked on SWE-bench: Practical Trade-offs for Local Inference
-
FlashAttention-4 Delivers 2.7x Faster Inference with 1613 TFLOPs/s on Blackwell GPUs
-
Llama.cpp ROCm 7 vs Vulkan Performance Benchmarks on AMD Mi50
-
MiniMax M2.7 Model to Be Released as Open Weights
-
Llama 8B Matches 70B Performance on Multi-Hop QA Using Structured Prompting
-
ik_llama.cpp Fork Delivers 26x Faster Prompt Processing on Qwen 3.5 27B
-
Multi-Token Prediction support coming to MLX-LM for Qwen 3.5
-
DeepSeek R1 RTX 4090 vs Apple M3 Max: Benchmark & Performance Guide
-
Build a $1,500 AI Server with DeepSeek-R1 on RTX 4090
-
Qwen 3.5 397B emerges as top-performing local coding model
-
Apple M5 Max 128GB real-world performance benchmarks for local inference
-
NVIDIA Nemotron Cascade 2 30B Delivers 120B-Class Performance in Compact Form Factor
-
Hugging Face Releases One-Liner for Automatic Hardware Detection and Model Selection
-
My Dinner with AI
-
I Ran Local LLMs on a 'Dead' GPU, and the Results Surprised Me
-
Qwen 3.5 4B Outperforms Nvidia Nemotron 3 4B in Local Benchmarks
-
Two Local Models Prove Competitive Enough to Replace ChatGPT, Gemini, and Copilot
-
Achieving 2000 Tokens Per Second with QWEN 3.5 27B on RTX-5090
-
P-EAGLE: Faster LLM Inference with Parallel Speculative Decoding in vLLM
-
Runpod Report: Qwen Has Overtaken Meta's Llama As The Most-Deployed Self-Hosted LLM
-
Apple M5 Max 128GB Benchmark Results for Local LLM Inference
-
Comprehensive MoE Backend Benchmarks for Qwen3.5-397B: Real Numbers vs Hype
-
Qwen 3.5-35B Uncensored GGUF Models Now Available
-
Simple Layer Duplication Technique Achieves Top Open LLM Leaderboard Performance
-
M5 Max and M5 Ultra Chipsets Demonstrate Significant Bandwidth Improvements for Local LLM Inference
-
HP OMEN MAX 16 Review: Is Local AI on a Laptop Viable in 2026?
-
Fine-Tuned Qwen SLMs (0.6–8B) Demonstrate Competitive Performance Against Frontier LLMs on Specialized Tasks
-
FretBench – Testing 14 LLMs on Reading Guitar Tabs Reveals Performance Gaps
-
Strix Halo (Ryzen AI Max+ 395) Achieves Strong Local Inference Performance with ROCm 7.2
-
Qwen 3.5 Family Benchmark Comparison Shows Strong Performance Across Smaller Models
-
AI Agent Reliability Tracker
-
Qwen 3.5 27B Achieves Strong Local Inference Performance
-
Benchmark: Local Open-Source LLMs Competitive in Real-Time Trading Applications
-
Qwen3-Coder-Next Achieves Top Ranking on SWE-bench at Pass@5
-
Qwen 3.5-35B-A3B Achieves 37.8% on SWE-bench Verified Hard
-
Qwen 3.5-27B Q4 Quantization Comparison and Analysis
-
Qwen 3.5 vs Qwen 3 Benchmark Analysis: Generational Performance Improvements Visualized
-
Local LLM Performance Improvements: A Year of Progress Since DeepSeek R1 Moment
-
Running Local AI Models on Mac Studio 128GB: 4B, 20B & 120B Tested
-
Google Research Finds Longer Chain-of-Thought Correlates Negatively With Accuracy
-
The ML.energy Leaderboard
-
Accuracy vs. Speed in Local LLMs: Finding Your Sweet Spot
-
Qwen3.5-35B Unsloth Dynamic GGUFs Achieve SOTA Across Nearly All Quantisation Levels
-
Qwen3.5-35B RTX 5080 Experiments Confirm KV q8_0 as Free Lunch, Q4_K_M Remains Optimal
-
Qwen 3.5 MoE Delivers 100K Context Window at 40+ TPS on RTX 5060 Ti
-
Qwen 3.5 Underperforms on Hard Coding Tasks—APEX Benchmark Analysis
-
DeepSeek Paper – DualPath: Breaking the Bandwidth Bottleneck in LLM Inference
-
Qwen3.5 Series Releases Comprehensive Model Lineup Across All Tiers
-
Show HN: 100% LLM Accuracy–No Fine-Tuning, JSON Only
-
Advanced Quantization Techniques Show Surprising Performance Gains Over Standard Methods
-
The Real AI Competition Is Closed-Source vs Open-Source, Not America vs China
-
Anthropic Has Never Open-Sourced an LLM: Implications for Local Deployment Strategy
-
Which Web Frameworks Are Most Token-Efficient for AI Agents?
-
How Do You Know Which SKILL.md Is Good?
-
Qwen3-Code-Next Proves Practical for Local Development: Real-World Coding Tasks on Mac Studio
-
GLM-5 Becomes Top Open-Weights Model on Extended NYT Connections Benchmark
-
How Slow Local LLMs Are on My Framework 13 AMD Strix Point
-
Asus ExpertBook B3 G2 with 50 TOPS AI Sets New Enterprise Standard
-
I Run Local LLMs in One of the World's Priciest Energy Markets, and I Can Barely Tell
-
Strix Halo Performance Benchmarks: Minimax M2.5, Step 3.5 Flash, Qwen3 Coder
-
Qwen3 Coder Next Remains Effective at Aggressive Quantization Levels
-
SanityBoard Adds 27 New Model Evaluations Including Qwen 3.5 Plus, GLM 5, and Gemini 3.1 Pro
-
Qwen3 Coder Next 8FP Demonstrates Exceptional Long-Context Performance on 128GB System
-
The Path to Ubiquitous AI (17k tokens/sec)
-
Free ASIC-Accelerated Llama 3.1 8B Inference at 16,000 Tokens/Second
-
Enhanced Quantization Visualization Methods for Understanding LLM Compression Trade-offs
-
Hardware Economics Shift: DDR5 RDIMM Pricing Now Comparable to GPUs for Local Inference
-
GPT4All Replaces Ollama On Mac After Quick Trial
-
Real-World Coding Benchmark Tests LLMs on 65 Production Codebase Tasks
-
Alibaba's Qwen3.5-397B Achieves #3 Position in Open Weights Model Rankings
-
Qwen 3.5-397B-A17B Now Available for Local Inference with Aggressive Quantisation
-
MiniMax-M2.5 230B MoE Model Released with GGUF Support for Local Deployment
-
Optimal llama.cpp Settings Found for Qwen3 Coder Next Loop Issues
-
MiniMax M2.5: 230B Parameter MoE Model Coming to HuggingFace
-
Running Your Own AI Assistant for €19/Month: Complete Self-Hosting Guide
-
OpenClaw with vLLM Running for Free on AMD Developer Cloud
-
Running Mistral-7B on Intel NPU Achieves 12.6 Tokens/Second
-
New Header-Only C++ Benchmark Tool for Predictive Models on Raw Binary Streams
-
Use Recursive Language Models to address huge contexts for local LLM