Tagged "fine-tuning"
-
Leveraging Local Small Language Models for Project-Specific Deployment
-
Strong Domain Adaptation Results with Qwen 3 4B Fine-Tuning
-
Teaching a Local LLM to Reason About a New Domain Through Continued Pretraining
-
NVIDIA Magpie TTS – Open-Weights Multilingual Voice Agents with Full Deployment Control
-
Meta's Muse Glimmer – Local, Agentic, Multimodal, and Open Source
-
TutorMoments: Research on When AI Should Intervene in Learning
-
Liquid AI Releases LFM2.5-2.6B: Powerful Agentic Model for Raspberry Pi and Edge Devices
-
Shrinking an AI Model 86% Doesn't Make It 86% Dumber: Compression Breakthroughs
-
LFM2.5-2.6B: On-Device Agentic Model With 128K Context and Tool Calling
-
Bubo: AI Code-Reviewer That Learns From Review Comments
-
Reinforcement Learning Fine-tuning Improves Local LLM Output Quality
-
NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework
-
NVIDIA AI Releases Molt: A PyTorch-Native Agentic Reinforcement Learning Framework
-
GPU Half-Idle: The Hundred-Billion-Dollar Race to Squeeze 10x Efficiency from Silicon
-
Testing Top Local LLMs Against ChatGPT and Claude Reveals Performance Gaps
-
Building a Dual V100 AI Workstation for Local LLMs
-
Sol-5.6 and Opus 5 Models Demonstrate Strong One-Shot Game Performance
-
NVIDIA Releases Molt: Agentic RL Training Framework Scaling to Trillion-Parameter Models
-
Brief notes on the OpenAI/Hugging Face incident
-
China State Media Says Support for Open AI Models Has Limits
-
Apertus 1.5: Swiss Open-Weight, Open-Source LLM Released
-
Don't Buy an Uncensored AI on a Flash Drive: What You Can Do Instead
-
A New Way of Debugging Open-Weight Models - IBM
-
MSI Pro Max Edge AI+ Mini PC Runs 120B Local AI Models With 128GB RAM
-
Shanghai Droi Technology Launches DroiClaw AI Operating System with Hybrid Edge-Cloud Architecture
-
My Local LLM Struggles with Big Questions—Here's What It's Actually Good At
-
Codeberg Updates Terms of Use to Prohibit LLM Model Training Extrusions
-
Qwen 3.8 with 2.4T Parameters Going Open-Weight Soon
-
'AI Code Is Insane Trash' – David Gerard on Code Generation Quality
-
My Local LLM Struggles With Big Questions—Here's What It's Actually Good At
-
Show HN: Senbonzakura – Remove Safety Guardrails from Open AI Models
-
Nvidia Showcases Nemotron Models for Japanese AI Development
-
AI-Assisted Development Exhaustion Highlights Need for Better Local Tooling
-
Mira Murati's Thinking Machines Launches Open-Weight AI Model
-
Rapid Rise of Open Source Models in the U.S.: Nvidia Nemotron Ultra Grows Quickly on Ollama
-
Building an AI Strength Coach: Local LLM Application with Research-Backed Training
-
WSL Transforms Windows Into a Viable Local LLM Development Platform
-
Record and Replay: Teach AI Agents Desktop Workflows by Showing Them Once
-
Building a Local LLM-as-Judge Pipeline for Image Dataset Curation
-
Tencent Open-Sources Hy3 295B MoE Model Built for STEM Reasoning
-
Open Source 1B LLM Trained from Scratch for $315 with Weights and Data Released
-
GLM-5.2's Code Reviews Are Only as Good as Your Prompt
-
A Guide on How to Run Nemotron 3 Super 120B Thinking on 2 Nvidia DGX Spark
-
Claude Opus 4.5 vs. GLM-5.2: Comparative Model Analysis
-
DeepSWE v1.1 – Updated Execution and Grading for Software Engineering Tasks
-
Why Small Local AI Models Get More Use Than Claude or Gemini
-
An Analysis on Why LLMs Perform Badly on Long Loop Tasks
-
Giving AI Human-Like Memory Limits (3–7 Words) Could Improve Language Learning
-
Mac Mini Positioned as Premier On-Device AI Computer for Local LLM Inference
-
Samsung's UFS 5.0 Addresses Critical Memory Bandwidth Bottleneck in Mobile AI Inference
-
Lessons from Building Evals for Financial AI Agents
-
Form Before Data: Addressing the Real Bottleneck in Physical AI Systems
-
DeepSWE Benchmark Updated with GLM 5.2 and Expanded Model Comparisons
-
The AI Definition of Done: Establishing Quality Standards Beyond Human Review
-
Agentic Systems Course: Learn to Build AI Agents with Live AI Coding
-
Repo-Slopscore: Detecting AI Contributions in Git Repositories via Commit Analysis
-
General-Purpose Large Language Models Outperform Specialized Clinical AI
-
It Is Beginning: AI Improves Itself
-
DeepSeek V4 Performance Analysis: 1.6T Day 0 to Day 43 Scaling Trends
-
Show HN: Veritrooper – find what your AI gets wrong about your own docs
-
I Replaced Cloud LLMs with Local Models Running Off a Proxmox LXC, and the Performance Trade-Off Was Worth It
-
SourceHut Disrupted by LLM Training Crawlers: Infrastructure and Data Concerns
-
Train Your Own LLM? Here's What Happens
-
Fine-tuning an LLM to Write Docs Like It's 1995
-
Nvidia Enters Windows Laptop Market, Taking on Intel and AMD
-
CNN sues Perplexity over alleged AI copyright theft
-
AI Guardrails Stripped From Meta and Google Models in Minutes
-
Show HN: An Open-Source Interactive AI Engineering Syllabus (1,100 Papers)
-
From Source Code to LLM Constraints: A Semantic Extractor for Python, SwiftUI, Lua
-
User Migration from LM Studio/Ollama to llama.cpp Shows Growing Preference
-
AMD's New Ryzen AI Max Pro 400 with 192GB LPDDR5X Memory
-
On-Device AI to Be in 80% of Wearables by 2032
-
Local LLMs Offer Unique Advantages That Cloud AI Services Cannot Match
-
The AI Layoff Receipts: Market Consolidation Accelerates Open-Source Model Adoption
-
Safety Paradox: How RLHF Creates the AI Psychosis Problem It's Meant to Prevent
-
MegaTrain: Full Precision Training of 100B+ Parameter LLMs on a Single GPU
-
How to Train Your GPT: Comprehensive Commented Training Guide
-
Claude Opus 4.7 System Prompt Leaks Raise Local Deployment Questions
-
Avocado Studio: Open-Source AI Content Editor for Next.js Sites
-
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
-
Legacy System Analysis with AI Reveals Modern Architecture Under the Hood
-
Discussion: Including New Mathematical Proofs in LLM Training Data for Rediscovery
-
Local LLM Rewrites Resume Better Than ChatGPT, and It's Not Even Close
-
I Replaced ChatGPT and Claude With This Powerful Local LLM and Saved Over $20 a Month While Gaining Full Control
-
NHS to Close-Source GitHub Repos Over AI and Security Concerns
-
Study: AI Models That Consider User Feelings Are More Likely to Make Errors
-
AI Coding Tools Are Silently Disagreeing with Each Other
-
IBM Introduces Granite 4.1 Family of Models for Local Deployment
-
Local AI Isn't Just Ollama—Here's the Ecosystem That Actually Makes It Useful
-
Unsloth's Custom Kernels Make LLM Fine-Tuning Viable on Consumer GPUs
-
Build Your Own Local AI Stack with 5 Docker Containers and Eliminate ChatGPT Subscriptions
-
Fixing Hallucination in LLM Prediction With Only One 48GB GPU
-
Using a Local LLM as a Zero-Shot Classifier
-
Mathesar 0.10.0
-
I Replaced My Local LLM With a Model Half Its Size and Got Better Results
-
AI Licensing Marketplaces: A Guide for Publishers and Content Creators
-
BibCrit – LLM Grounded in ETCBC Corpus Data for Biblical Textual Criticism
-
Laimark – 8B LLM That Self-Improves on Consumer GPUs
-
When Should AI Step Aside?: Teaching Agents When Humans Want to Intervene
-
The Case for Out-of-Process Enforcement for AI Agents
-
LLM Personalization Breaks Down in High-Stakes Finance
-
GBrain – System to Make Your AI Agent Better Reflect You
-
Fine-Tuned Qwen3.5-0.8B for OCR Outperforms Previous 2B Release
-
Minisforum N5 MAX AI NAS Delivers 126 TOPS with 200TB Storage for Local LLM Workloads
-
Developer Shares Golden Stack for Local Coding Assistant Integration Directly Inside Code Editors
-
Abliterated Local LLM Models Show Distinct Behavioral Characteristics Compared to Standard Variants
-
MiniMax-M2.7 Delivers Exceptional Performance on Consumer Hardware
-
MiniMax M2.7 Open-Sources Globally as Industry's First Self-Improving Model
-
MiniMax M2.7 Is Now Open Source
-
MiniMax M2.7 Released: New Model Available for Local Deployment
-
Self-Hosted LLMs Transform Personal Knowledge Management Systems
-
5 Open-Source Projects Running Transformers on CPUs to GPUs in Pure Java
-
Quansloth Using Google's Turboquant Breaks the VRAM Wall for Local LLMs
-
Apple Research Shows Self-Distillation Significantly Improves Local Code Generation
-
Google Launches Gemma 4 For Advanced On-Device AI
-
Autonet: Decentralized AI Training with Constitutional Governance
-
Local AI Ecosystem Extends Far Beyond Ollama
-
Does RAG Help AI Coding Tools?
-
Unsloth Studio Beta Ships 50+ New Features for Local Model Training and Inference
-
LLM Neuroanatomy II: Modern LLM Hacking and Hints of a Universal Language
-
Qwen3.5-27B Emerges as Sweet Spot for Single-GPU Local Deployment
-
Building a Production AI Receptionist: Practical Local LLM Deployment Case Study
-
Why You Should Use Both ChatGPT and Local LLMs: A Practical Hybrid Approach
-
Llama 8B Matches 70B Performance on Multi-Hop QA Using Structured Prompting
-
Self-Hosted AI Code Review with Local LLMs: Secure Automation Guide
-
Ultra-Compact 28M Parameter Models Show Promise for Specialized Domain Tasks
-
Cursor's Composer 2 Model Analysis – Fine-Tuned Variant of Kimi K2.5
-
Tether's QVAC Introduces Cross-Platform Bitnet LoRA Framework for On-Device AI Training
-
Unsloth Studio: Open-Source Web UI for Training and Running LLMs Locally
-
On-Device AI: Tether's QVAC Fabric Enables Local Training
-
Mistral Releases Small 4 Open-Source Model Under Apache 2.0
-
Mistral Releases Leanstral: First Open-Source Code Agent for Lean 4 Proof Assistant
-
Researcher Discovers Universal "Danger Zone" in Transformer Model Architecture at 50% Depth
-
KAIST Develops World's First Hyper-Personalized On-Device AI Chip
-
Show HN: Generate, Clean, and Prepare LLM Training Data, All-in-One
-
NVIDIA Updates Nemotron 3 122B License, Removes Deployment Restrictions
-
StepFun Releases SFT Dataset Used to Train Step 3.5 Flash for Community Fine-Tuning
-
OpenClaw vs Eigent vs Claude Cowork: Comparing Open-Source AI Collaboration Platforms
-
Fine-Tuned 14B Model Outperforms Claude Opus 4.6 on Ada Code Generation
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
LMF – LLM Markup Format
-
Show HN: AIWatermarkDetector: Detect AI Watermarks in Text or Code
-
Experiment: 0.8B Model Self-Improvement on MacBook Air Yields Surprising Results
-
Texas Instruments Launches NPU-Powered MCUs for Low-Power Edge AI
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
.ispec: Runtime Specification Validation for AI System Consistency
-
Fish Audio Open-Sources S2: Expressive Text-to-Speech with Natural Language Control and 100ms Latency
-
Fine-Tuned Qwen SLMs (0.6–8B) Demonstrate Competitive Performance Against Frontier LLMs on Specialized Tasks
-
FretBench – Testing 14 LLMs on Reading Guitar Tabs Reveals Performance Gaps
-
Sarvam Open-Sources 30B and 105B Reasoning Models
-
Apple Launches MacBook Neo with A18 Pro Chip for Affordable Local AI Inference
-
Sarvam AI Releases 30B and 105B Open-Source Models Trained from Scratch
-
Show HN: BoardMint – A PCB Review Tool That Avoids AI Hallucinations
-
Incrmd: Incremental AI Coding by Editing PROJECT.md
-
Qwen 3.5 vs Qwen 3 Benchmark Analysis: Generational Performance Improvements Visualized
-
Jan Releases Code-Tuned 4B Model for Efficient Local Code Generation and Development Tasks
-
Change Intent Records: The Missing Artifact in AI-Assisted Development
-
Apple Neural Engine Reverse-Engineered for Local Model Training on Mac Mini M4
-
Nummi – AI Companion with Memory and Daily Guidance
-
Google Research Finds Longer Chain-of-Thought Correlates Negatively With Accuracy
-
Extracting 100K Concepts from an 8B LLM
-
Researchers Develop Persistent Memory System for Local LLMs—No RAG Required
-
Show HN: 100% LLM Accuracy–No Fine-Tuning, JSON Only
-
Comparing Manual vs. AI Requirements Gathering: 2 Sentences vs. 127-Point Spec
-
Anthropic Has Never Open-Sourced an LLM: Implications for Local Deployment Strategy
-
nanollama: Open-Source Framework for Training Llama 3 from Scratch with One-Command GGUF Export
-
Wave Field LLM Achieves O(n log n) Scaling: 825M Model Trained to 1B Parameters in 13 Hours
-
O-TITANS: Orthogonal LoRA Framework for Gemma 3 with Google TITANS Memory Architecture
-
CPU-Trained Language Model Outperforms GPU Baseline After 40 Hours
-
Can We Leverage AI/LLMs for Self-Learning?
-
Matmul-Free Language Model Trained on CPU in 1.2 Hours
-
Cohere Releases Tiny Aya: Efficient 3.3B Multilingual Model for 70+ Languages
-
GPU-Accelerated DataFrame Library for Local Inference Workloads
-
Developer Creates Custom Local AI Headshot Generator After Commercial Solutions Fail