Tagged "evaluation"
- Testing Top Local LLMs Against ChatGPT and Claude Reveals Performance Gaps
- Enprompta: Prompt Registry, LLM Evals, and Observability for Production AI Apps
- Round-Trip Correctness: New Metric for Generative AI Process Modeling
- Agentic Test Processes and LLM Benchmarks: Evaluating Local AI Agents
- GPT-5.6 Sol vs. Claude Fable 5 in CNC Red Alert 2 Benchmark
- 'AI Code Is Insane Trash' – David Gerard on Code Generation Quality
- Making AI Code Review Measurable
- Lessons from Building Evals for Financial AI Agents
- Show HN: Tail Panic – a multiplayer game designed for AI agents
- Show HN: Veritrooper – find what your AI gets wrong about your own docs
- LLM Memory Systems Benchmark: High Recall, Near-Zero Precision for Tested Systems
- Exploration Got Cheap. Human Review Did Not
- LLM Hallucinations in the Wild
- Anthropic Develops Tool to Detect When Claude Recognizes It's Being Tested
- Eval Skills for AI Agents
- Control AI Risk with Pre-Built Frameworks and Ready-to-Run Evaluations
- How to Test AI Agents When They Never Give the Same Answer Twice
- LLM Personalization Breaks Down in High-Stakes Finance
- Show HN: SkillCompass – Open-Source Quality Evaluator for Your AI Skills
- SkillCompass – Diagnose and Improve AI Agent Skills Across 6 Dimensions
- FretBench – Testing 14 LLMs on Reading Guitar Tabs Reveals Performance Gaps
- AI Agent Reliability Tracker
- No, Local LLMs Can't Replace ChatGPT or Gemini — I Tried
- How Do You Know Which SKILL.md Is Good?