wiki / raw / hamel-husain-do-automated-evals-work-2026

Do Automated Evals Work?

updated 2026-08-24

Original source: https://hamel.dev/ SHA256: 2abfbf6161c7174f9fe900490af15cbd3dd6cb2ed96ca5aa3c8eab988078c884

Do Automated Evals Work?

Hamel Husain’s Blog – Hamel’s Blog I’m a machine learning engineer with 20+ years of experience. I’ve worked at Airbnb and GitHub, including early LLM research used by OpenAI for code understanding. I’ve also led and contributed to popular open-source machine-learning tools. I’m currently working to bring data science back to AI: helping teams debug, analyze, and measure their systems. I call this “evals,” and it’s the focus of my teaching, consulting, and writing. 🎓 Learn With Me I co-teach AI Evals for Engineers and PMs, with over 4,500 students from 500+ companies (including OpenAI, Anthropic, and Google). Check the page for upcoming cohorts. 📮 Blog I often share my experience building AI products. Below is a selection of my long-form writing. Subscribe To My Newsletter Date Title 8/12/26 AI Product Engineering Notes 7/11/26 Do Automated Evals Work? 6/29/26 “It’s Hard to Eval” Is a Product Smell 3/26/26 The Revenge of the Data Scientist 3/2/26 Evals Skills for Coding Agents 1/18/26 Why I Stopped Using nbdev 10/1/25 Selecting The Right AI Evals Tool 7/11/25 Stop Saying RAG Is Dead 6/23/25 Inspect AI, An OSS Python Library For LLM Evals 5/28/25 LLM Evals: Everything You Need to Know 3/24/25 A Field Guide to Rapidly Improving AI Products 1/19/25 Thoughts On A Month With Devin 12/13/24 nbsanity - Share Notebooks as Polished Web Pages in Seconds 11/30/24 Building an Audience Through Technical Writing: Strategies and Mistakes 10/29/24 Using LLM-as-a-Judge For Evaluation: A Complete Guide 10/10/24 Concurrency Foundations For FastHTML 7/29/24 An Open Course on LLMs, Led by Practitioners 6/1/24 What We’ve Learned From A Year of Building with LLMs 4/12/24 Debugging AI With Adversarial Validation 3/29/24 Your AI Product Needs Evals 3/27/24 Is Fine-Tuning Still Valuable? 2/14/24 Fuck You, Show Me The Prompt. 1/11/24 How To Debug Axolotl 1/9/24 Dokku: my favorite personal serverless platform 12/17/23 Tokenization Gotchas 11/15/23 Tools for curating LLM Data 10/28/23 vLLM & Large Models 10/15/23 Optimizing LLM latency 5/30/23 On commercializing nbdev 1/16/23 Why Should ML Engineers Learn Kubernetes? 7/28/22 nbdev + Quarto: A new secret weapon for productivity 2/9/22 Notebooks in production with Metaflow 12/18/20 ghapi, a new third-party Python client for the GitHub API 11/20/20 Nbdev: A literate programming environment that democratizes software engineering best practices 9/1/20 fastcore: An Underrated Python Library 9/1/20 Data Science Meets Devops: MLOps with Jupyter, Git, & Kubernetes 3/6/20 GitHub Actions: Providing Data Scientists With New Superpowers. 2/21/20 Introducing fastpages, An easy to use blogging platform with extra features for Jupyter Notebooks. 2/5/20 Python Concurrency: The Tricky Bits 9/20/19 CodeSearchNet Challenge: Evaluating the State of Semantic Code Search 4/10/19 How to Automate Tasks on GitHub With Machine Learning for Fun and Profit 5/29/18 How To Create Natural Language Semantic Search for Arbitrary Objects With Deep Learning 1/18/18 How To Create Magical Data Products Using Sequence-to-Sequence Models 12/16/17 How Docker Can Help You Become A More Effective Data Scientist 5/10/17 Automated Machine Learning — A Paradigm Shift That Accelerates Data Scientist Productivity @ Airbnb No matching items 📬 Follow Me The best way to follow my activity is X (Twitter).