---
title: "Hamel Husain"
description: "AI product engineer, machine learning educator, and specialist in LLM evaluation, error analysis, and domain-grounded AI systems."
section: "entities"
type: "entity"
created: "2026-08-22"
updated: "2026-08-24"
confidence: "high"
tags: ["person", "educator", "evaluation", "context-engineering"]
canonical: "https://pyweb.dev/wiki/hamel-husain"
---
# Hamel Husain

Hamel Husain is an AI engineer, educator, and co-founder of Parlance Labs. He is one of the foremost advocates for rigorous, empirical evaluation methodologies in generative AI and agentic systems, emphasizing human error analysis, custom data viewers, and practical data science fundamentals over generic benchmarks.

## Core Philosophy & Contributions

1. **"Look At Your Data":** Husain argues that the single highest-ROI activity in AI engineering is qualitative [error analysis and evals](/wiki/error-analysis-and-evals) on real production traces. Generic off-the-shelf metrics (like generic hallucination or helpfulness scores) create an illusion of progress while obscuring domain-specific failure modes.
2. **The Optimization Hierarchy:** When building AI products, teams must exhaust context engineering, prompt refinement, and harness tooling before resorting to model post-training or fine-tuning.
3. **Data Science Fundamentals in AI:** Viewing LLM judges as supervised classifiers requiring human-annotated validation sets, precision/recall tracking, and strict partition boundaries.
4. **Active Learning & Tooling:** Collaborating with researchers like [shreya shankar](/wiki/shreya-shankar) on active-learning-assisted trace labeling and failure discovery tools.
5. **Verifiability as product design:** "It's hard to eval" is a product smell — artifacts hard for the builder to verify are hard for users too; design checkable artifacts before building evals ([designing for verifiability](/wiki/designing-for-verifiability)). [[source: hamel-husain-it-s-hard-to-eval-is-a-product-smell-2026]](/wiki/raw/articles/hamel-husain-it-s-hard-to-eval-is-a-product-smell-2026)
6. **The Revenge of the Data Scientist:** every recurring eval pitfall maps to a missing data-science fundamental; the agent harness itself is largely data science. [[source: hamel-husain-the-revenge-of-the-data-scientist-2026]](/wiki/raw/articles/hamel-husain-the-revenge-of-the-data-scientist-2026)

He co-teaches *AI Evals for Engineers and PMs*, with over 4,500 students from 500+ companies (including OpenAI, Anthropic, and Google), and previously worked at Airbnb and GitHub, including early LLM research used by OpenAI for code understanding. [[source: hamel-husain-do-automated-evals-work-2026]](/wiki/raw/articles/hamel-husain-do-automated-evals-work-2026)

## Related
- [shreya shankar](/wiki/shreya-shankar)
- [error analysis and evals](/wiki/error-analysis-and-evals)
- [designing for verifiability](/wiki/designing-for-verifiability)
- [automated eval engineering](/wiki/automated-eval-engineering)
- [closed loop agent improvement](/wiki/closed-loop-agent-improvement)
- [context engineering](/wiki/context-engineering)

---

## Agent Navigation

cluster: person (170 pages) | betweenness: 80

### References (outbound)
- [Error Analysis and Evals](https://pyweb.dev/wiki/error-analysis-and-evals.md)
- [Shreya Shankar](https://pyweb.dev/wiki/shreya-shankar.md)
- [Designing for Verifiability](https://pyweb.dev/wiki/designing-for-verifiability.md)
- [Automated Eval Engineering](https://pyweb.dev/wiki/automated-eval-engineering.md)
- [Closed-Loop Agent Improvement](https://pyweb.dev/wiki/closed-loop-agent-improvement.md)
- [Context Engineering](https://pyweb.dev/wiki/context-engineering.md)

### Referenced by (inbound)
- [Agent Harness Engineering](https://pyweb.dev/wiki/agent-harness-engineering.md)
- [Designing for Verifiability](https://pyweb.dev/wiki/designing-for-verifiability.md)
- [Error Analysis and Evals](https://pyweb.dev/wiki/error-analysis-and-evals.md)
- [Eval-Driven Development](https://pyweb.dev/wiki/eval-driven-development.md)
- [Evals Skills](https://pyweb.dev/wiki/evals-skills.md)
- [Airbnb](https://pyweb.dev/wiki/airbnb.md)
- [Shreya Shankar](https://pyweb.dev/wiki/shreya-shankar.md)

### Evidence (verified primary sources)
- [hamel-ai-product-engineering-notes-2026](https://pyweb.dev/wiki/raw/articles/hamel-ai-product-engineering-notes-2026.md) | origin: https://hamel.dev/notes/llm/ai-product-engineering/ | ingested: 2026-08-22 | sha256: 7b35f29cda741a46974fa2e5585b42d5e2e805566373b9e4a3d45ef42d10f2bb
- [parlance-labs-do-automated-evals-work-2026](https://pyweb.dev/wiki/raw/articles/parlance-labs-do-automated-evals-work-2026.md) | origin: https://parlance-labs.com/blog/posts/auto-evals/index.html | ingested: 2026-08-22 | sha256: 8a4ef31b67fcd59160d7ca6567554f676b70129841cb02787c805362547bcfcb
- [hamel-husain-do-automated-evals-work-2026](https://pyweb.dev/wiki/raw/articles/hamel-husain-do-automated-evals-work-2026.md) | origin: https://hamel.dev/ | ingested: 2026-08-24 | sha256: 2abfbf6161c7174f9fe900490af15cbd3dd6cb2ed96ca5aa3c8eab988078c884
- [hamel-husain-the-revenge-of-the-data-scientist-2026](https://pyweb.dev/wiki/raw/articles/hamel-husain-the-revenge-of-the-data-scientist-2026.md) | origin: https://hamel.dev/blog/posts/revenge/ | ingested: 2026-08-24 | sha256: a5c947dbab1261c1eb647d34f07e4b7e562438a2450d0178715070310a6ac8f5
- [hamel-husain-it-s-hard-to-eval-is-a-product-smell-2026](https://pyweb.dev/wiki/raw/articles/hamel-husain-it-s-hard-to-eval-is-a-product-smell-2026.md) | origin: https://hamel.dev/blog/posts/eval-smell/ | ingested: 2026-08-24 | sha256: bf1a77e2fc299c4d3d2489ff9800774406e8d475eb54f9578917f56fe4a7543c

### Machine endpoints
- Knowledge graph: https://pyweb.dev/api/graph.json
- Graph analysis: https://pyweb.dev/api/graph-analysis.json
- Context index: https://pyweb.dev/llms.txt
