wiki / entities / hamel-husain

Hamel Husain

high confidence updated 2026-08-24 person · educator · evaluation · context-engineering

Hamel Husain

Hamel Husain is an AI engineer, educator, and co-founder of Parlance Labs. He is one of the foremost advocates for rigorous, empirical evaluation methodologies in generative AI and agentic systems, emphasizing human error analysis, custom data viewers, and practical data science fundamentals over generic benchmarks.

Core Philosophy & Contributions

  1. “Look At Your Data”: Husain argues that the single highest-ROI activity in AI engineering is qualitative error analysis and evals on real production traces. Generic off-the-shelf metrics (like generic hallucination or helpfulness scores) create an illusion of progress while obscuring domain-specific failure modes.
  2. The Optimization Hierarchy: When building AI products, teams must exhaust context engineering, prompt refinement, and harness tooling before resorting to model post-training or fine-tuning.
  3. Data Science Fundamentals in AI: Viewing LLM judges as supervised classifiers requiring human-annotated validation sets, precision/recall tracking, and strict partition boundaries.
  4. Active Learning & Tooling: Collaborating with researchers like shreya shankar on active-learning-assisted trace labeling and failure discovery tools.
  5. Verifiability as product design: “It’s hard to eval” is a product smell — artifacts hard for the builder to verify are hard for users too; design checkable artifacts before building evals (designing for verifiability). [source: hamel-husain-it-s-hard-to-eval-is-a-product-smell-2026]
  6. The Revenge of the Data Scientist: every recurring eval pitfall maps to a missing data-science fundamental; the agent harness itself is largely data science. [source: hamel-husain-the-revenge-of-the-data-scientist-2026]

He co-teaches AI Evals for Engineers and PMs, with over 4,500 students from 500+ companies (including OpenAI, Anthropic, and Google), and previously worked at Airbnb and GitHub, including early LLM research used by OpenAI for code understanding. [source: hamel-husain-do-automated-evals-work-2026]