---
title: "Shreya Shankar"
description: "Computer science researcher at UC Berkeley focusing on data management, ML systems, and active-learning tooling for LLM error analysis and evaluation."
section: "entities"
type: "entity"
created: "2026-08-22"
updated: "2026-08-22"
confidence: "high"
tags: ["person", "educator", "evaluation"]
canonical: "https://pyweb.dev/wiki/shreya-shankar"
---
# Shreya Shankar

Shreya Shankar is a computer science researcher at UC Berkeley specializing in machine learning systems, data management, and operational tooling for AI evaluation.

## Key Research & Systems

1. **Error Discovery & Active Learning:** Pioneered active-learning workflows where an AI assistant observes real-time human trace annotations, updates a dynamic failure taxonomy, and proactively queries similar unlabeled records.
2. **Criteria Drift in Evaluation:** Demonstrated empirically that humans often cannot specify complete evaluation rubrics up front; criteria naturally emerge through the iterative act of labeling and observing model behaviors.
3. **Model Cascades (BARGAIN):** Created algorithms for optimal model cascading, routing high-confidence queries to cheap, small models while reserving frontier models for ambiguous edge cases to cut inference costs up to 86% without sacrificing quality.
4. **Data Agent Benchmark (DAB):** Designed benchmarks reflecting realistic, messy multi-database environments to measure planning and execution failures in analytical agents.

## Related
- [hamel husain](/wiki/hamel-husain)
- [error analysis and evals](/wiki/error-analysis-and-evals)
- [automated eval engineering](/wiki/automated-eval-engineering)
- [closed loop agent improvement](/wiki/closed-loop-agent-improvement)

---

## Agent Navigation

cluster: person (170 pages) | betweenness: 44.6

### References (outbound)
- [Hamel Husain](https://pyweb.dev/wiki/hamel-husain.md)
- [Error Analysis and Evals](https://pyweb.dev/wiki/error-analysis-and-evals.md)
- [Automated Eval Engineering](https://pyweb.dev/wiki/automated-eval-engineering.md)
- [Closed-Loop Agent Improvement](https://pyweb.dev/wiki/closed-loop-agent-improvement.md)

### Referenced by (inbound)
- [Error Analysis and Evals](https://pyweb.dev/wiki/error-analysis-and-evals.md)
- [Eval-Driven Development](https://pyweb.dev/wiki/eval-driven-development.md)
- [Anthropic](https://pyweb.dev/wiki/anthropic.md)
- [Hamel Husain](https://pyweb.dev/wiki/hamel-husain.md)
- [Johann Rehberger](https://pyweb.dev/wiki/johann-rehberger.md)

### Evidence (verified primary sources)
- [parlance-labs-do-automated-evals-work-2026](https://pyweb.dev/wiki/raw/articles/parlance-labs-do-automated-evals-work-2026.md) | origin: https://parlance-labs.com/blog/posts/auto-evals/index.html | ingested: 2026-08-22 | sha256: 8a4ef31b67fcd59160d7ca6567554f676b70129841cb02787c805362547bcfcb
- [hamel-ai-product-engineering-notes-2026](https://pyweb.dev/wiki/raw/articles/hamel-ai-product-engineering-notes-2026.md) | origin: https://hamel.dev/notes/llm/ai-product-engineering/ | ingested: 2026-08-22 | sha256: 7b35f29cda741a46974fa2e5585b42d5e2e805566373b9e4a3d45ef42d10f2bb

### Machine endpoints
- Knowledge graph: https://pyweb.dev/api/graph.json
- Graph analysis: https://pyweb.dev/api/graph-analysis.json
- Context index: https://pyweb.dev/llms.txt
