wiki / entities / model-evaluation-and-threat-research
Model Evaluation & Threat Research
Machine ingest — raw context
loading…
~… tokensappend .md to any wiki URL for this view
Model Evaluation & Threat Research
Model Evaluation & Threat Research (METR) studies frontier AI capabilities and their real-world effects. Its task-completion horizon work models agent success against human expert task duration, while its developer-productivity randomized trial found that early-2025 AI tools increased completion time by 19% for experienced maintainers working in familiar mature repositories. [source: metr-task-completion-time-horizons-2026]^[raw/papers/metr-developer-productivity-rct-2025.md]
Related
- agentic quality evidence — benchmark and productivity evidence
- error analysis and evals — evaluation design and error analysis
- agentic code quality — proposed factory controls
Evidence — verified primary sources
| metr-task-completion-time-horizons-2026 | https://metr.org/time-horizons/ | ingested 2026-08-27 sha256:01fa095450fa… |
| raw/papers/metr-developer-productivity-rct-2025.md | internal workspace doc |