wiki / entities / tessl
Tessl
Machine ingest — raw context
loading…
~… tokensappend .md to any wiki URL for this view
Tessl
AI developer tools and software engineering research organization based in London, UK. Tessl conducts empirical research on agentic software development, benchmark alignment, and autonomous system harnesses. [source: tessl-coding-benchmarks-misaligned-system-harness-2026]
Research & Frameworks
- Position on Coding Benchmarks: Tessl researchers (Maria I. Gorinova, Dru Knox, Amy Heineike, et al.) authored the foundational 2026 ACM SIGKDD position paper arguing that pre-agent coding benchmarks (SWE-Bench, HumanEval) fail to evaluate real-world agentic software engineering because they conflate raw model capability with the composite system harness. [source: tessl-coding-benchmarks-misaligned-system-harness-2026]
- NS2 System Harness: An open-source, issue-driven multi-agent harness treating GitHub issues as the state machine for task decomposition, execution, test quality validation, and multi-tier feedback loops. [source: tessl-coding-benchmarks-misaligned-system-harness-2026]
Cross-links
- agentic code quality — feedback signal tiers and verification gates
- agent harness engineering — composite system harness design
- designing for verifiability — objective verification criteria over reference solutions
- error analysis and evals — benchmark and eval methodology
Evidence — verified primary sources
| tessl-coding-benchmarks-misaligned-system-harness-2026 | https://arxiv.org/abs/2606.17799 | ingested 2026-08-27 sha256:4b68e983226a… |
Graph context
References (4)
Agentic Code Qualityfeedback signal tiers and verification gatesAgent Harness Engineeringcomposite system harness designDesigning for Verifiabilityobjective verification criteria over reference solutionsError Analysis and Evalsbenchmark and eval methodology Referenced by (1)
Agentic Software Factorysystem harness benchmarks and NS2 orchestration