---
title: "Tessl"
description: "AI software development company researching developer platforms, agentic software engineering benchmarks, and system harness architectures."
section: "entities"
type: "entity"
created: "2026-08-27"
updated: "2026-08-27"
confidence: "high"
tags: ["company", "evaluation"]
canonical: "https://pyweb.dev/wiki/tessl"
---
# Tessl

AI developer tools and software engineering research organization based in London, UK. Tessl conducts empirical research on agentic software development, benchmark alignment, and autonomous system harnesses. [[source: tessl-coding-benchmarks-misaligned-system-harness-2026]](/wiki/raw/articles/tessl-coding-benchmarks-misaligned-system-harness-2026)

## Research & Frameworks

- **Position on Coding Benchmarks:** Tessl researchers (Maria I. Gorinova, Dru Knox, Amy Heineike, et al.) authored the foundational 2026 ACM SIGKDD position paper arguing that pre-agent coding benchmarks (SWE-Bench, HumanEval) fail to evaluate real-world agentic software engineering because they conflate raw model capability with the composite **system harness**. [[source: tessl-coding-benchmarks-misaligned-system-harness-2026]](/wiki/raw/articles/tessl-coding-benchmarks-misaligned-system-harness-2026)
- **NS2 System Harness:** An open-source, issue-driven multi-agent harness treating GitHub issues as the state machine for task decomposition, execution, test quality validation, and multi-tier feedback loops. [[source: tessl-coding-benchmarks-misaligned-system-harness-2026]](/wiki/raw/articles/tessl-coding-benchmarks-misaligned-system-harness-2026)

## Cross-links
- [agentic code quality](/wiki/agentic-code-quality) — feedback signal tiers and verification gates
- [agent harness engineering](/wiki/agent-harness-engineering) — composite system harness design
- [designing for verifiability](/wiki/designing-for-verifiability) — objective verification criteria over reference solutions
- [error analysis and evals](/wiki/error-analysis-and-evals) — benchmark and eval methodology

---

## Agent Navigation

cluster: person (170 pages) | betweenness: 14.3

### References (outbound)
- [Agentic Code Quality](https://pyweb.dev/wiki/agentic-code-quality.md)
- [Agent Harness Engineering](https://pyweb.dev/wiki/agent-harness-engineering.md)
- [Designing for Verifiability](https://pyweb.dev/wiki/designing-for-verifiability.md)
- [Error Analysis and Evals](https://pyweb.dev/wiki/error-analysis-and-evals.md)

### Referenced by (inbound)
- [Agentic Software Factory](https://pyweb.dev/wiki/agentic-software-factory.md)

### Evidence (verified primary sources)
- [tessl-coding-benchmarks-misaligned-system-harness-2026](https://pyweb.dev/wiki/raw/articles/tessl-coding-benchmarks-misaligned-system-harness-2026.md) | origin: https://arxiv.org/abs/2606.17799 | ingested: 2026-08-27 | sha256: 4b68e983226a2731dc5e89da3b27b4097495393d2208e4ad9d63c5aa68be3812

### Machine endpoints
- Knowledge graph: https://pyweb.dev/api/graph.json
- Graph analysis: https://pyweb.dev/api/graph-analysis.json
- Context index: https://pyweb.dev/llms.txt
