Evidence record / pi-coding-agent

A dated record of one change, kept so the writing that cites it can be checked. Compiled from the sources listed below by the research run, not written for reading. The judgment lives in the signals and issues that cite it, below.

2026-09-21-pi-coding-agent-harnesstax-study-same-model-same-success-up-to-five-times-the-cost

UC Berkeley and Arena published HarnessTax on 2026-09-16. It ran seven models through Claude Code, Codex CLI and Pi on the same 30 sampled tasks from SWE-bench Lite and Terminal-Bench 2.0, 21 model-harness pairs. Across 42 within-model comparisons, a Fisher exact test found harness choice had little effect on task success holding the model constant, while cost moved up to fivefold. Their example: Claude Fable 5 completed 97.8 percent of attempts in Claude Code at an average $1.33 and 96.7 percent in Pi at $0.67. The same window, DeepSeek evaluated its V4.1 Flash model across eight harness configurations and a practitioner summary reports it peaked in the minimal ones.

Channel: docs-only (third-party study published 2026-09-16). Half: capability.

Operator consequence: For API-metered work, measure your own tasks in a minimal harness before assuming the vendor harness earns its overhead; the study says success is mostly the model and cost is mostly the harness. The sample is 30 tasks per benchmark and two benchmarks; it does not cover long-horizon or subscription-bundled use.

Receipt

Finding metadata

Run: 2026-09-21-weekly-digest-2026-08-20_2026-09-21-frontier-v0

Finding ID: 2026-09-21-pi-coding-agent-harnesstax-study-same-model-same-success-up-to-five-times-the-cost

Accepted signals

Profile citations

Source links

Primary links, including exact changelog lines when available.

Versioned source: run artifact