~ / coding tests

AI Coding Agent Tests

Static scores are not enough. AgentRanks now treats strict Real Code Work as the primary coding signal: DeepSWE first when available, then SWE-bench Verified, Terminal-Bench, cost-to-finish, and user-reported failure risk.

Real Code Work evidence

Strict benchmark signals to watch. DeepSWE is the preferred primary input because it is closer to long-horizon engineering work.

01 DeepSWE strict coding benchmarkPreferred AgentRanks primary input when available: original long-horizon software-engineering tasks across real repositories. A 70% score is strong; many models drop far lower under this pressure. Primary TestDeepSWE SourceDataCurve Statusstrict 02 GPT-5.6 Sol / Terra / LunaUnknown series for AgentRanks until public users can reproduce coding results. Apply a 50% hype haircut to any vendor or invite-only preview claim. Unknown TestPreview claim Sourceunknown / preview Statusevidence 03 Mythos-class claimsMythos is not a guaranteed rank. Treat as an unverified frontier label unless community tests prove real-world stability. Unknown TestPersona / model-class claim Sourceunknown / no reliable data Statusevidence 04 ChatGPT Codex + GPT-5.5Reported as leading Terminal-Bench v2.1 in June 2026 coverage; needs direct source verification before becoming an AgentRanks official score. 83.4% TestTerminal-Bench v2.1 Sourcereported Statusevidence 05 Claude Code + Fable 5Very strong reported score, but Fable 5 access was suspended by Anthropic on June 12, 2026. 83.1% TestTerminal-Bench v2.1 Sourcereported / unavailable Statuscaution 06 Claude Fable 5Model-level coding signal, not an agent workflow score. 95.0% TestSWE-bench Verified Sourcethird-party snapshot Statusevidence 07 Claude Opus 4.8Useful model-level signal for Claude Code-style stacks. 88.6% TestSWE-bench Verified Sourcethird-party snapshot Statusevidence 08 GPT-5.5Useful model-level signal for Codex-style stacks. 82.6% TestSWE-bench Verified Sourcethird-party snapshot Statusevidence

Current score audit

Legacy ARscore is useful for discovery, but Real Code Work is now the primary signal. DeepSWE-style strict coding scores beat broad marketing claims.

primary_score = strict_real_code_work first; legacy_ARscore remains proxy evidence only
Verdict: 70% on a strict benchmark is genuinely strong. Most models are still much lower under messy, long-horizon code work.
#Current top proxy stacksARscoreAgentModel SWEEvidence status
1Claude Code + Opus 4.8
Proxy only: Claude Code architecture score times Opus 4.8 SWE-bench
76.28688.6%needs AgentRanks run
2Claude Code + Opus 4.7
Proxy only: Claude Code architecture score times Opus 4.7 SWE-bench
75.38687.6%needs AgentRanks run
3ChatGPT Codex + GPT-5.5
Proxy only: ChatGPT Codex architecture score times GPT-5.5 SWE-bench
69.48482.6%needs AgentRanks run
4Claude Code + Sonnet 4.6
Proxy only: Claude Code architecture score times Sonnet 4.6 SWE-bench
68.58679.6%needs AgentRanks run
5ChatGPT Codex + GPT-5.3 Codex
Proxy only: ChatGPT Codex architecture score times GPT-5.3 Codex SWE-bench
63.28475.2%needs AgentRanks run
6Cursor + Opus 4.8
Proxy only: Cursor architecture score times Opus 4.8 SWE-bench
62.97188.6%needs AgentRanks run
7Cursor + GPT-5.5
Proxy only: Cursor architecture score times GPT-5.5 SWE-bench
58.67182.6%needs AgentRanks run
8Cursor + DeepSeek V4 Pro
Proxy only: Cursor architecture score times DeepSeek V4 Pro SWE-bench
56.27180.6%needs AgentRanks run

AgentRanks test pack

The product upgrade path: run identical coding tasks against every stack and publish pass rate, cost, time, retries, and code-quality notes.

Bug fix with hidden tests

Patch a real failing issue in a small repo, add regression coverage, and pass hidden tests.

30% weight

Terminal autonomy

Inspect a repo, run commands, diagnose failures, and produce a working fix without manual file hints.

20% weight

Feature build

Implement a small UI/API feature from product requirements with tests and no unrelated churn.

15% weight

Refactor safety

Refactor a shared module while preserving behavior and avoiding over-broad edits.

15% weight

Cost and latency

Measure tokens, wall time, retries, and cost per accepted solution.

10% weight

Maintainability review

Score code clarity, test quality, minimalism, and ease of future modification.

10% weight

Benchmark sources

Primary and secondary sources that should feed the evidence layer.

DeepSWE

Preferred primary source for Real Code Work because it targets original long-horizon software engineering tasks across real repositories.

primary benchmark

SWE-bench Verified

Real GitHub issue resolution with tests; strong signal, but increasingly saturated and source quality varies by run.

official benchmark

Terminal-Bench 2.0

Closer to real coding-agent work because the system must inspect files, run commands, debug, and finish tasks.

official benchmark

SWE-bench / Vals AI

Useful secondary snapshot with recent Fable 5, Opus 4.8, and GPT-5.5 figures, but should not replace primary benchmark links.

third-party snapshot