Which coding model actually finishes the work?
AgentRanks scores coding models and agent stacks on real software-engineering tasks, then shows what each one costs to run. Benchmark first, opinion second.
As of 21 September 2026, Muse Spark 1.3 leads the AgentRanks coding leaderboard with 75.4 DeepSWE 1.1 pass@1, ahead of DeepSeek V4.1 Flash (74.2) and GPT-6 Astra (74.1). DeepSWE scores the model alone on long-horizon software-engineering tasks.
Common questions about AI coding model rankings
Which AI model is best for coding right now?
Muse Spark 1.3 is first at 75.4 DeepSWE 1.1 pass@1, followed by DeepSeek V4.1 Flash at 74.2 and GPT-6 Astra at 74.1 (21 September 2026 refresh). DeepSWE measures how many long-horizon software-engineering tasks a model finishes on the first attempt, so it rewards seeing work through rather than producing plausible-looking code.
Why does AgentRanks rank by DeepSWE instead of SWE-bench or vibes?
DeepSWE runs long-horizon software-engineering tasks and scores pass@1, so a model has to finish the job unaided. SWE-bench Verified, Terminal-Bench, cost-to-finish, user failure votes and openness are kept as supporting signals rather than the headline number.
Is the best coding model the same as the best coding setup?
No. DeepSWE scores the model on its own; the agent harness decides how much of that capability survives into your repo, and the monthly membership rules decide what it costs you. AgentRanks ranks models, agents and subscription plans separately for that reason.
What is the best open-weight model for coding?
DeepSeek V4.1 Flash is the highest-scoring open-weight model on the table at 74.2 DeepSWE pass@1. Open weights are not the same as runnable at home: several open-weight leaders need server or cluster hardware, which the local hardware table on the leaderboard sets out tier by tier.
Is there a cheap model that still codes well?
DeepSeek V4.1 Flash scores 74.2 DeepSWE 1.1 pass@1 at the low end of the pricing table. Price per token is only half the question — the leaderboard also lists cost-to-finish, because a cheaper model that retries more can end up costing more per completed task.
List API price · bar length is output price · lower is better
DeepSeek V4 Flash$0.28$0.14 in
DeepSeek V4.1 Flash$0.60$0.15 in
DeepSeek V4 Pro$0.87$0.435 in
GLM-5.2$3$0.95 in
Gemini 3.7 Flash$3.75$0.75 in
Gemini 3.8 Flash$3.75$0.75 in
Muse Spark 1.3$4.25$1.25 in
Muse Spark 1.2$4.25$1.25 in
Grok 4.6$6$2 in
Qwen3.8 Max$6$2 in
GPT-5.6 Terra$15$2.50 in
Kimi K3$15$3 in
Claude Opus 5$25$5 in
GPT-5.6 Sol$30$5 in
Claude Fable 5$50$10 in
GPT-6 Astra$50$10 in
DeepSeek V4 Flash is 178x cheaper on output than Claude Fable 5, so its bar is near-invisible at this scale. Grok 4.6 doubles to $4/$12 above a 200K prompt, GPT-6 Astra to $20/$75 above 272K. Muse Spark 1.3 drops to $0.10/$0.20 if the buyer lets Meta train on their prompts — a data discount, not a price. Gemini 3.5 and 3.6 Flash publish no list price.
Model performance — DeepSWE 1.1
The model on its own, no agent attached. Top two per brand from the independent DeepSWE 1.1 table of 21 September 2026. The 1.1 refresh re-measured every model on harder long-horizon tasks and reordered the board. Hover a bar for context size and price.
Muse Spark 1.3
DeepSeek V4.1 Flash
GPT-6 Astra
Gemini 3.8 Flash
GPT-5.6 Sol
Claude Fable 5
Kimi K3
Claude Opus 5
GLM-5.3
Grok 4.6
Gemini 3.7 Flash
GLM-5.3-Flash
DeepSeek V4 Pro
Muse Spark 1.2
Qwen3.8 Flash
Qwen3.8 Max
Grok 4.5
Kimi K2.7 Code
Highlighted marks open weights you can download and run yourself today. Everything unhighlighted is API-only. Both models that were still pending in the last refresh have shipped since: Qwen3.8 Max published its weights on 2 September under Apache 2.0, and Z.ai published GLM-5.3 weights on 29 August.
Every bar is an independently measured DeepSWE 1.1 result; no maker figure is ranked anywhere on this page. The 1.1 table moved things: Claude Opus 5 led the retired 1.0 board at 73.6 and re-measures at 68.8, while DeepSeek V4.1 Flash went from an unranked reported claim on 12 September to a measured 74.2 and second place on 21 September. Hy4 preview remains unscored: its 64.3 is still a maker blind-test checkpoint. Where this page shows a self claim, not from third party it says so on the row.
Monthly memberships, credit for credit
Same-tier comparison of the coding agents you buy by membership, not by token: what the fee buys in quota, and what 100M tokens of agent work costs on each. The open harness at the foot of the table is the pay-as-you-go baseline every membership has to beat.
| Product & tier | Monthly price | Quota rule (as published) | Est. cost per 100M tokens | Promo right now |
|---|---|---|---|---|
| Window plans — you buy a time-slice, not a token budget | ||||
| ChatGPT Codex Plus / Pro 5x / Pro 20x | $20 / $100 / $200 | Rolling 5-hour message windows (Plus: GPT-6 Astra 5–45, Sol 10–100, Terra 25–200, Luna 250–2,000); the Pro tiers multiply that by 5 and 20; weekly caps apply at an undisclosed count; unused windows bank for 30 days | not convertible — no token equivalent published | none current |
| Claude Code Pro / Max 5x / Max 20x | $20 / $100 / $200 | 5-hour session allowance plus a weekly cap; on Max plans Fable models draw at most half the weekly allowance, on Pro they run on usage credits only since September; overage falls back to API-rate credits | not convertible; community testing puts Max 20x at 6–8 times Pro, not 20 — a class action filed 14 September makes that argument in court | none current |
| Grok Build (SuperGrok) | $30 SuperGrok / $300 SuperGrok Heavy | One shared weekly pool across chat, voice and Build; xAI publishes no per-action credit rates — the dashboard percentage is the only gauge; Grok 4.6 open to all SuperGrok since 13 August; Heavy buys the maximum weekly allocation, fastest queue and the Grok Bot cloud workspace | not convertible — no published rates | double included allocation for the first seven days on the official CLI or Cursor |
| Credit plans — the quota is a number, and the number converts | ||||
| Cursor (Anysphere) | $20 Pro / $60 Pro+ / $200 Ultra | The fee ships as API credit — Pro includes $20 worth, Pro+ $70, Ultra $400; Auto mode is unlimited on any paid plan, manually selected frontier models burn the credit at API cost, Max mode and Composer burn it fastest | not convertible while Auto mode carries unlimited — on manual models Ultra is effectively a 2x credit top-up: $400 of frontier tokens for $200 | prices rose roughly 60% in September; the credit value still exceeds the fee on Pro+ and Ultra |
| Zcode / GLM Coding Plan (Z.ai) | $18 Lite … $168 top tier (international pricing) | Transparent credit system since the 31 July relaunch with a published deduction table; weekend hours deduct at lower multipliers | ≈ $95 on Lite (community-measured ~19M tokens a month) falling to ≈ $6 on the top tier (≈ 2.7B tokens) — the only plan family here with a checkable token allowance | weekend multipliers live; the 30% annual / 20% quarterly promo ended 15 August |
| Qoder / QoderWork (Alibaba) | $20 Pro / $60 Pro+ / $200 Ultra (international pricing) | 2,000 / 6,000 / 20,000 credits per cycle, no rollover; an agent request runs 7–12 credits, Quest agent around 50; one shared pool covers the IDE, CLI and the QoderWork desktop agent | ≈ $200 at 10 credits per request and ~50K tokens per request | no limited-time discount live on the pricing page |
| Trae Work (ByteDance) | $10 Pro (international pricing) | 600 priority requests a month on built-in models, unlimited basic queue; bring-your-own-model requests cost nothing; the credit-metered CN tiers are a separate price list | ≈ $33 at 600 requests ≈ 30M tokens a month on the same 50K-token request | first month $3; $7.50/mo on annual billing |
| WorkBuddy (Tencent) | ¥99 / ¥199 / ¥999 personal; ¥198–¥316 per seat team — no international tier yet | Free tier carries 500 credits a month; team seats carry 2,000 credits a month, poolable inside the team; deduction follows model coefficients the pricing page does not list | not derivable until the coefficient table is published | Hy4 preview free on WorkBuddy and CodeBuddy through 10 October, capped daily |
| No membership at all — the open-harness baseline | ||||
| DeepSeek Harness (dsh) | $0 — MIT open source | Agent harness launched 13 August as a developer preview, v0.1.5 added V4.1 Flash support on 10 September; CLI plus local web UI, provider-agnostic across roughly 40 model backends; you pay only the underlying API tokens | ≈ $29 off-peak / ≈ $57 peak per 100M tokens at DeepSeek V4.1 Flash API rates — the cheapest measured frontier-adjacent coding tokens on this page, with no subscription between you and them | free is the promo; the harness itself has no paid tier yet |
Method: the 100M column divides the monthly fee by the token allowance a vendor or its community has actually measured; where only credits are published, the conversion assumes ~50K tokens per agent request. Chinese products are priced at their international USD tiers where one exists; WorkBuddy still sells only in RMB. The split in this table is the finding. Window plans (OpenAI, Anthropic, xAI) sell a time-slice and do not say how many tokens fit inside it, so value can only be argued in messages — and Anthropic's 20x claim is now a lawsuit. Credit plans publish a number, and at 100M tokens Zcode's top tier (≈ $6) is the cheapest membership coding deal here by an order of magnitude, with Trae (≈ $33) ahead of Qoder (≈ $200) and the US window plans unpriceable in token terms at all. WorkBuddy is the outlier in the other direction: a credit currency with no published exchange rate is not a price. And the last row is the awkward truth for every membership here: an MIT harness plus V4.1 Flash tokens undercuts all of them per 100M. Every row on this table moves weekly — recheck the vendor page before buying.
Running open models on your own machine
Open weights and runnable at home are not the same thing. Sorted by what the hardware demands, because that decides the question before any score does.
| Model | DeepSWE | Other score | Parameters | Memory at 4-bit | Speed | Licence |
|---|---|---|---|---|---|---|
| Runs on one consumer GPU | ||||||
| Qwen 3.8 27B | 42.2 | — | 28B dense | 14–17 GB (Q4_K_M) | fits a 24 GB card; the Qwen3.8-27B guide | Apache 2.0 |
| Devstral Small 2 | not tested | SWE-bench Verified 68 | 24B | ~16 GB | 15–25 tok/s on RTX 4090 | Apache 2.0 |
| Codestral 22B | not tested | fill-in-the-middle tuned | 22B | ~14 GB | 20–30 tok/s on RTX 4090 | Apache 2.0 |
| Muse Glimmer 30B | not tested | SWE-Bench Pro 51.2 | 29.6B | under 20 GB (K-Quant) | not published | Apache 2.0 |
| Qwen3 8B | not tested | HumanEval ~76 | 8B | ~5 GB | 30–45 tok/s on RTX 4090 | Apache 2.0 |
| StarCoder 2 15B | not tested | — | 15B | ~30 GB | not published | OpenRAIL |
| MiniMax H3 video generation | not applicable | 2K video, 15s, stereo audio | — | 12–16 GB (pruned INT4) | not published | H3 Community |
| 128 GB unified memory desktop — the one model class it genuinely runs is listed here | ||||||
| NVIDIA DGX Spark GB10 | hardware | — | — | 128 GB unified pool | 273 GB/s bandwidth | ~$4,000 |
| AMD Ryzen AI Max+ 395 Strix Halo | hardware | — | — | ~96 GB to GPU on Windows, ~108 GB on Linux | 256 GB/s bandwidth | varies by build |
| Laguna S 2.1 Poolside | 40.4 | — | 118B MoE, 8B active | ~59 GB (4-bit) | runs on one 128 GB machine per community reports | open weights |
| Needs a server | ||||||
| DeepSeek V4 Flash | 53.3 | — | 284B, 13B active | server class | not published | MIT |
| GLM-5.2 | 43.8 | SWE-Bench Pro 62.1 | 753B, 40B active | ~370 GB, 4x H100 | not published | MIT |
| MiniMax M3 | not tested | LiveBench Coding 68.2 | 428B | ~233 GB, 3-4x H100 | not published | MIT |
| DeepSeek V4 Pro | not tested | LiveBench Coding 70.0 | 1.6T, 49B active | ~430 GB | not published | MIT |
| MiMo-V2.5-Pro | not tested | SWE-bench Verified 78.9 | 1.02T, 42B active | multi-GPU | not published | MIT |
| Kimi K2.7 Code | 30.5 | — | — | server class | not published | open weights |
| Cluster only | ||||||
| GLM-5.3 | 69.0 | — | 744B | ~370 GB class, 4x H100 | not published | open weights, published 29 August |
| Hy4 preview | 64.3 maker figure | Terminal-Bench 2.1 85.4 | 770B, 49B active | cluster | not published | Apache 2.0 |
| Kimi K3 | 68.5 | LiveBench Coding 81.5 | 2.8T, 104B active | ~1.4 TB, 64+ accelerators | not published | Kimi K3 licence |
| Qwen3.8 Max | 57.5 | DeepSWE 1.1 69.3 (maker) | 2.4T, 95B active | cluster | not published | Apache 2.0, weights published 2 September |
Mixture-of-experts models are the exception worth understanding, because only the active experts are read per token: a large MoE with a small active set can run several times faster than a dense model of the same file size. That is what these machines are good at.
What they do not unlock is the top of this table. DeepSeek V4 Flash needs roughly 142 GB at 4-bit and only fits under aggressive quantisation, if at all; GLM-5.2 at ~370 GB and MiniMax M3 at ~233 GB stay out of reach entirely. And the honest limit is not capacity but speed: a 27B dense coder at 4-bit is a ~16 GB weight set, so 273 GB/s streams it near sixteen tokens per second — a fraction of a 4090 and well under a 5090. Neither box is a good host for the Qwen3.8-27B class however it fits on paper. The cards that run these models are an RTX 5090 with 32 GB, or an RTX PRO 5000 / 6000 Blackwell when FP8 weights and multi-agent KV headroom are needed; the 128 GB tier buys longer context on MoE models, not a usable dense-coder loop.
The two halves barely overlap: every model that fits on one GPU is untested on DeepSWE, and every model with a strong DeepSWE score needs server or cluster hardware. Scores from different benchmarks are not comparable to each other. Throughput figures are community measurements at 4-bit quantisation and vary with context length, quantisation and speculative decoding; treat them as ranges, not guarantees.
Where models still break
Eight failure modes ranked by user votes, compared on average severity rather than raw traffic.
Full rankings
Stacks, agents, models, and verified skills. Filter by tab or search.
Best Entry Points
Useful search landing pages, not generic directory filler.
Qwen3.8-27B on one GPU
What the card has to be, how to start it, and the measured score gap against Opus 5, Kimi K3, GPT-5.5 and DeepSeek V4 Pro.
What Is Vibe Coding?
The 2025 breakout term, its creator, what it means for non-coders, and the embedded model rank.
Verified AI skill ranks
Real GitHub-verified skills by domain. Fake open-source rows are blocked instead of padded.
Highest score stacks
The strongest agent + model combinations by ARscore.
Lowest cost stacks
Free agents and low-cost models for budget-conscious workflows.
Coding tests
Real task evidence and source-labeled benchmark signals.
24/7 usage loop
Visual workflow for using ChatGPT Codex and Claude Code five-hour sessions without wasting context.
Open source agents
Provider choice, local-model support, and transparent agent behavior.
VS Code agents
IDE workflows for autocomplete, context, and fast multi-file edits.
Compare agents
Side-by-side radar and architecture comparisons.







