Vibe coding for everyone

Which coding model actually finishes the work?

AgentRanks scores coding models and agent stacks on real software-engineering tasks, then shows what each one costs to run. Benchmark first, opinion second.

32Models scored
63Stacks ranked
DeepSWEPrimary benchmark
Sep 12Latest release

As of 21 September 2026, Muse Spark 1.3 leads the AgentRanks coding leaderboard with 75.4 DeepSWE 1.1 pass@1, ahead of DeepSeek V4.1 Flash (74.2) and GPT-6 Astra (74.1). DeepSWE scores the model alone on long-horizon software-engineering tasks.

Common questions about AI coding model rankings

Which AI model is best for coding right now?

Muse Spark 1.3 is first at 75.4 DeepSWE 1.1 pass@1, followed by DeepSeek V4.1 Flash at 74.2 and GPT-6 Astra at 74.1 (21 September 2026 refresh). DeepSWE measures how many long-horizon software-engineering tasks a model finishes on the first attempt, so it rewards seeing work through rather than producing plausible-looking code.

Why does AgentRanks rank by DeepSWE instead of SWE-bench or vibes?

DeepSWE runs long-horizon software-engineering tasks and scores pass@1, so a model has to finish the job unaided. SWE-bench Verified, Terminal-Bench, cost-to-finish, user failure votes and openness are kept as supporting signals rather than the headline number.

Is the best coding model the same as the best coding setup?

No. DeepSWE scores the model on its own; the agent harness decides how much of that capability survives into your repo, and the monthly membership rules decide what it costs you. AgentRanks ranks models, agents and subscription plans separately for that reason.

What is the best open-weight model for coding?

DeepSeek V4.1 Flash is the highest-scoring open-weight model on the table at 74.2 DeepSWE pass@1. Open weights are not the same as runnable at home: several open-weight leaders need server or cluster hardware, which the local hardware table on the leaderboard sets out tier by tier.

Is there a cheap model that still codes well?

DeepSeek V4.1 Flash scores 74.2 DeepSWE 1.1 pass@1 at the low end of the pricing table. Price per token is only half the question — the leaderboard also lists cost-to-finish, because a cheaper model that retries more can end up costing more per completed task.

Price per 1M Tokens

List API price · bar length is output price · lower is better

  • DeepSeek V4 Flash$0.28$0.14 in
  • DeepSeek V4.1 Flash$0.60$0.15 in
  • DeepSeek V4 Pro$0.87$0.435 in
  • GLM-5.2$3$0.95 in
  • Gemini 3.7 Flash$3.75$0.75 in
  • Gemini 3.8 Flash$3.75$0.75 in
  • Muse Spark 1.3$4.25$1.25 in
  • Muse Spark 1.2$4.25$1.25 in
  • Grok 4.6$6$2 in
  • Qwen3.8 Max$6$2 in
  • GPT-5.6 Terra$15$2.50 in
  • Kimi K3$15$3 in
  • Claude Opus 5$25$5 in
  • GPT-5.6 Sol$30$5 in
  • Claude Fable 5$50$10 in
  • GPT-6 Astra$50$10 in

DeepSeek V4 Flash is 178x cheaper on output than Claude Fable 5, so its bar is near-invisible at this scale. Grok 4.6 doubles to $4/$12 above a 200K prompt, GPT-6 Astra to $20/$75 above 272K. Muse Spark 1.3 drops to $0.10/$0.20 if the buyer lets Meta train on their prompts — a data discount, not a price. Gemini 3.5 and 3.6 Flash publish no list price.

Model performance — DeepSWE 1.1

The model on its own, no agent attached. Top two per brand from the independent DeepSWE 1.1 table of 21 September 2026. The 1.1 refresh re-measured every model on harder long-horizon tasks and reordered the board. Hover a bar for context size and price.

Color by
  • 75.4
    Muse Spark 1.3
  • 74.2
    DeepSeek V4.1 Flash
  • 74.1
    GPT-6 Astra
  • 73.7
    Gemini 3.8 Flash
  • 73.0
    GPT-5.6 Sol
  • 70.0
    Claude Fable 5
  • 69.0
    Kimi K3
  • 68.8
    Claude Opus 5
  • 66.9
    GLM-5.3
  • 65.9
    Grok 4.6
  • 65.3
    Gemini 3.7 Flash
  • 63.4
    GLM-5.3-Flash
  • 62.8
    DeepSeek V4 Pro
  • 59.3
    Muse Spark 1.2
  • 58.7
    Qwen3.8 Flash
  • 56.6
    Qwen3.8 Max
  • 54.0
    Grok 4.5
  • 31.0
    Kimi K2.7 Code
OpenAI Anthropic Moonshot xAI Meta Z.ai Google Alibaba DeepSeek

Highlighted marks open weights you can download and run yourself today. Everything unhighlighted is API-only. Both models that were still pending in the last refresh have shipped since: Qwen3.8 Max published its weights on 2 September under Apache 2.0, and Z.ai published GLM-5.3 weights on 29 August.

Every bar is an independently measured DeepSWE 1.1 result; no maker figure is ranked anywhere on this page. The 1.1 table moved things: Claude Opus 5 led the retired 1.0 board at 73.6 and re-measures at 68.8, while DeepSeek V4.1 Flash went from an unranked reported claim on 12 September to a measured 74.2 and second place on 21 September. Hy4 preview remains unscored: its 64.3 is still a maker blind-test checkpoint. Where this page shows a self claim, not from third party it says so on the row.

Monthly memberships, credit for credit

Same-tier comparison of the coding agents you buy by membership, not by token: what the fee buys in quota, and what 100M tokens of agent work costs on each. The open harness at the foot of the table is the pay-as-you-go baseline every membership has to beat.

Product & tierMonthly priceQuota rule (as published)Est. cost per 100M tokensPromo right now
Window plans — you buy a time-slice, not a token budget
ChatGPT Codex Plus / Pro 5x / Pro 20x$20 / $100 / $200Rolling 5-hour message windows (Plus: GPT-6 Astra 5–45, Sol 10–100, Terra 25–200, Luna 250–2,000); the Pro tiers multiply that by 5 and 20; weekly caps apply at an undisclosed count; unused windows bank for 30 daysnot convertible — no token equivalent publishednone current
Claude Code Pro / Max 5x / Max 20x$20 / $100 / $2005-hour session allowance plus a weekly cap; on Max plans Fable models draw at most half the weekly allowance, on Pro they run on usage credits only since September; overage falls back to API-rate creditsnot convertible; community testing puts Max 20x at 6–8 times Pro, not 20 — a class action filed 14 September makes that argument in courtnone current
Grok Build (SuperGrok)$30 SuperGrok / $300 SuperGrok HeavyOne shared weekly pool across chat, voice and Build; xAI publishes no per-action credit rates — the dashboard percentage is the only gauge; Grok 4.6 open to all SuperGrok since 13 August; Heavy buys the maximum weekly allocation, fastest queue and the Grok Bot cloud workspacenot convertible — no published ratesdouble included allocation for the first seven days on the official CLI or Cursor
Credit plans — the quota is a number, and the number converts
Cursor (Anysphere)$20 Pro / $60 Pro+ / $200 UltraThe fee ships as API credit — Pro includes $20 worth, Pro+ $70, Ultra $400; Auto mode is unlimited on any paid plan, manually selected frontier models burn the credit at API cost, Max mode and Composer burn it fastestnot convertible while Auto mode carries unlimited — on manual models Ultra is effectively a 2x credit top-up: $400 of frontier tokens for $200prices rose roughly 60% in September; the credit value still exceeds the fee on Pro+ and Ultra
Zcode / GLM Coding Plan (Z.ai)$18 Lite … $168 top tier (international pricing)Transparent credit system since the 31 July relaunch with a published deduction table; weekend hours deduct at lower multipliers≈ $95 on Lite (community-measured ~19M tokens a month) falling to ≈ $6 on the top tier (≈ 2.7B tokens) — the only plan family here with a checkable token allowanceweekend multipliers live; the 30% annual / 20% quarterly promo ended 15 August
Qoder / QoderWork (Alibaba)$20 Pro / $60 Pro+ / $200 Ultra (international pricing)2,000 / 6,000 / 20,000 credits per cycle, no rollover; an agent request runs 7–12 credits, Quest agent around 50; one shared pool covers the IDE, CLI and the QoderWork desktop agent≈ $200 at 10 credits per request and ~50K tokens per requestno limited-time discount live on the pricing page
Trae Work (ByteDance)$10 Pro (international pricing)600 priority requests a month on built-in models, unlimited basic queue; bring-your-own-model requests cost nothing; the credit-metered CN tiers are a separate price list≈ $33 at 600 requests ≈ 30M tokens a month on the same 50K-token requestfirst month $3; $7.50/mo on annual billing
WorkBuddy (Tencent)¥99 / ¥199 / ¥999 personal; ¥198–¥316 per seat team — no international tier yetFree tier carries 500 credits a month; team seats carry 2,000 credits a month, poolable inside the team; deduction follows model coefficients the pricing page does not listnot derivable until the coefficient table is publishedHy4 preview free on WorkBuddy and CodeBuddy through 10 October, capped daily
No membership at all — the open-harness baseline
DeepSeek Harness (dsh)$0 — MIT open sourceAgent harness launched 13 August as a developer preview, v0.1.5 added V4.1 Flash support on 10 September; CLI plus local web UI, provider-agnostic across roughly 40 model backends; you pay only the underlying API tokens≈ $29 off-peak / ≈ $57 peak per 100M tokens at DeepSeek V4.1 Flash API rates — the cheapest measured frontier-adjacent coding tokens on this page, with no subscription between you and themfree is the promo; the harness itself has no paid tier yet

Method: the 100M column divides the monthly fee by the token allowance a vendor or its community has actually measured; where only credits are published, the conversion assumes ~50K tokens per agent request. Chinese products are priced at their international USD tiers where one exists; WorkBuddy still sells only in RMB. The split in this table is the finding. Window plans (OpenAI, Anthropic, xAI) sell a time-slice and do not say how many tokens fit inside it, so value can only be argued in messages — and Anthropic's 20x claim is now a lawsuit. Credit plans publish a number, and at 100M tokens Zcode's top tier (≈ $6) is the cheapest membership coding deal here by an order of magnitude, with Trae (≈ $33) ahead of Qoder (≈ $200) and the US window plans unpriceable in token terms at all. WorkBuddy is the outlier in the other direction: a credit currency with no published exchange rate is not a price. And the last row is the awkward truth for every membership here: an MIT harness plus V4.1 Flash tokens undercuts all of them per 100M. Every row on this table moves weekly — recheck the vendor page before buying.

Running open models on your own machine

Open weights and runnable at home are not the same thing. Sorted by what the hardware demands, because that decides the question before any score does.

ModelDeepSWEOther scoreParametersMemory at 4-bitSpeedLicence
Runs on one consumer GPU
Qwen 3.8 27B42.228B dense14–17 GB (Q4_K_M)fits a 24 GB card; the Qwen3.8-27B guideApache 2.0
Devstral Small 2not testedSWE-bench Verified 6824B~16 GB15–25 tok/s on RTX 4090Apache 2.0
Codestral 22Bnot testedfill-in-the-middle tuned22B~14 GB20–30 tok/s on RTX 4090Apache 2.0
Muse Glimmer 30Bnot testedSWE-Bench Pro 51.229.6Bunder 20 GB (K-Quant)not publishedApache 2.0
Qwen3 8Bnot testedHumanEval ~768B~5 GB30–45 tok/s on RTX 4090Apache 2.0
StarCoder 2 15Bnot tested15B~30 GBnot publishedOpenRAIL
MiniMax H3 video generationnot applicable2K video, 15s, stereo audio12–16 GB (pruned INT4)not publishedH3 Community
128 GB unified memory desktop — the one model class it genuinely runs is listed here
NVIDIA DGX Spark GB10hardware128 GB unified pool273 GB/s bandwidth~$4,000
AMD Ryzen AI Max+ 395 Strix Halohardware~96 GB to GPU on Windows, ~108 GB on Linux256 GB/s bandwidthvaries by build
Laguna S 2.1 Poolside40.4118B MoE, 8B active~59 GB (4-bit)runs on one 128 GB machine per community reportsopen weights
Needs a server
DeepSeek V4 Flash53.3284B, 13B activeserver classnot publishedMIT
GLM-5.243.8SWE-Bench Pro 62.1753B, 40B active~370 GB, 4x H100not publishedMIT
MiniMax M3not testedLiveBench Coding 68.2428B~233 GB, 3-4x H100not publishedMIT
DeepSeek V4 Pronot testedLiveBench Coding 70.01.6T, 49B active~430 GBnot publishedMIT
MiMo-V2.5-Pronot testedSWE-bench Verified 78.91.02T, 42B activemulti-GPUnot publishedMIT
Kimi K2.7 Code30.5server classnot publishedopen weights
Cluster only
GLM-5.369.0744B~370 GB class, 4x H100not publishedopen weights, published 29 August
Hy4 preview64.3 maker figureTerminal-Bench 2.1 85.4770B, 49B activeclusternot publishedApache 2.0
Kimi K368.5LiveBench Coding 81.52.8T, 104B active~1.4 TB, 64+ acceleratorsnot publishedKimi K3 licence
Qwen3.8 Max57.5DeepSWE 1.1 69.3 (maker)2.4T, 95B activeclusternot publishedApache 2.0, weights published 2 September
What 128 GB of unified memory actually buys you. A DGX Spark or a Strix Halo machine holds far more weights than any consumer graphics card, but capacity is not speed. Both sit around 256–273 GB/s of memory bandwidth against roughly 1,000 GB/s on an RTX 4090, and token generation is bandwidth-bound: a dense model reads its whole weight set for every token produced. A 70B dense model at 4-bit occupies about 40 GB, so 273 GB/s caps it near seven tokens per second before any overhead. Published throughput claims for these machines vary by an order of magnitude between sources, and the high ones do not survive that arithmetic.

Mixture-of-experts models are the exception worth understanding, because only the active experts are read per token: a large MoE with a small active set can run several times faster than a dense model of the same file size. That is what these machines are good at.

What they do not unlock is the top of this table. DeepSeek V4 Flash needs roughly 142 GB at 4-bit and only fits under aggressive quantisation, if at all; GLM-5.2 at ~370 GB and MiniMax M3 at ~233 GB stay out of reach entirely. And the honest limit is not capacity but speed: a 27B dense coder at 4-bit is a ~16 GB weight set, so 273 GB/s streams it near sixteen tokens per second — a fraction of a 4090 and well under a 5090. Neither box is a good host for the Qwen3.8-27B class however it fits on paper. The cards that run these models are an RTX 5090 with 32 GB, or an RTX PRO 5000 / 6000 Blackwell when FP8 weights and multi-agent KV headroom are needed; the 128 GB tier buys longer context on MoE models, not a usable dense-coder loop.
Not every open-weight licence gives you the same rights. Most of the models above ship under MIT (DeepSeek V4 Flash and Pro, GLM-5.2, MiniMax M3, MiMo-V2.5-Pro) or Apache 2.0 (Qwen 3.6, Devstral Small 2, Codestral, Muse Glimmer), which permit commercial use, modification and redistribution. Kimi K3 is the exception: its weights are downloadable, but under a bespoke Kimi K3 licence rather than a standard open-source one, and it carries a revenue-triggered condition for running the model commercially as a hosted service. Downloading it and serving it to paying customers are two different questions. Read the licence before you build on it, and treat the highlight on this page as "weights are published", not "you may do anything with them".

The two halves barely overlap: every model that fits on one GPU is untested on DeepSWE, and every model with a strong DeepSWE score needs server or cluster hardware. Scores from different benchmarks are not comparable to each other. Throughput figures are community measurements at 4-bit quantisation and vary with context length, quantisation and speculative decoding; treat them as ranges, not guarantees.

Where models still break

Eight failure modes ranked by user votes, compared on average severity rather than raw traffic.

Open the 8 concern ranks

Full rankings

Stacks, agents, models, and verified skills. Filter by tab or search.

63 results
01 Claude Code + Mythos 5 limitedFrontier Claude coding stack; restricted trusted-access lane, priced like Fable. 88.2 Context1M Input/Output$10/$50 02 Claude Code + Fable 5 officialWidely released Mythos-class Claude model for deep coding and long-horizon agent work. 86.8 Context1M Input/Output$10/$50 03 ChatGPT Codex + GPT-5.6 Sol previewOpenAI's limited-preview flagship for coding, computer use, and long agentic work. 85.7 Context1M Input/Output$5/$30 04 ChatGPT Codex + GPT-5.6 Terra previewBalanced GPT-5.6 tier: lower cost, expected GPT-5.5-class or better daily coding. 82.9 Context1M Input/Output$2.50/$15 05 ChatGPT Codex + GPT-5.6 Luna previewFastest GPT-5.6 tier; speculative best value for routine ChatGPT Codex tasks. 79.6 Context1M Input/Output$1/$6 06 Cursor + Fable 5 officialPremium Claude model inside an IDE workflow; expensive but attractive for big refactors. 78.1 Context1M Input/Output$10/$50 07 Cursor + GPT-5.6 Sol speculativeLikely high-end IDE stack if Cursor routes or BYOK access opens up. 75.5 Context1M Input/Output$5/$30 08 Claude Code + Opus 4.8 Stable premium Claude Code baseline while Fable/Mythos access settles. 76.2 Context1M Input/Output$5/$25 09 ChatGPT Codex + GPT-5.5 Current widely usable ChatGPT Codex baseline before GPT-5.6 broad access. 69.4 Context258K Input/Output$6/$25 10 Cursor + Gemini 3.1 Pro Strong non-OpenAI closed model option for IDE-first workflows. 57.2 Context1M+ Input/Output$1.5/$6 11 Cursor + GPT-5.4 IDE agent / daily IDE coding with autocomplete and composer workflows 56.2 Context1M Input/Output$5/$20 12 Cursor + DeepSeek V4 Pro IDE agent / daily IDE coding with autocomplete and composer workflows 56.2 Context1M Input/Output$0.44/$0.87 13 Cursor + DeepSeek V4 Flash IDE agent / daily IDE coding with autocomplete and composer workflows 56.1 Context1M Input/Output$0.14/$0.28 14 Cursor + Gemini 3 Flash IDE agent / daily IDE coding with autocomplete and composer workflows 54.1 Context1M+ Input/Output$0.5/$3 15 Devin + GPT-5.5 Cloud agent / long-running autonomous cloud tasks 50.4 Context258K Input/Output$6/$25 16 Cline + Kimi K2.6OS Extension agent / MCP-heavy VS Code workflows and bring-your-own-model setups 48.9 Context256K Input/Output$0.6/$2.5 17 Windsurf + GPT-5.5 IDE agent / AI-native IDE workflows with Cascade-style multi-file edits 48.7 Context258K Input/Output$6/$25 18 Devin + Sonnet 4.6 Cloud agent / long-running autonomous cloud tasks 48.6 Context1M Input/Output$2/$8 19 Cline + DeepSeek V4 ProOS Extension agent / MCP-heavy VS Code workflows and bring-your-own-model setups 48.3 Context1M Input/Output$0.44/$0.87 20 Cline + DeepSeek V4 FlashOS Extension agent / MCP-heavy VS Code workflows and bring-your-own-model setups 48.2 Context1M Input/Output$0.14/$0.28 21 Aider + Kimi K2.6OS CLI agent / git-native terminal editing and experienced open-source users 48.1 Context256K Input/Output$0.6/$2.5 22 Aider + DeepSeek V4 ProOS CLI agent / git-native terminal editing and experienced open-source users 47.5 Context1M Input/Output$0.44/$0.87 23 Aider + DeepSeek V4 FlashOS CLI agent / git-native terminal editing and experienced open-source users 47.4 Context1M Input/Output$0.14/$0.28 24 Windsurf + Sonnet 4.6 IDE agent / AI-native IDE workflows with Cascade-style multi-file edits 47 Context1M Input/Output$2/$8 25 Continue + Sonnet 4.6 Extension agent / custom context providers across VS Code and JetBrains 47 Context1M Input/Output$2/$8 26 Continue + DeepSeek V4 ProOS Extension agent / custom context providers across VS Code and JetBrains 46.7 Context1M Input/Output$0.44/$0.87 27 Windsurf + DeepSeek V4 Pro IDE agent / AI-native IDE workflows with Cascade-style multi-file edits 46.7 Context1M Input/Output$0.44/$0.87 28 Copilot + GPT-5.5 IDE agent / GitHub-centered teams and IDE-native suggestions 43 Context258K Input/Output$6/$25 31 GLM-5.2OS Open-weight coding model baseline / DeepSWE 44% +-2 / listed here because Open-source now includes models, not only stacks. 44 Context1M BasisDeepSWE 29 OpenCode + Kimi K2.6OS CLI agent / provider-flexible terminal coding with open-source control 42.5 Context256K Input/Output$0.6/$2.5 30 OpenCode + DeepSeek V4 ProOS CLI agent / provider-flexible terminal coding with open-source control 41.9 Context1M Input/Output$0.44/$0.87 01 Claude CodeCLI/Desktop / deep terminal workflows and complex multi-step refactors 86 Price$20/$200 TypeCLI/Desktop 02 ChatGPT CodexCLI/Desktop / ChatGPT Codex workflows, terminal automation, and CI-style editing 84 Price$20/$200 TypeCLI/Desktop 03 CursorIDE / daily IDE coding with autocomplete and composer workflows 71 Price$20/mo+ TypeIDE 04 DevinCloud / long-running autonomous cloud tasks 61 Price$500/mo TypeCloud 05 ClineOSExtension / MCP-heavy VS Code workflows and bring-your-own-model setups 61 PriceFree agent TypeExtension 06 AiderOSCLI / git-native terminal editing and experienced open-source users 60 PriceFree agent TypeCLI 07 WindsurfIDE / AI-native IDE workflows with Cascade-style multi-file edits 59 Price$15/mo+ TypeIDE 08 ContinueOSExtension / custom context providers across VS Code and JetBrains 59 PriceFree agent TypeExtension 09 OpenCodeOSCLI / provider-flexible terminal coding with open-source control 53 PriceFree agent TypeCLI 10 CopilotIDE / GitHub-centered teams and IDE-native suggestions 52 Price$10/mo+ TypeIDE 12 Oh-My-CodexOSCLI / open-source Codex enhancements 47 PriceFree agent TypeCLI 12 Qwen CodeOSCLI / Qwen-first open-source CLI workflows 45 PriceFree agent TypeCLI 13 GooseOSCLI / MCP-compatible local automation 43 PriceFree agent TypeCLI 14 Gemini CLICLI / legacy Google CLI workflows 40 PriceFree TypeCLI 15 Hermes AgentOSCLI / skills-based terminal experimentation 39 PriceFree agent TypeCLI 16 Amazon Q DeveloperIDE / AWS-heavy enterprise teams 38 PriceSubscription TypeIDE 17 Antigravity CLICLI / Google terminal-agent exploration 35 PriceTBD TypeCLI 18 Pi AgentOSCLI / MIT harness: four tools, sub-1K-token prompt, bring your own model 35 PriceFree agent TypeCLI 19 Replit AgentBrowser / browser-based app building 34 PriceSubscription TypeBrowser 20 TabnineIDE / enterprise code completion and private deployments 28 Price$12/mo+ TypeIDE 21 Bolt.newBrowser / browser-based full-stack prototypes 25 PriceSubscription TypeBrowser
01 Muse Spark 1.3 DeepSWEMeta / supersedes Muse Spark 1.2 at unchanged pricing / $1.25 / $4.25 per 1M, falling to $0.10 / $0.20 if you let Meta train on your prompts / tops the independent DeepSWE 1.1 table of 21 September; its first-week claim was contested as missing from the public leaderboard and underwhelming in hands-on runs before both independent trackers confirmed the number 75.4 Context1M BasisDeepSWE
02 DeepSeek V4.1 Flash OSDeepSWEDeepSeek / open weights / off-peak $0.15 / $0.60 per 1M, peak $0.30 / $1.20, the cheapest ranked price on the board / replaced V4 Pro in the DeepSeek API around 12 September / independently measured at 74.2 in the 21 September DeepSWE 1.1 table, second on the board 74.2 Contextnot listed BasisDeepSWE
03 GPT-6 Astra DeepSWEOpenAI / OpenAI next-generation flagship / $10 / $50 per 1M, rising to $20 / $75 above a 272K prompt / 74.1 on the independent DeepSWE 1.1 table 74.1 Context272K+ BasisDeepSWE
04 Gemini 3.8 Flash DeepSWEGoogle / third Flash tier in six weeks / $0.75 / $3.75 per 1M introductory through 31 December, rising to $1.50 / $7.50 / 73.7 on the independent DeepSWE 1.1 table 73.7 Context1M BasisDeepSWE
05 GPT-5.6 Sol DeepSWEOpenAI / ChatGPT Codex selectable model / $5 / $30 per 1M / avg task cost $8.39 73.0 Context1.1M BasisDeepSWE
06 Claude Fable 5 DeepSWEAnthropic / leads Terminal-Bench 2.1 inside Claude Code / $10 / $50 per 1M / avg task cost $21.63 70.0 Context1M BasisDeepSWE
07 GPT-5.6 Terra DeepSWEOpenAI / balanced GPT-5.6 tier / $2.50 / $15 per 1M / avg task cost $4.95 70.0 Context1.1M BasisDeepSWE
08 Kimi K3 OSDeepSWEMoonshot AI / 2.8T open weights but needs roughly 1.4TB of accelerator memory / $3 / $15 per 1M 69.0 Context1M BasisDeepSWE
09 Claude Opus 5 DeepSWEAnthropic / frontier tier at unchanged Opus pricing / $5 / $25 per 1M / re-measured at 68.8 on DeepSWE 1.1; the 73.6 that led this board was the 1.0 table 68.8 Context1M BasisDeepSWE
10 GPT-5.6 Luna DeepSWEOpenAI / fastest GPT-5.6 tier / $1 / $6 per 1M / avg task cost $3.03 67.0 Context1.1M BasisDeepSWE
11 GPT-5.5 DeepSWEOpenAI / ChatGPT Codex menu baseline / $6 / $25 per 1M 67.0 Context258K BasisDeepSWE
12 GLM-5.3 OSDeepSWEZ.ai / open weights published 29 August / 744B, same base as GLM-5.2 with gains from post-training alone / CyberGym 84.5 is a security benchmark, not a coding one 66.9 Contextnot listed BasisDeepSWE
13 Grok 4.6 DeepSWExAI / shipped 12 August / $2 / $6 per 1M, doubling above a 200K prompt 65.9 Contextnot listed BasisDeepSWE
14 Gemini 3.7 Flash DeepSWEGoogle / shipped 13 August, three weeks after 3.6 Flash / $0.75 / $3.75 per 1M introductory, rising to $1.50 / $7.50 on 1 January 2027 65.3 Contextnot listed BasisDeepSWE
15 GLM-5.3-Flash OSDeepSWEZ.ai / weights published 26 August, three days before the 744B flagship 63.4 Contextnot listed BasisDeepSWE
16 DeepSeek V4 Pro 0813 OSDeepSWEDeepSeek / MIT open weights, 1.6T with 49B active / $0.435 / $0.87 per 1M / replaced in the DeepSeek API by V4.1 Flash on 12 September 62.8 Context1M BasisDeepSWE
17 Muse Spark 1.2 DeepSWEMeta / superseded by 1.3 / $1.25 / $4.25 per 1M 59.3 Context1M BasisDeepSWE
18 Claude Opus 4.8 DeepSWEAnthropic / previous premium baseline / $5 / $25 per 1M / avg task cost $13.22 59.0 Context1M BasisDeepSWE
19 Qwen3.8 Flash DeepSWEAlibaba / fast Qwen3.8 tier / no list price published here 58.7 Contextnot listed BasisDeepSWE
20 Qwen3.8 Max OSDeepSWEAlibaba / 2.4T flagship / $2 / $6 per 1M / Apache 2.0 weights and the 0902 checkpoint published 2 September / the maker claims 69.3 for 0902, unranked until an independent 1.1 retest of that build 56.6 Context1M BasisDeepSWE
21 Claude Sonnet 5 DeepSWEAnthropic / mid tier / avg task cost $26.40 / long traces, high spend 54.0 Context1M BasisDeepSWE
22 Grok 4.5 DeepSWExAI / high reasoning tier, superseded by Grok 4.6 / $2 / $6 per 1M 54.0 Context500K BasisDeepSWE
23 Muse Spark 1.1 DeepSWEMeta / superseded by 1.2 and 1.3 / $1.25 / $4.25 per 1M 53.0 Context1M BasisDeepSWE
24 GPT-5.4 DeepSWEOpenAI / ChatGPT Codex fallback model / $5 / $20 per 1M / avg task cost $5.65 52.0 Context1M BasisDeepSWE
25 Gemini 3.6 Flash DeepSWEGoogle / fast tier at high reasoning effort / no list price published here 49.0 Contextnot listed BasisDeepSWE
26 GLM-5.2 OSDeepSWEZ.ai / MIT open weights, roughly 370GB at 4-bit / $0.95 / $3 per 1M / avg task cost $3.92 44.0 Context1M BasisDeepSWE
27 Qwen3.8-27B OSDeepSWEAlibaba / Apache 2.0, 28B dense / 14–17 GB at 4-bit, runs on one consumer GPU / the independent 1.1 figure confirms Alibaba's model-card claim to the tenth 42.2 Context262K BasisDeepSWE
28 Laguna S 2.1 OSDeepSWEPoolside / open-weights 118B MoE with 8B active / fits a single 128 GB machine at 4-bit / the strongest score here that runs on desktop silicon 40.4 Context1M BasisDeepSWE
29 Gemini 3.5 Flash DeepSWEGoogle / fast tier at medium reasoning effort / no list price published here 37.0 Contextnot listed BasisDeepSWE
30 Kimi K2.7 Code OSDeepSWEMoonshot AI / previous-generation open weights / no list price published here 31.0 Contextnot listed BasisDeepSWE
31 Claude Sonnet 4.6 DeepSWEAnthropic / previous-generation mid tier at high effort 30.0 Context1M BasisDeepSWE
32 Gemini 3.1 Pro previewGoogle / preview build at high effort / weakest long-horizon result on the table 12.0 Context1M+ BasisDeepSWE
33 Hy4 preview OSno DeepSWETencent / Apache 2.0, 770B MoE with 49B active / $0.83 / $2.50 per 1M / the maker blind-test DeepSWE checkpoint of 64.3 is not ranked until independently measured / free trial on WorkBuddy and CodeBuddy through 10 October, capped daily Context1M BasisNot yet tested
01 GPT Image 2 (medium)Current text-to-image leader on AI Arena public voting. 1386 MetricAI Arena Score SourceArena AI 02 Reve 2.0High-scoring image model on AI Arena public comparisons. 1272 MetricAI Arena Score SourceArena AI 03 Gemini 3.1 Flash Image PreviewGoogle image model with a top-three AI Arena public score. 1270 MetricAI Arena Score SourceArena AI 04 MAI-Image-2.5Microsoft image model ranked fourth on AI Arena public voting. 1257 MetricAI Arena Score SourceArena AI 01 MiniMax H3Open weights since 3 August / 2K video, 15s, native stereo audio in 11 languages / runs locally from 12GB VRAM at pruned INT4 OS MetricOpen weights SourceMiniMax 05 Gemini Omni FlashCurrent AI Arena text-to-video leader by public model score. 1527 MetricAI Arena Score SourceArena AI 06 Dreamina Seedance 2.0 720pTop text-to-video and image-to-video contender from ByteDance. 1466 MetricAI Arena Score SourceArena AI 07 HappyHorse-1.0Third-ranked AI Arena text-to-video model by public score. 1437 MetricAI Arena Score SourceArena AI 08 Veo 3.1 Audio 1080pHigh-ranked Google video model with audio support on AI Arena. 1369 MetricAI Arena Score SourceArena AI 09 Qwen3-TTSOSOpen-source voice design, cloning, multilingual, and streaming speech generation. 12311 MetricGitHub Stars SourceQwenLM GitHub 10 Kokoro 82M v1.0OSSmall open-weight TTS model with Apache-licensed weights. 7811 MetricGitHub Stars Sourcehexgrad GitHub 11 Voxtral TTSOSMistral official TTS model card: multilingual, streaming, and voice cloning. 900 MetricOfficial Signal SourceMistral Docs 12 Step Audio EditXOSOfficial StepFun open model page for speech editing and zero-shot TTS. 700 MetricOfficial Signal SourceHugging Face 01 OpenCodeOSVerified open-source coding agent skill lane; CLI workflow for vibe coding and agentic edits. 184K DomainVibe coding Updated2026-07-12 02 Stable Diffusion WebUIOSVerified creative generation skill; Python app with GPU-recommended local image generation. 164K DomainMedia Updated2026-03-02 03 LangChainOSVerified agent framework for tools, chains, memory, retrieval, and app orchestration. 141K DomainAgent framework Updated2026-07-12 04 ComfyUIOSVerified node-based media workflow; useful for visual assets and local generation pipelines. 120K DomainMedia Updated2026-07-12 05 Gemini CLIOSVerified terminal coding agent; requires Gemini authentication and a local workspace. 105K DomainVibe coding Updated2026-07-12 06 WhisperOSVerified speech transcription skill; Python plus ffmpeg, GPU recommended for speed. 104K DomainMedia Updated2026-04-15 07 MCP ServersOSVerified tool connector collection for MCP-compatible agents; per-server requirements vary. 88K DomainMCP tools Updated2026-07-10 08 OpenHandsOSVerified coding agent platform; Docker strongly recommended for full agent execution. 80K DomainVibe coding Updated2026-07-12 09 ElasticsearchVerified search and analytics engine for retrieval pipelines and data-heavy agent workflows. 77K DomainData/RAG Updated2026-07-12 10 ClineOSVerified VS Code agent workflow with model-provider keys and local tool permissions. 64K DomainVibe coding Updated2026-07-11

Best Entry Points

Useful search landing pages, not generic directory filler.

Qwen3.8-27B on one GPU

What the card has to be, how to start it, and the measured score gap against Opus 5, Kimi K3, GPT-5.5 and DeepSeek V4 Pro.

local setup

What Is Vibe Coding?

The 2025 breakout term, its creator, what it means for non-coders, and the embedded model rank.

core guide

Verified AI skill ranks

Real GitHub-verified skills by domain. Fake open-source rows are blocked instead of padded.

verified only

Highest score stacks

The strongest agent + model combinations by ARscore.

63 stacks

Lowest cost stacks

Free agents and low-cost models for budget-conscious workflows.

price focused

Coding tests

Real task evidence and source-labeled benchmark signals.

evidence layer

24/7 usage loop

Visual workflow for using ChatGPT Codex and Claude Code five-hour sessions without wasting context.

workflow guide

Open source agents

Provider choice, local-model support, and transparent agent behavior.

9 agents

VS Code agents

IDE workflows for autocomplete, context, and fast multi-file edits.

IDE focused

Compare agents

Side-by-side radar and architecture comparisons.

decision tool