AI Coding Agent News
Recent market notes for ranking changes, pricing shifts, model releases, and benchmark risk, refreshed 21 September 2026. Rumors and source quality are clearly marked.
Sep 21, 2026AgentRanks moves the board to DeepSWE 1.1: Muse Spark 1.3 leads at 75.4, Claude Opus 5 re-measures at 68.8The 1.1 refresh re-measured every model on harder long-horizon tasks and reordered the top ten: DeepSeek V4.1 Flash, unranked as a reported claim nine days ago, is now a measured 74.2 and second place; GPT-6 Astra third at 74.1, Gemini 3.8 Flash fourth at 73.7. Opus 5, the 1.0 leader at 73.6, sits ninth at 68.8. The full re-ranked table is on the leaderboard.TypeBenchmarkSourceIndependentSep 20, 2026StepFun announces Step 5 Preview: 600B MoE flagship with 1M context, weights promised for 15 OctoberOpen-weight frontier pressure keeps coming from China: Step 5 Preview joins Hy4 and GLM-5.3 on the same release cadence. Nothing is scored until weights are public. On the radar.TypeModelSourceOfficialSep 18, 2026Anonymous free 262K multimodal model Union Alpha appears on OpenRouter; community sleuths point at Z.aiStealth drops have preceded major releases before. Treat the provenance as unconfirmed and the free window as marketing until a named model ships.TypeModelSourceRumorSep 15, 2026Hands-on tests of DeepSeek V4.1 Flash describe parity-grade benchmarks alongside logic missesThe 74.2 figure comes from a single third-party run, not the independent table, the 21 September 1.1 refresh measured it independently at 74.2 and ranked it second after all. Practical reports split: near-frontier on long tasks, more surprises in both directions than Opus 5 or Astra at the same price tier.TypeBenchmarkSourceReportedSep 14, 2026DeepSeek retires V4 Pro from its API in favour of V4.1 Flash: off-peak $0.15/$0.60 per 1M, peak $0.30/$1.20The cheapest seat near the top of the leaderboard moves by replacing its own flagship, not by raising prices. V4 Pro 0813 keeps its measured 62.8 in the table.TypePricingSourceOfficialSep 12, 2026DeepSeek V4.1 Flash ships as open weights alongside DeepSeek Harness v0.1.5KV-cache compressed toward 890 bytes per token per the release notes; serving economics, not just scores, are the story. Already on third-party hosts including DigitalOcean.TypeModelSourceOfficialSep 11, 2026Sakana AI ships Fugu Max and Fugu Ultra v2 on a low-cost multi-model orchestration pitchA different axis of competition: routing several models per task instead of buying one frontier seat. Not yet scored on DeepSWE.TypeModelSourceOfficialSep 9, 2026Abliteration.ai starts selling a commercial API of a refusal-reduced GLM-5.3 forkZ.ai Apache 2.0 weights permit modification, so the open question is distribution norms, not licence law. Relevant to anyone evaluating agent safety baselines: agent safety guide.TypeSecuritySourceReportedSep 5, 2026Tencent extends the free Hy4 preview window on WorkBuddy and CodeBuddy to 10 October, with daily token capsTwo weeks of frontier-class free access, server-crushingly popular on day one. The cap, not the price tag, is the thing to plan around. WorkBuddy profile.TypePricingSourceOfficialSep 5, 2026Muse Spark 1.3 tops the DeepSWE 1.1 claims and spends its first week being argued withMeta's 75.4 did not appear on the public leaderboard for days and hands-on runs underperformed lower-ranked models, which is exactly the pattern this board punishes. By 21 September both independent trackers list it first; the row still carries the dispute in its note.TypeBenchmarkSourceReportedSep 3, 2026Community blind tests put Hy4 preview around GLM and Kimi class at 2.99, but self-hosting wants 8x B300770B MoE with 49B active keeps inference sane per token; the weight set does not fit anywhere near a single consumer GPU. Vendor blind-test margins were tight.TypeBenchmarkSourceReportedSep 2, 2026Alibaba publishes Qwen3.8 Max weights under Apache 2.0; the 0902 API build follows the same dayThe announced-not-released marker that sat on this row for three weeks is retired. $2/$6 per 1M with $0.25 cache reads. The maker DeepSWE 1.1 figure of 69.3 does not enter the ranking: the board keeps the independently measured figure of 56.6 on the 1.1 table.TypeModelSourceOfficialSep 1, 2026GPT-6 Astra and Gemini 3.8 Flash enter the independent DeepSWE table within days of each otherBoth land within a tenth of the frontier on the 1.0 table and neither publishes a list price here; the 21 September 1.1 refresh places them at 74.1 and 73.7.TypeBenchmarkSourceIndependentAug 29, 2026Z.ai releases GLM-5.3 open weights, 744B, one day past its stated safety-hardening windowThe independent 1 September table scores it at 69.0, sixth on the board and the best independently measured open-weight model. CyberGym 84.5 circulated alongside it is a security benchmark, not a coding one.TypeModelSourceOfficialAug 28, 2026Tencent releases and open-sources Hy4 preview: Apache 2.0, ~1M context, vendor DeepSWE of 64.3Terminal-Bench 2.1 at 85.4 is the headline the maker leads with; the DeepSWE number is its own checkpoint, not a third-party measurement, so the board lists it unscored until measured.TypeModelSourceOfficialAug 26, 2026GLM-5.3-Flash weights land three days before the flagshipZ.ai release split had people downloading the wrong checkpoint; the Flash tier scores an independent 63.4, ahead of DeepSeek V4 Pro measured 62.8.TypeModelSourceOfficialAug 21, 2026GLM-5.3 score jump credited to post-training on an unchanged base model; distillation questions followIf the same 744B base with different post-training changes the ranking this much, per-vendor fine-tune economics get very interesting. Unresolved reporting so far.TypeBenchmarkSourceReportedAug 14, 2026ChatGPT free tier goes unlimited on text after the 13 August board refreshPressure on sub-$20 subscriptions from the consumer side while model makers fight at the API edge. No ranking change; relevant to low-cost stacks.TypeMarketSourceOfficialAug 13, 2026Alibaba ships Qwen3.8-27B: the hot single-card coder, 28B dense under Apache 2.014-17 GB at 4-bit means a 24 GB card runs it; 262K native context. Its DeepSWE figure of 42.2, published on the model card, was later confirmed to the tenth by the independent 1.1 table. Bandwidth, not capacity, is why 128 GB unified-memory boxes are a poor host. Setup guide.TypeModelSourceOfficialAug 13, 2026AgentRanks 13 August refresh measures Grok 4.6, Gemini 3.7 Flash and DeepSeek V4 Pro 0813All three had shipped on vendor figures; the refresh moved Grok 4.6 from a claimed 65.9 to a measured 66.7 and DeepSeek V4 Pro from 62.7 to 62.8.TypeBenchmarkSourceIndependentAug 13, 2026Claude Opus 5 takes the DeepSWE lead at 73.6 in the 13 August tableAhead of GPT-5.6 Sol (72.7) and Claude Fable 5 (69.7) at unchanged Opus pricing of $5/$25 per 1M. Superseded by the 21 September refresh.TypeModelSourceIndependentAug 12, 2026xAI ships Grok 4.6 at $2/$6 per 1M, doubling above a 200K promptLong-context surcharges keep landing mid-table; cost-to-finish, not list price, decides leaderboard placement for long-horizon agent runs.TypeModelSourceOfficialAug 7, 2026Qwen3.8 Max announced with an open-weight promise; the board labels it announced, not releasedKept as the marker of how long a gap this was: five weeks between announcement and published weights, during which the promise was not a score.TypeMarketSourceOfficial For traffic and trust, AgentRanks should treat news as ranked impact notes: what changed, which agents/models it affects, and whether the source is official, third-party, directory-listed, or speculative.