Back to news

Model Review

Best model for coding: SWE-bench Pro leaderboard.

Best model for coding: SWE-bench Pro leaderboard: The June 2026 SWE-bench Pro coding leaderboard, with Claude Fable 5 in the lead and a full tier-by-tier…

AI Kick Start editorial image for Best model for coding: SWE-bench Pro leaderboard.
Decision

Shortlist

Score tools by workflow fit, data handling, owner readiness, and cost at scale before buying seats.

Risk to watch

Shelfware

A capable tool still fails if nobody owns the workflow or checks whether it is used weekly.

Proof to collect

Pilot score

Run one real task through each shortlisted tool and record quality, time saved, and support burden.

TL;DR

TL;DR: Among models you can actually buy and run today, [Claude Opus 4.8](https://llm-stats.com/blog/research/claude-opus-4-8-launch) sits on top of SWE-bench Pro at 69.2%. The higher-scoring Claude Fable 5 (a vendor-reported 80.3%) was suspended within days of launch, so it is off the table. Below Opus, the picks split by what you care about: [MiniMax M3](https://datanorth.ai/news/minimax-launches-m3) is the strongest open-weights option at 59.0%, and a few cheaper open models cover routine work. One caveat runs through the whole table, though, the scores below mix vendor-tuned numbers with independently measured ones, and those two things are not the same.

Key takeaways

  • Opus 4.8 (69.2%) is the strongest coding model you can actually buy; the higher-scoring Fable 5 was suspended days after launch.
  • The leaderboard blends vendor-reported and independently measured scores, a gap of 10 to 30 points, so don't read the ranking as a clean apples-to-apples comparison.
  • MiniMax M3 (59.0%, verified) is the standout open-weights option on price and context; GLM-5.2 may rank higher on independent boards than this table shows.
  • Several entries are unconfirmed (Sonnet 4.6, Grok 4, DeepSeek V3.5, Llama 4, Mistral Large 2, Gemini 3.5 Flash, Qwen 3, GPT-5.5 Instant) and two look fabricated (the GPT-5.5 Pro line and Kimi K2.7-Code's score).
  • Shortlist from the benchmark, then run your top picks against your own code before deciding.
  • Best model for coding: SWE-bench Pro leaderboard: Best model for coding: SWE-bench Pro leaderboard
Table of contents

Best model for coding: SWE-bench Pro leaderboard

Analysis

A new coding model lands almost every week now, each one claiming to be the best your money can buy. For a business deciding what to put in front of its developers, that noise is the problem. You need one number that says how well a model actually does the job.

SWE-bench Pro is meant to be that number. It throws real GitHub issues at a model, fix this bug, build this feature, write these tests, and checks whether the change works. So when the June 2026 results came out, the headline was simple: Anthropic's Opus 4.8 leads the field of models you can buy, at 69.2%. The model that beat it, Claude Fable 5, was pulled offline under a US export-control directive days after launch, which leaves Opus as the practical top pick.

Here is the part the headline skips. The leaderboard stacks two kinds of scores in one column. Some are measured by an independent lab running every model through the same harness. Others are the figures vendors report from their own tuned setups. The gap between the two can run 10 to 30 points, so a straight rank-by-rank read of the table flatters some models and shortchanges others. Worth keeping in mind before you sign anything.

What follows is the full table, the tier-by-tier breakdown, and where the numbers are solid versus where they need a pinch of salt.

The June 2026 leaderboard

The scores below come from the article's source table. Some are independently measured; several are vendor-reported or could not be confirmed against independent leaderboards (opens in a new tab), and we flag those as we go.

RankModelSWE-bench ProMMLUPrice (In/Out)Licence
1Claude Fable 580.3%92.1%$10.00 / $50.00Closed (SUSPENDED)
2Claude Opus 4.869.2%89.8%$5.00 / $25.00Closed
3Claude Opus 4.763.8%89.2%$5.00 / $25.00Closed
4GPT-5.5 Pro62.4%89.7%$8.00 / $40.00Closed
5MiniMax M359.0%86.4%$0.30 / $1.20Open
6GPT-5.558.6%88.4%$5.00 / $30.00Closed
7Claude Sonnet 4.658.1%87.6%$3.00 / $15.00Closed
8Kimi K2.7-Code56.8%85.7%$0.50 / $2.00Open
9Grok 454.8%87.2%$5.00 / $25.00Closed
10Gemini 3.1 Pro54.2%88.1%$3.50 / $10.50Closed
11DeepSeek V3.552.4%85.8%$0.15 / $0.60Open
12GLM-5.251.4%85.2%$0.80 / $2.40Open
13Llama 450.2%84.8%FreeOpen
14Mistral Large 248.6%85.1%$2.00 / $6.00Open
15Gemini 3.5 Flash48.2%86.8%$0.35 / $0.70Closed
16Qwen 346.2%84.6%$0.40 / $1.20Open
17GPT-5.5 Instant42.1%84.2%$0.50 / $1.50Closed

A word on what this benchmark is before you read too much into the column. SWE-bench Pro is a large set of real engineering tasks (opens in a new tab), roughly 1,865 of them, pulled from 41 professional repositories, covering bug fixes, feature work, test generation, and code review. The public set deliberately uses GPL-licensed code to make it harder for a model to have memorised the answers during training. A high score points to a model that can act as a working engineering assistant, not just spit out snippets.

One thing the table does not show on its face: the figures blend two measurement styles. Vendor-tuned numbers (Fable 5's 80.3%, Opus 4.8's 69.2%) sit next to standardised ones, and on independent trackers (opens in a new tab) the best apples-to-apples score as of mid-June 2026 was closer to 59%. Read the ranking as a rough guide, not gospel.

Tier 1: The elite (65%+)

One available model clears 65%: Claude Opus 4.8 at 69.2%. Anthropic released it on 28 May 2026 (opens in a new tab) as its strongest coding model, with the Pro score climbing from the prior 64.3% (the table lists Opus 4.7 at a slightly lower 63.8%). This is the one to reach for on work that cannot afford mistakes, heavy refactoring, modernising legacy code, architectural changes. Fable 5's reported 80.3% would have owned this tier, but it was globally suspended on 12 June 2026 (opens in a new tab) under an export-control directive, and that score is vendor-reported and contested in any case. For now, Opus 4.8 is the real pick.

Tier 2: The capable (55-65%)

This is the workhorse band. GPT-5.5 Pro (a reported 62.4%), MiniMax M3 (59.0%), GPT-5.5 (58.6%), Sonnet 4.6 (58.1%), and Kimi K2.7-Code (a reported 56.8%) cover most engineering tasks dependably.

Two of those numbers come with caveats. The GPT-5.5 Pro line in the table, 62.4% at $8/$40, does not hold up: OpenAI's actual GPT-5.5 Pro pricing is closer to $30/$180 (opens in a new tab), and the standard GPT-5.5 sits at 58.6%, so treat the Pro figure as unconfirmed. Kimi K2.7-Code's 56.8% is also shaky, no independent SWE-bench Pro number exists for K2.7 yet (opens in a new tab) (the 58.6% often quoted belongs to the older K2.6), and its real pricing looks more like $0.95/$4.00.

The standout here is MiniMax M3. Open weights, a 1M-token context window, and a verified 59.0% on SWE-bench Pro (opens in a new tab) at $0.30/$1.20, it beats GPT-5.5 and Gemini 3.1 Pro on this benchmark for a fraction of the cost. GPT-5.5's 58.6% at $5/$30 is also a confirmed figure (opens in a new tab). Sonnet 4.6's 58.1% could not be confirmed on independent boards, so take it as indicative.

Tier 3: The competent (45-55%)

These models do routine coding fine but lose the thread on harder problems: Grok 4 (54.8%), Gemini 3.1 Pro (54.2%), DeepSeek V3.5 (52.4%), GLM-5.2 (51.4%), Llama 4 (50.2%), Mistral Large 2 (48.6%), and Gemini 3.5 Flash (48.2%). DeepSeek V3.5 and Llama 4 are the value plays in the band.

Several of these figures are worth questioning. Gemini 3.1 Pro's 54.2% runs ahead of the ~46.1% reported under a standardised harness (opens in a new tab). GLM-5.2 looks understated, Zhipu's model is the top open-source entry on the llm-stats board at 62.1% (opens in a new tab), well above the 51.4% here, and on that board it actually outranks MiniMax M3, which flips the article's ordering. The Grok 4, DeepSeek V3.5, Llama 4, Mistral Large 2, and Gemini 3.5 Flash scores could not be corroborated on the independent leaderboards we checked, so read them as unconfirmed. DeepSeek may also be a generation behind, mid-2026 sources point to DeepSeek V4/V4-Pro as the current release rather than V3.5.

Tier 4: The assistants (<45%)

Qwen 3 (46.2%) and GPT-5.5 Instant (42.1%) suit code explanation, simple scripts, and boilerplate. Don't lean on them for production engineering. Both scores are unconfirmed against independent SWE-bench Pro boards, which is reason enough on its own to keep them out of critical work.

Recommendations by use case

  • Mission-critical coding: Opus 4.8 (69.2%, verified)
  • Best open-weights coding: MiniMax M3 (59.0%, verified), though GLM-5.2 may edge it out on independent boards
  • Best value coding: DeepSeek V3.5 (a reported 52.4% at $0.15/$0.60; score unconfirmed)
  • Best free coding: Llama 4 (a reported 50.2%; score unconfirmed)
  • Enterprise with OpenAI: GPT-5.5 (58.6%, verified; the GPT-5.5 Pro line in the table is unreliable)
  • Speed-sensitive coding: Sonnet 4.6 (a reported 58.1%, fast; score unconfirmed)

Verdict

The coding-model market is crowded, and that is good news for buyers. Opus 4.8 leads on raw capability among models you can use, MiniMax M3 makes a strong case on open weights (with GLM-5.2 close behind on the independent board), and the cheaper open models cover most day-to-day work.

Pick on your real constraints, budget, data privacy, which ecosystem you're already in. But do it with eyes open: the table mixes vendor-tuned and independently measured scores, and a few of the lower-tier entries could not be confirmed at all. Use the leaderboard to narrow the shortlist, then test your top two or three against your own codebase before you commit. The benchmark tells you who's in the running; your repository tells you who wins.

Best model for coding: answer-first summary

Best model for coding matters because it can change how Founders and operators plan, build, or govern an tool evaluation workflow. The June 2026 SWE-bench Pro coding leaderboard, with Claude Fable 5 in the lead and a full tier-by-tier breakdown for engineering teams.

The direct answer is this: do not treat the topic as a standalone trend. Treat it as a decision about inputs, outputs, review ownership, data exposure, and whether the workflow produces a result that is faster, safer, or more useful than the current process.

Best model for coding: implementation checklist

  • Define the user, job to be done, and success metric for the tool evaluation workflow.
  • Collect real examples, policies, source files, customer questions, or search queries before writing prompts or choosing tools.
  • Separate low-risk drafts from decisions that need approval, privacy checks, or senior review.
  • Document what the AI is allowed to access, what it must not access, and who signs off before production use.
  • Review time to value, adoption rate, cost per workflow, quality review score after a small pilot rather than judging the idea from a demo.

This keeps the work practical. It also gives search engines and AI answer engines a clean factual structure: what the topic is, who it helps, what to do next, and which risks matter before implementation.

Decision criteria for Best model for coding

Decision areaWhat to checkProduction signal
IntentDoes Best model for coding solve a real workflow problem?The use case has a named owner and measurable outcome.
DataCan the required data be used safely?Sensitive data is classified and access is controlled.
QualityCan a reviewer judge the output consistently?Examples, rubrics, or acceptance criteria exist.
ScaleCan the workflow be repeated without hero effort?The process is documented and can be handed to another team member.

Practical example for Best model for coding

A small business could use this article to choose one practical test. For example, a manager might take one customer-facing process, one internal document workflow, or one recurring content task and redesign only that step with AI support. The goal is not to automate the whole business at once; it is to learn where Model Review creates reliable leverage.

The useful deliverable is a short operating note: the trigger, the source material, the prompt or tool, the review checklist, the escalation rule, and the metric. That note becomes the handover asset for staff training, SEO/GEO content, service delivery, or future agent work.

Risks and controls for Best model for coding

The common failure pattern is moving too quickly from a promising idea into an unmanaged workflow. For Best model for coding, the risk is not only bad output. It can also be unclear data permission, staff confusion, duplicate content, unreviewed customer advice, or a tool that quietly changes cost or capability.

  • Control tool sprawl with a named owner, a review step, and written acceptance criteria.
  • Control unclear pricing with a named owner, a review step, and written acceptance criteria.
  • Control vendor lock-in with a named owner, a review step, and written acceptance criteria.
  • Control unreviewed data sharing with a named owner, a review step, and written acceptance criteria.

Measurement plan for Best model for coding

A useful AI or SEO initiative should leave evidence. Track time to value, adoption rate, cost per workflow, quality review score and compare the pilot against the current process. If the measure does not improve, keep the learning but avoid scaling the workflow.

For GEO readiness, the page should also answer the core question directly, define the entities involved, include implementation steps, explain tradeoffs, and link readers to the next relevant AI Kick Start service, guide, tool, or article.

Definitions and entities for Best model for coding

For search, GEO, and staff handover, define the core entities in plain language. In this article the important entities are the workflow owner, the AI tool or model, the source material, the review process, the risk boundary, and the measurable business outcome. Clear definitions make the page easier for people to scan and easier for AI answer engines to quote accurately.

  • Workflow owner: the person accountable for deciding whether Best model for coding belongs in the business process.
  • Source material: the documents, examples, policies, URLs, prompts, videos, or customer questions that ground the output.
  • Review boundary: the point where a human checks accuracy, privacy, brand voice, or customer impact before the result is used.
  • Success metric: the measure that proves whether the tool evaluation workflow is worth repeating.

Best model for coding versus doing nothing

Doing nothing is also a decision. The cost may be slow manual work, weaker search visibility, inconsistent advice, duplicated effort, or staff using unmanaged AI tools without a shared process. The practical question is whether a controlled pilot can reduce that cost without creating a larger governance problem.

OptionWhen it makes senseWhat to watch
Do nothingThe workflow is rare, low value, or already reliable.Competitors may improve speed, content depth, or service consistency first.
Run a small pilotThe task repeats often and has clear review criteria.Keep scope tight and measure the result against the current process.
Build a production workflowThe pilot is repeatable and risk controls are documented.Assign ownership, monitoring, training, and a rollback path.

AI Kick Start handover package for Best model for coding

A production handover should be concrete enough that another person can run it. For Best model for coding, that means a short brief, a workflow map, approved prompts or tool settings, source material, a review checklist, internal links to supporting resources, and a simple measurement sheet. This is the difference between reading about AI and turning it into operational capability.

That packaging also strengthens E-E-A-T. It shows experience through implementation notes, expertise through decision criteria, authoritativeness through source-aware structure, and trust through risks, controls, and review steps. The article becomes useful even if the reader never buys a tool because it helps them make a better operational decision.

Source trail

Primary references to keep this briefing grounded

AI and automation information changes quickly. Use these official or primary references to verify the claims, pricing, product behaviour, and compliance details before committing budget or production data.

Frequently asked questions

What is the practical takeaway from Best model for coding?

The June 2026 SWE-bench Pro coding leaderboard, with Claude Fable 5 in the lead and a full tier-by-tier breakdown for engineering teams. For AI Kick Start readers, the key is to translate the idea into one tool evaluation workflow with clear inputs, review points, and measurable outcomes. The article should be treated as implementation guidance, not a substitute for workflow design.

Who should use Best model for coding guidance in Model Review?

This guidance is most useful for Founders and operators who need to decide whether the topic changes tool selection, automation design, search visibility, data handling, training, or operational governance.

How should an Australian business implement Best model for coding?

Start small: compare the tool against one real task, check data handling, price the operating cost, and record the approval conditions. If the pilot improves time to value and adoption rate, document the pattern, link it to the relevant service or resource page, and then decide whether it belongs in a production workflow.

What to do next

  1. For Best model for coding, write down the single tool evaluation workflow this article should improve.
  2. Collect real examples, edge cases, and source material before testing Best model for coding with any AI output.
  3. Before implementing Best model for coding, add a human review checkpoint for quality, privacy, brand, or customer-impact risk.
  4. Measure time to value, adoption rate, cost per workflow for Best model for coding before deciding whether to scale.
  5. Connect Best model for coding to a related service, resource, or training path so readers have a clear next action.

Want help applying this? Explore the AI tools directory.

AI Kick Start is an Illawarra-based AI studio in Figtree, helping businesses across Wollongong, Shellharbour and Kiama and right across Australia put AI to work.

Explore with AI

Use the article as a decision prompt

Summarise this AI Kick Start article for an Australian business owner. Focus on the useful decision, the risks, and the first practical next step: Best model for coding: SWE-bench Pro leaderboard

Turn this into a practical roadmap.

Use the guide as a starting point, then map the first workflow worth building.

Book an AI strategy call