The $1-per-million club: Cheapest capable models
Analysis
A year ago, running a million tokens through a decent model cost real money. Now the floor has dropped so far that "good enough" has stopped being expensive. That's the story underneath all the numbers: the cheap end of the market has caught up to where the frontier sat not long ago, and for a lot of everyday business work, you no longer need to pay flagship prices.
That's the genuine shift. The catch is the specifics. A list has been doing the rounds claiming exactly four models now sit under $1 per million input tokens, each with neat benchmark scores to match. When we went looking, most of those figures didn't line up with what the vendors actually publish. Some of the models don't appear to exist under the names given. So read what follows as a map of how to weigh cheap models against each other, not as a verified shopping list.
Here's why it matters for your team: if you're doing bulk document work, monitoring, classification, or first-draft generation, the cost difference between the cheap tier and a flagship model is large enough to change what's worth automating at all. The trick is matching the model to the job and confirming today's price yourself.
The contenders
| Model | Input Price | Output Price | SWE-bench Pro | MMLU | Context |
|---|---|---|---|---|---|
| DeepSeek V3.5 | $0.15 / 1M | $0.60 / 1M | 52.4% | 85.8% | 1M |
| Gemini 3.5 Flash | $0.35 / 1M | $0.70 / 1M | 48.2% | 86.8% | 1M |
| Qwen 3 | $0.40 / 1M | $1.20 / 1M | 46.2% | 84.6% | 128K |
| GPT-5.5 Instant | $0.50 / 1M | $1.50 / 1M | 42.1% | 84.2% | 128K |
A caution before you read the table as gospel: we couldn't confirm most of these figures, and a few clash with what the vendors publish. There's no "DeepSeek V3.5" in DeepSeek's own change log (opens in a new tab), the documented line runs V3, V3.1, V3.2 and V4, with V3.2 priced nearer $0.28 input / $0.42 output on a roughly 131K context. Gemini 3.5 Flash is real, but reported I/O 2026 pricing (opens in a new tab) lands closer to $1.50 input, which would put it above the sub-$1 line, not under it. And while GPT-5.5 shipped in April 2026 (opens in a new tab), its standard API pricing is reported around $5 / $30 per million on a 1M context, nowhere near $0.50 / $1.50. The SWE-bench Pro scores below are also unconfirmed and sit lower than the public leaderboard (opens in a new tab) numbers we'd expect for named flagships. So take the ranking as reasoning about value, and verify the live numbers before you spend anything.
Ranking by value
1. DeepSeek V3.5, Best overall value. On the figures quoted ($0.15/$0.60, 1M context, 52.4% SWE-bench Pro, 85.8% MMLU), this would be the cheapest capable option with the longest context, and the open licence sweetens it further. Worth flagging: a model under exactly that name and price doesn't appear in DeepSeek's docs, so the real-world equivalent is more likely V3.2 or V4. Either way, for bulk document processing, monitoring and analysis, DeepSeek's cheap tier is hard to beat on cost per token.
2. Gemini 3.5 Flash, Best balance. The pitch is a small premium over DeepSeek in exchange for a higher MMLU (86.8%) and Google's production reliability. The $0.35/$0.70 pricing quoted here is the part to double-check, reported figures are several times higher, which would knock it out of the sub-$1 club entirely. If the cheap price holds where you are, Flash is a sensible default for production. If it doesn't, the reliability argument still stands, just at a higher cost.
3. Qwen 3, Best for multilingual. At a quoted $0.40/$1.20, Qwen 3 reads as the cheapest route for Asian-language work. The exact SKU and the 128K context are unconfirmed, the current Qwen line tends to ship with much larger windows, but the multilingual strength is the real draw here. If your workload leans into non-English content, this is the one to trial.
4. GPT-5.5 Instant, Best ecosystem integration. Instant is pitched as the priciest and weakest of the four on paper, earning its place through tight integration with OpenAI's platform. The $0.50/$1.50, 128K-context SKU quoted here isn't one we could verify against OpenAI's published pricing, which runs far higher. If you're already inside OpenAI's tooling, it's the path of least friction, just don't assume the cheap price tag.
The price-performance curve
The headline point survives the messy details: the gap between cheap and capable has narrowed sharply. The quoted scores put all four above 42% on SWE-bench Pro and above 84% on MMLU. The MMLU range is broadly believable for capable 2026 models; the SWE-bench numbers we couldn't confirm. The MMLU framing is the safer one to lean on, that level of general knowledge would have read as frontier performance a couple of years back, and now it's showing up at budget prices. The cheapening of AI is real even if these particular figures aren't nailed down.
Verdict
If you want maximum value, start by trialling DeepSeek's cheapest current model, but check which SKU is actually live, since "V3.5" may not be it. Step up to Gemini 3.5 Flash if you need Google's infrastructure or a bit more general knowledge, and confirm the price first, because it may sit above the sub-$1 line. Reach for Qwen 3 for multilingual work, and pick GPT-5.5 Instant only if you're already committed to OpenAI's stack. Across all of them, the rule is the same: verify today's pricing on the vendor's own page before you build anything on top of it.
Best value: DeepSeek V3.5
The $1-per-million club: answer-first summary
The $1-per-million club matters because it can change how Founders and operators plan, build, or govern an tool evaluation workflow. Which models do real work under $1 per million input tokens?
The direct answer is this: do not treat the topic as a standalone trend. Treat it as a decision about inputs, outputs, review ownership, data exposure, and whether the workflow produces a result that is faster, safer, or more useful than the current process.
The $1-per-million club: implementation checklist
- Define the user, job to be done, and success metric for the tool evaluation workflow.
- Collect real examples, policies, source files, customer questions, or search queries before writing prompts or choosing tools.
- Separate low-risk drafts from decisions that need approval, privacy checks, or senior review.
- Document what the AI is allowed to access, what it must not access, and who signs off before production use.
- Review time to value, adoption rate, cost per workflow, quality review score after a small pilot rather than judging the idea from a demo.
This keeps the work practical. It also gives search engines and AI answer engines a clean factual structure: what the topic is, who it helps, what to do next, and which risks matter before implementation.
Decision criteria for The $1-per-million club
| Decision area | What to check | Production signal |
|---|---|---|
| Intent | Does The $1-per-million club solve a real workflow problem? | The use case has a named owner and measurable outcome. |
| Data | Can the required data be used safely? | Sensitive data is classified and access is controlled. |
| Quality | Can a reviewer judge the output consistently? | Examples, rubrics, or acceptance criteria exist. |
| Scale | Can the workflow be repeated without hero effort? | The process is documented and can be handed to another team member. |
Practical example for The $1-per-million club
A small business could use this article to choose one practical test. For example, a manager might take one customer-facing process, one internal document workflow, or one recurring content task and redesign only that step with AI support. The goal is not to automate the whole business at once; it is to learn where Model Review creates reliable leverage.
The useful deliverable is a short operating note: the trigger, the source material, the prompt or tool, the review checklist, the escalation rule, and the metric. That note becomes the handover asset for staff training, SEO/GEO content, service delivery, or future agent work.
Risks and controls for The $1-per-million club
The common failure pattern is moving too quickly from a promising idea into an unmanaged workflow. For The $1-per-million club, the risk is not only bad output. It can also be unclear data permission, staff confusion, duplicate content, unreviewed customer advice, or a tool that quietly changes cost or capability.
- Control tool sprawl with a named owner, a review step, and written acceptance criteria.
- Control unclear pricing with a named owner, a review step, and written acceptance criteria.
- Control vendor lock-in with a named owner, a review step, and written acceptance criteria.
- Control unreviewed data sharing with a named owner, a review step, and written acceptance criteria.
Measurement plan for The $1-per-million club
A useful AI or SEO initiative should leave evidence. Track time to value, adoption rate, cost per workflow, quality review score and compare the pilot against the current process. If the measure does not improve, keep the learning but avoid scaling the workflow.
For GEO readiness, the page should also answer the core question directly, define the entities involved, include implementation steps, explain tradeoffs, and link readers to the next relevant AI Kick Start service, guide, tool, or article.
Definitions and entities for The $1-per-million club
For search, GEO, and staff handover, define the core entities in plain language. In this article the important entities are the workflow owner, the AI tool or model, the source material, the review process, the risk boundary, and the measurable business outcome. Clear definitions make the page easier for people to scan and easier for AI answer engines to quote accurately.
- Workflow owner: the person accountable for deciding whether The $1-per-million club belongs in the business process.
- Source material: the documents, examples, policies, URLs, prompts, videos, or customer questions that ground the output.
- Review boundary: the point where a human checks accuracy, privacy, brand voice, or customer impact before the result is used.
- Success metric: the measure that proves whether the tool evaluation workflow is worth repeating.
The $1-per-million club versus doing nothing
Doing nothing is also a decision. The cost may be slow manual work, weaker search visibility, inconsistent advice, duplicated effort, or staff using unmanaged AI tools without a shared process. The practical question is whether a controlled pilot can reduce that cost without creating a larger governance problem.
| Option | When it makes sense | What to watch |
|---|---|---|
| Do nothing | The workflow is rare, low value, or already reliable. | Competitors may improve speed, content depth, or service consistency first. |
| Run a small pilot | The task repeats often and has clear review criteria. | Keep scope tight and measure the result against the current process. |
| Build a production workflow | The pilot is repeatable and risk controls are documented. | Assign ownership, monitoring, training, and a rollback path. |
AI Kick Start handover package for The $1-per-million club
A production handover should be concrete enough that another person can run it. For The $1-per-million club, that means a short brief, a workflow map, approved prompts or tool settings, source material, a review checklist, internal links to supporting resources, and a simple measurement sheet. This is the difference between reading about AI and turning it into operational capability.
That packaging also strengthens E-E-A-T. It shows experience through implementation notes, expertise through decision criteria, authoritativeness through source-aware structure, and trust through risks, controls, and review steps. The article becomes useful even if the reader never buys a tool because it helps them make a better operational decision.





