The AI Model Tier List, September 2026

AI Model Tier List
Last updated: September 1, 2026Hover a model, or tab to it, for the one-line reason it ranks where it does.- S+Daily driver
- Fable 5.1$20.00/M
- Opus 5$10.00/M
- AReach for it
- Grok 4.6$3.00/M
- gpt-5.6-sol$8.00/M
- glm-5.3$2.15/M
- kimi k3$6.00/M
- BGood, not first
- glm-5.3-flash$0.24/M
- Gemini 3.7 Flash$1.50/M
- gpt-5.6-luna$0.45/M
- deepseek v4 flash$0.66/M
- CSituational
- Muse Spark 1.2$2.00/M
- DNiche at best
- Composer 2.5$1.00/M
- gpt-5.6-terra$4.50/M
- Sonnet 5$4.00/M
- deepseek v4 pro$1.98/M
- FSkip it
- Gemini 3.1 Pro$4.50/M
Every price is the standard published rate from the vendor's own pricing page, retrieved September 1, 2026, blended at three parts input to one part output. Batch, cached and off-peak rates are excluded. One of these numbers is promotional, another is dated, and a third is brand new, which is its own section below.
S+: What Actually Earns The Top
Only three points separate these models on the one index that scores both (Artificial Analysis), but their blended prices are $10 apart. Here, unlike everywhere else on the list, the more expensive model ranks higher.
Fable 5.1
Released on September 1, 2026, it arrived on the date this list reflects and is both the highest-ranked and most expensive model here. Artificial Analysis scores it at 66, three points ahead of Opus 5 and four ahead of the Fable 5 model it replaces.
Its list price remains unchanged at $10 per million input tokens and $50 per million output tokens. The real price cut is in cache reads, which fell from $1 to $0.25 per million tokens. Anthropic now charges 0.025 times the base input rate, compared with 0.1 times for every other model on its pricing page. Because this tier list excludes cached rates, Anthropic's largest price reduction of the year doesn't appear on the price chip.
Opus 5
Opus 5 gives up only three points to Fable 5.1 while costing half as much, making it the strongest value this close to the top. Anthropic released it on July 24, 2026, and Artificial Analysis found that it delivered comparable intelligence to the previous Fable model at twenty-six percent lower cost per task.
The discount comes with a real tradeoff. In the same evaluation, Opus 5 scored lower than the Fable line on factual knowledge and hallucinated roughly half the time. For long, fact-sensitive work, that is reason enough to pay for Fable 5.1.
The reception in r/ClaudeCode is much colder than the benchmark. One active thread describes Opus 5 as difficult to steer and exhausting to supervise, with users objecting to its long, convoluted answers. These are user reports, not controlled evaluations, but the gap between the score and the day-to-day experience is too wide to ignore.
What the thread says: The complaint is less about benchmark quality than communication. Simple answers become long, convoluted explanations, even after requests to be concise.
A: Four Models, One Point Apart
Benchmarks barely separate these four models. Price does. Their blended rates range from $2.15 to $8 per million tokens.
Grok 4.6
The best price-to-score ratio anywhere above B, and the reason this row isn't just the leftovers. It matches gpt-5.6-sol's index score exactly at $3 against $8. It also has the flattest input-to-output ratio on the list, which makes it unusually cheap for long generation on top of that.
gpt-5.6-sol
Five points behind Fable 5.1, two behind Opus 5, and it would still have a case for S+ if Grok 4.6 didn't match it for a third of the money. It's the flagship of the GPT-5.6 family OpenAI shipped on July 9, 2026, which replaced the old unsuffixed naming with three tiers: Sol at the top, Terra in the middle, Luna at the bottom. Sol's own model page says it "roughly corresponds to the unsuffixed model tier used in earlier GPT-5 families", which is the clearest thing anyone has written about that naming scheme.
The $8 is promotional, and the section below says so, but the placement survives the expiry. The rate it was cut from was $5 and $30, which blends to $11.25. What it doesn't survive is a long document: Sol's window is 400K against the million that Fable 5.1, Opus 5 and kimi k3 all give you, and every GPT-5.6 model bills the whole request at double input above 272K tokens.
glm-5.3
glm-5.3 is the clear bargain in this row: it matches kimi k3's index score for roughly a third of the price. Z.ai released the model on August 18, 2026 and published its weights ten days later. Its licence is more restrictive than the plain MIT terms given to glm-5.3-flash, however. Model-as-a-service providers with more than ten billion dollars in revenue must first pass a Z.ai security review.
kimi k3
Three points off the top of the list at thirty percent of the price, and it keeps the full million-token window that Sol gives up. It always reasons, which is why its output rate is five times its input.
It's also, with glm-5.3 beside it, one of two models in this row you can download and run. Moonshot published open weights for a 2.8-trillion-parameter model, which is a strange sentence to be able to write, and it means the price on the chip is a convenience rate rather than the only way to get it.
Two of these four ship downloadable weights. That’s not a rounding error in a row this close to the top.
B: The Tier That Should Change How You Budget
This is the most interesting row on the list, because all four models in it cost less than two dollars per million tokens blended.
glm-5.3-flash
The cheapest thing on this list, at roughly half of Luna. Z.ai published its weights on August 26, 2026 under a plain MIT licence, eight days after the larger glm-5.3 and on much friendlier terms.
What keeps it out of the rows above is the word in its name. Flash is slow: around 42 output tokens a second, against a median near 65 for open-weight models its size, and verbose enough that a benchmark suite pulls about a third more output out of it than average. Cheap tokens, expensive minutes. If someone is waiting on it, the price stops being the number that matters.
Gemini 3.7 Flash
A capable budget option priced alongside Luna and DeepSeek Flash. Its input rate is dated, which is its own section below.
gpt-5.6-luna
The cheapest thing OpenAI sells, and until last week the cheapest thing here. Its price was cut by eighty percent on July 30, 2026, which means every comparison written before that date has it at five times what it now costs.
deepseek v4 flash
The third-cheapest blended price on the list, which makes it the DeepSeek sibling worth buying, and the spec sheet below explains exactly how much cheaper that sibling argument is.
The lesson of this row isn’t that cheap models are good. It’s that the gap between the cheap models and the expensive ones has closed enough that the expensive ones now have to justify a multiple, not a margin.
C: The Model That Lands Between The Rows
Muse Spark 1.2
Meta's, released August 5, 2026. It scores 57, which is the same number gpt-5.6-terra carries two rows below it and glm-5.3-flash carries one row above, and its price sits between the two of them. Landing exactly between the rows on both axes is what C is for.
D: Paying For The Wrong Thing
Four models, and between them they carry the two placements most likely to get argued with.
Composer 2.5
The cheapest model in this row by some way, which is a strange thing to find in D until you notice it isn't really a general model: Cursor's own launch post says it's built on the same open-source checkpoint as Composer 2, Moonshot's Kimi K2.5.
It's also the one model with no number in the index column below. Artificial Analysis rates it on their coding agent index instead, where it placed third. That's a real result, and it's measuring a different thing than every other row here.
gpt-5.6-terra
The middle tier of the GPT-5.6 family, and the score column is unkind to it. Two models match its 57 exactly: Muse Spark 1.2 for less than half the price, and glm-5.3-flash for a nineteenth of it.
Sonnet 5
In the same row, and its price has its own story. The $2 in and $10 out rate launched as introductory pricing due to expire on August 31, 2026 and rise to $3 and $15 on September 1. That increase was cancelled. Anthropic's pricing page now states the old rate is the standard one and the scheduled rise "will not occur." If you read almost any comparison written before this week, it has Sonnet 5 at $3 and $15.
deepseek v4 pro
It costs exactly three times its own sibling: DeepSeek v4 Flash is $0.44 in and $1.32 out. Three times the input, three times the output, from the same vendor, released two weeks apart. The index column says what the premium buys: 53 against 52. One point.
Ranking the Pro two rows below the Flash is still the boldest call on this list, and it’s the easiest one to check, because both models are from the same company and you can run them against each other in an afternoon. The index column is the first outside evidence I have found that agrees with it.
F: One Model I Can’t Justify
Gemini 3.1 Pro
The oldest model here by a distance. It was released on February 19, 2026, nearly three months before anything else on this list, and it still prices identically to gpt-5.6-terra. A model holding its launch price while everything around it gets cheaper is a statement about where its vendor's attention has moved.
It also carries the lowest score on the list. Every single model above it is cheaper, better, or both. That's what F is for.
A couple gotchas
Four numbers on this list aren’t what they look like, and all four would be wrong in a post written a month from now.
gpt-5.6-sol’s $4 and $20 is promotional. OpenAI’s own changelog says, in the same entry that announces the cut, that “GPT-5.6 Sol’s promotional pricing is available at least through November 21, 2026.” Every comparison I found quoted the first half of that sentence and dropped the second. Treat $8 blended as a floor with an expiry date on it. The rate it was cut from was $5 and $30, which blends to $11.25, and OpenAI hasn’t said whether that’s where it goes back to.
Gemini 3.7 Flash’s $0.75 input is dated. Google’s pricing page gives it as $0.75 through December 31, 2026 and $1.50 starting January 1, 2027. The blended price on the chip doubles on New Year’s Day.
glm-5.3-flash’s $0.24 is six days old, and for the six days before that it was zero. It spent its first week on OpenRouter unnamed and free, listed under a “Stealth” provider label as ox-alpha, and reached the top of that site’s usage rankings before anyone knew which lab had built it. Z.ai claimed it on August 26, 2026. It’s also the only number on this tier list that is currently too high: a launch promotion halves the list rate until September 9, 2026. Every leaderboard placement this model has was earned while it cost nothing and carried no brand, which is either the cleanest blind test anyone has run on a frontier model or the best launch marketing of the year.
Every GPT-5.6 model has a long-context surcharge. Prompts above 272K input tokens are billed at double input and one and a half times output for the entire request. If your workload is long documents, the three OpenAI numbers on this list are the wrong ones for you.
| Model | Input | Output | Blended | Context | Released | Index |
|---|---|---|---|---|---|---|
| Fable 5.1 | 10.00 | 50.00 | 20.00 | 1M | Sep 1, 2026 | 66 |
| Opus 5 | 5.00 | 25.00 | 10.00 | 1M | Jul 24, 2026 | 63 |
| Grok 4.6 | 2.00 | 6.00 | 3.00 | 500K | Aug 12, 2026 | 61 |
| gpt-5.6-sol | 4.00 | 20.00 | 8.00 | 400K | Jul 9, 2026 | 61 |
| glm-5.3 | 1.40 | 4.40 | 2.15 | 200K | Aug 18, 2026 | 60 |
| kimi k3 | 3.00 | 15.00 | 6.00 | 1M | Jul 27, 2026 | 60 |
| glm-5.3-flash | 0.15 | 0.50 | 0.24 | 1M | Aug 26, 2026 | 57 |
| Gemini 3.7 Flash | 0.75 | 3.75 | 1.50 | 1M | Aug 13, 2026 | 56 |
| gpt-5.6-luna | 0.20 | 1.20 | 0.45 | 400K | Jul 9, 2026 | 52 |
| deepseek v4 flash | 0.44 | 1.32 | 0.66 | 1M | Jul 31, 2026 | 52 |
| Muse Spark 1.2 | 1.25 | 4.25 | 2.00 | 512K | Aug 5, 2026 | 57 |
| Composer 2.5 | 0.50 | 2.50 | 1.00 | 200K | May 18, 2026 | — |
| gpt-5.6-terra | 2.00 | 12.00 | 4.50 | 400K | Jul 9, 2026 | 57 |
| Sonnet 5 | 2.00 | 10.00 | 4.00 | 1M | Jun 30, 2026 | 55 |
| deepseek v4 pro | 1.32 | 3.96 | 1.98 | 1M | Aug 13, 2026 | 53 |
| Gemini 3.1 Pro | 2.00 | 12.00 | 4.50 | 1M | Feb 19, 2026 | 48 |
Prices retrieved September 1, 2026 from each vendor's own pricing documentation. Batch, cached-read and off-peak rates are excluded, as are the promotional and long-context adjustments described above. The index column is Artificial Analysis' Intelligence Index v4.1.1, read the same day, at each model's best published configuration. That scale scores configurations rather than models: Opus 5 is 63 at maximum effort and 52 at minimum. So the highest scores in the column were set at settings that burn far more tokens than the blended price beside them assumes.