Price, Speed & Limits

This Week's AI Model Scorecard: Winners for Writing, Coding, Research, and Cost

This Week's AI Model Scorecard: Winners for Writing, Coding, Research, and Cost
Four models split the week's wins: Claude Opus 5.5 tops LMArena writing at 1,509 Elo while cutting em dashes 95%, and the cheaper Sonnet 5.5 ($2/$10 per million tokens) beat Opus on Terminal-Bench 4.0 at 70.6%.

This was the most consequential week for AI model releases since the GPT-4 era. Anthropic shipped Claude Opus 5.5 and Sonnet 5.5 within six days. OpenAI cut GPT-6 Sol and Luna prices in half. Google expanded Gemini 3.8 across voice and enterprise. No single model won the week. Four different models won four different categories. Here is where each one stands and what it means for your work.

Week in review: what shipped

September 22: Anthropic released Claude Opus 5.5. OpenAI released GPT-6 Sol and Luna roughly 90 minutes later.

September 28: Anthropic released Claude Sonnet 5.5 at $2/$10 per million tokens — half the price of Opus 5.5.

September 29: OpenAI shipped GPT-6.1 Sol, a mid-cycle refresh that scored within 1 point of GPT-6 Astra at one-fifth the task cost.

Ongoing: Google's Gemini 3.8 Flash and Live models, launched earlier in September, gained enterprise availability for Live Avatar across 97 languages.

The result: the frontier moved in four directions at once — higher intelligence, lower cost, faster speed, and broader modality.

Writing: Claude Opus 5.5

The winner: Claude Opus 5.5
Runner-up: Claude Sonnet 5.5
Surprise contender: Gemini 3.8 Flash on LMArena

Anthropic's Opus 5.5 leads LMArena's blind text vote at 1,509 Elo, ahead of Gemini 3.8 Flash at 1,492 and GPT-6 Astra at 1,478. In a category decided by human preference, that gap is significant.

But the more interesting writing story this week is what Anthropic removed. Opus 5.5 cut em dash usage by 95% compared to Opus 5 — from 15.2 per 1,000 words down to 0.8. It also reduced semicolon usage by 73%. These are the structural tells that make AI writing sound like AI writing. Anthropic rebuilt the model to sound less like a machine and more like a person who happens to be very well-informed.

Sonnet 5.5 inherits much of this. At $2/$10 per million tokens, it offers near-Opus writing quality at half the price. For anyone whose writing task is well-scoped — an email, a summary, a client update — Sonnet 5.5 may be the smarter buy.

The tradeoff: Opus 5.5 remains the best model for open-ended, voice-sensitive writing. Sonnet 5.5 is the best value. Neither has a clear competitor from OpenAI or Google on tone and naturalness.

Best for: Anyone whose writing will be read by a human who notices when something sounds off.

Coding: Claude Sonnet 5.5 (with a caveat)

The winner: Claude Sonnet 5.5
Runner-up: Claude Opus 5.5
Best value: Claude Sonnet 5.5 at high effort

This is the most surprising result of the week. Claude Sonnet 5.5 — a mid-tier model priced at half of Opus 5.5 — beat its more expensive sibling on Terminal-Bench 4.0, an agentic coding benchmark that tests whether a model can complete real professional tasks by typing commands on its own.

Benchmark

Sonnet 5.5

Opus 5.5

Sonnet 5

Terminal-Bench 4.0

70.6%

66.4%

10.3%

CursorBench 4.0

55.5%

57.8%

34.1%

GDPval-AA v2.1

1,844

1,846

1,449

Source: Anthropic's official evaluations

Artificial Analysis independently confirmed the finding: Sonnet 5.5 scored 63.6% on its Terminal-Bench run versus 59.6% for Opus 5.5 and 59.1% for GPT-6 Astra.

The cost story is where it gets complicated. Anthropic says Sonnet 5.5 at High effort matches GPT-6 Sol on FrontierCode for about one-fifth the cost per task. But at maximum effort, Sonnet 5.5 generates approximately 193,000 output tokens per task — the highest Artificial Analysis has ever measured, roughly 60% more than Opus 5.5. That pushed its cost per Intelligence Index task to $7.60, about 50% higher than Sonnet 5's $5.09.

The practical takeaway: Sonnet 5.5 is the best coding model for well-scoped tasks at default or high effort. At maximum effort on complex tasks, Opus 5.5 remains more cost-efficient. Anthropic's migration guidance recommends re-running your effort sweep rather than carrying over settings from Sonnet 5.

Best for: Developers who want near-frontier agentic coding at half the per-token price, and who are willing to tune effort settings to their workload.

Research: Claude Opus 5.5 and GPT-6 Astra split the crown

Best for knowledge work: Claude Opus 5.5
Best for pure reasoning: GPT-6 Astra
Best open-weight alternative: Kimi K3

Research splits into two subcategories this week, and each has a different winner.

Knowledge work and factual accuracy: Claude Opus 5.5 leads the Artificial Analysis Intelligence Index at 57.6, ahead of GPT-6 Astra at 52.7 and Gemini 3.8 Flash at 40.9. It also leads the AA-Omniscience factual accuracy benchmark at 46.4%, followed by Claude Fable 5.1 at 43.5% and GPT-6 Astra at 43.4%.

Pure reasoning and hard math: GPT-6 Astra holds the edge. On GPQA Diamond, a graduate-level reasoning benchmark, GPT-6 Astra scores 96% against Gemini 3.8 Flash's 95.3%. On ARC-AGI-3, Astra solved 62.7% of puzzles through the standard harness — the highest published score among tested models.

Humanity's Last Exam: Claude Fable 5.1 leads at 65%, followed by Claude Opus 5 at 64.7% and Claude Mythos 5 at 64.5%. Anthropic holds the top three positions on this benchmark.

Open-weight: Kimi K3 scores 52.5 composite — a 7.2-point gap behind Claude Opus 5.5's 59.7, but practically bridgeable for many workloads. GLM-5.3, Qwen3.8 Max, and DeepSeek V4.1-Flash all score above 51.

The tradeoff: Opus 5.5 is better at research tasks that require synthesizing facts across sources and producing a usable deliverable. Astra is better at tasks that require novel reasoning where no precedent exists. Choose based on whether your research is retrieval-heavy or reasoning-heavy.

Best for: Opus 5.5 for literature reviews, financial analysis, and document synthesis. Astra for mathematical research, novel problem-solving, and scientific hypothesis generation.

Cost: GPT-6 Luna (and a price warning)

Cheapest overall: GPT-6 Luna at $0.10/$0.50 per million tokens
Best cost-per-intelligence: Gemini 3.8 Flash at $0.75/$3.75
Best cost-adjusted coding value: Ministral 3 3B at $0.10 per million tokens

OpenAI's price cuts this week were structural, not promotional. GPT-6 Sol dropped from $4/$20 to $2/$10. GPT-6 Luna dropped from $0.20/$1.20 to $0.10/$0.50. Both are permanent.

At $0.10 input / $0.50 output, GPT-6 Luna is the cheapest model from a major provider that can still handle professional-grade text work. For high-volume extraction, classification, and summarization, Luna is the rational default.

But the headline cost story this week is not just about low prices. It is about a price increase coming in three months.

Gemini 3.8 Flash's introductory pricing — $0.75 input / $3.75 output — expires January 1, 2027. Starting then, rates double to $1.50 and $7.50. If you are planning a high-volume deployment on Gemini 3.8 Flash, the cost advantage disappears in Q1 2027. Build your budget accordingly.

Cost per Intelligence Index task:

Model

Price (in/out per 1M)

Cost per task

Claude Opus 5.5

$4 / $20

~$2.80 (high effort)

GPT-6 Sol

$2 / $10

~$1.20

Gemini 3.8 Flash

$0.75 / $3.75

~$0.58

GPT-6 Luna

$0.10 / $0.50

~$0.15

Sources: Artificial Analysis, OpenAI, Google

Small business owner calculating AI API costs in supermarket office.

The tradeoff: The cheapest model is rarely the best model for complex work. But for the 70% of professional tasks that are extraction, classification, and routine generation, Luna and Flash are more than sufficient.

Best for: GPT-6 Luna for high-volume, low-stakes text processing. Gemini 3.8 Flash for multimodal tasks at scale, before the January price increase.

The scorecard

Category

Winner

Runner-up

Key metric

Price

Writing

Claude Opus 5.5

Claude Sonnet 5.5

LMArena Elo: 1,509

$4 / $20

Coding

Claude Sonnet 5.5

Claude Opus 5.5

Terminal-Bench 4.0: 70.6%

$2 / $10

Research (knowledge)

Claude Opus 5.5

Claude Fable 5.1

AA Intelligence Index: 57.6

$4 / $20

Research (reasoning)

GPT-6 Astra

Claude Opus 5.5

GPQA Diamond: 96%

$10 / $50

Cost (overall)

GPT-6 Luna

Gemini 3.8 Flash

$0.10 in / $0.50 out

$0.10 / $0.50

Cost (multimodal)

Gemini 3.8 Flash

GPT-6 Luna

$0.75 in / $3.75 out

$0.75 / $3.75

Sources: Artificial Analysis, LMArena, Anthropic, OpenAI, Google

What this means for model selection

The cost-performance frontier moved twice this week. GPT-6.1 Sol scored within 1 point of GPT-6 Astra on the Intelligence Index at one-fifth the task cost. Claude Sonnet 5.5 beat Opus 5.5 on agentic coding at half the per-token price. The gap between "frontier" and "good enough" has narrowed to the point where the premium model is often the wrong default.

Three of four category winners are Anthropic models. Claude Opus 5.5 and Sonnet 5.5 took writing, coding, and knowledge-work research. GPT-6 Astra took pure reasoning. Google took cost and multimodal. The week belonged to Anthropic on task-specific performance.

The cheapest model got cheaper, and the expensive model got more expensive. GPT-6 Luna at $0.10/$0.50 is now cheap enough to run extraction and classification at scales that were cost-prohibitive six months ago. Claude Sonnet 5.5 at max effort costs $7.60 per task — 50% more than its predecessor. Cost optimization is no longer just about picking the cheapest model. It is about picking the right effort level for each task.

The January price cliff is the most important planning item. Gemini 3.8 Flash's cost advantage over GPT-6 Sol and Claude Sonnet 5.5 disappears on January 1, 2027. If you are building a deployment on Flash's current pricing, model the doubled rate into your Q1 budget.

Model routing is now the highest-leverage optimization. The cost-accuracy study data shows that cherry-picking the best model per task improves average accuracy by 11 percentage points over using any single model. With four models winning four different categories this week, routing is no longer optional. It is the default strategy.

FAQ

Which AI model is the overall winner this week?

There is no overall winner. Claude Opus 5.5 leads the Artificial Analysis Intelligence Index at 57.6. Claude Sonnet 5.5 won agentic coding. GPT-6 Astra won pure reasoning. GPT-6 Luna won cost. Four models, four categories, four different answers to "which is best."

Is Claude Sonnet 5.5 actually better than Opus 5.5 for coding?

On Terminal-Bench 4.0, yes — 70.6% versus 66.4% on Anthropic's evaluations, and 63.6% versus 59.6% on Artificial Analysis's independent run. But at maximum effort, Sonnet 5.5 costs more per task than Opus 5.5 because it generates far more tokens. It wins on capability at the same effort level. It loses on cost-efficiency at the highest effort settings.

Should I switch from GPT-6 to Claude Opus 5.5?

If your work is writing-heavy or knowledge-work research, yes. Opus 5.5 leads on both, and it costs less per token than GPT-6 Astra ($4/$20 vs $10/$50). If your work is math-heavy or requires computer-use capabilities, Astra remains the better pick.

Is GPT-6 Luna good enough for professional work?

For extraction, classification, summarization, and routine generation, yes. At $0.10/$0.50 per million tokens, it is the cheapest professional-grade text model available. For writing that will be read by clients or published, use a model with better prose quality.

What happens to Gemini 3.8 Flash pricing in January?

The introductory rate of $0.75 input / $3.75 output doubles to $1.50 / $7.50 on January 1, 2027. Google has not announced whether it will extend the promotional pricing.

What should I watch next week?

Claude Haiku 5.5 is expected to join the Claude 5.5 family "in the coming weeks," which would bring the generation's improvements to the lowest price tier. Google's Gemini 4 Pro is in late-stage testing under a disguised benchmark name, with leaked scores that would place it above GPT-6 Astra and Claude Fable 5.1 on agentic tasks. Neither has an official release date.

Last updated · 2026-10-01 23:20

Letters

No letters yet — be the first to write.

Leave a letter
© 2026 modelmatchdesk.com. All rights reserved. — grown slowly, toward the light —