How we tested
We used the consumer web interface for each model, default settings, fresh sessions, and identical prompts. No regenerations, no follow-up prompts, no cherry-picking the best of five attempts. Output was evaluated on tone, structure, factual accuracy, and how much human editing would be required before sending.
ChatGPT | Claude | Gemini | |
|---|---|---|---|
Model tested | GPT-6 Sol | Claude Opus 5.5 | Gemini 3.8 Flash |
Interface | Web, default | Web, default | Web, default |
Pricing (consumer) | $20/month (Plus) | $20/month (Pro) | $19.99/month (AI Pro) |
API price (1M tokens) | $2 input / $10 output | $4 input / $20 output | $0.75 input / $3.75 output |
Email test: canceling a phone contract

The prompt: Write a firm, polite email to a customer service agent asking them to terminate a mobile contract. Do not leave room for counteroffers. Do not use numbered lists, bolded headers, or sterile corporate greetings. Sound like a tired human who wants this to be over.
This is one of the hardest everyday writing tasks because it requires three things at once: firmness, courtesy, and the absence of any formatting that signals “AI wrote this.”
Claude Opus 5.5: the natural winner
Claude produced a single cohesive block of text that read like something a real person would type after a long day. It stated the cancellation request directly, gave a clear reason without over-explaining, and closed without begging for retention. No bullet points. No bolded headers. No “I hope this email finds you well.”
Anthropic specifically rebuilt Opus 5.5 to reduce the dense, over-structured writing style that plagued Opus 5. The company cut em dash usage by 95%, shortened average sentence length from 12.14 words to 10.03, and trained the model to “put the key point first” and “use less jargon”. The email test confirmed those improvements translated into natural-sounding prose.
The tradeoff: Claude’s average response length was 481 words—the longest of any Opus model tested. You may need to specify a word limit or you will get more context than you asked for.
GPT-6 Sol: structured and reliable, but stiff
GPT-6 Sol delivered a clean, send-ready draft on the first attempt. Every element was present: subject line, greeting, context, request, closing. The tone was professional and polite. The problem was that it read like a template. One reviewer described ChatGPT’s email output as “correct but boring”—a safe template that would work in almost any situation but had “little personality, conviction, or fresh note”.
GPT-6 Sol improved on earlier ChatGPT models by cutting fluff and producing shorter, more direct answers. But in the cancellation email test, that directness still came wrapped in a layer of corporate politeness that Claude avoided.
The tradeoff: GPT-6 Sol is the fastest and most reliable model for volume email work. If you send 30 routine emails a day, its predictability is an asset. If you send five emails that matter, its stiffness is a liability.
Gemini 3.8 Flash: practical tips, weak tone
Gemini produced a competent email but finished a clear third. Its output had what one reviewer called “a faint corporate stiffness” and lacked the conversational subtlety the task required.
Gemini’s strength was not the email itself but the extras: it appended practical tips for personalization and follow-up that the other two models did not offer. If you want a model that coaches you on how to improve the email rather than just writing it, Gemini has value. If you just want the email, it is the weakest of the three.
The tradeoff: Gemini 3.8 Flash costs roughly one-fifth of Claude Opus 5.5 per task and runs at ~300 output tokens per second. For routine correspondence where “good enough” is genuinely good enough, the cost savings are real.
Research test: a multi-source briefing on a regulatory change
The prompt: Using the three attached source documents, produce a 400-word briefing on how the proposed SEC climate disclosure rule would affect mid-sized public companies. Include the compliance timeline, the three highest-cost requirements, and one paragraph on what remains uncertain. Cite the specific document and page for each factual claim.
This task tests long-document comprehension, multi-source synthesis, factual accuracy, and instruction-following under constraints.
GPT-6 Sol: the research winner
GPT-6 Sol produced the most usable briefing. It correctly identified the compliance timeline from Document 1, pulled the three costliest requirements from Document 2, and flagged the open questions from Document 3. Every factual claim was cited. The structure was clean and the prose was dense but readable.
This aligns with GPT-6 Sol’s design priorities: OpenAI explicitly optimized Sol for “professional work, coding, automation, and computer-use tasks” rather than pushing the reasoning frontier. On Terminal-Bench Science, GPT-6 Astra scored 64.6% while Sol scored 22.4%—Astra is the science model, Sol is the workhorse. But for a research task that requires reading, extracting, and organizing—not novel scientific reasoning—Sol’s strengths were the right fit.
The tradeoff: Sol’s factual density means the output reads like a professional memo, not an explainer. If your audience is non-specialist, you will need to rewrite.
Claude Opus 5.5: strong analysis, longer output
Claude produced a briefing that was analytically sharper than GPT-6 Sol’s. It caught a nuance in Document 3 that Sol missed—a distinction between “proposed” and “adopted” requirements that changed the compliance timeline for one category of company. The prose was natural and the citations were accurate.
The problem was length. Claude’s briefing ran to roughly 700 words against Sol’s 420. The extra words were not padding; they were additional analysis. But the prompt specified 400 words, and Claude overshot it by 75%. If you need a model that follows length constraints precisely, Claude is not the best choice.
The tradeoff: Claude Opus 5.5 leads the Artificial Analysis Intelligence Index at 58, the highest score ever recorded on that platform at launch. That intelligence shows in the quality of its analysis. It does not show in its willingness to be brief.
Gemini 3.8 Flash: fast and cheap, but shallow
Gemini produced a briefing that was accurate on the surface but missed the synthesis. It extracted facts from each document correctly but did not connect them. The section on “what remains uncertain” was a list of open questions copied from Document 3 rather than an analysis of what those questions mean for the reader.
Gemini 3.8 Flash scores 59 on the Artificial Analysis Intelligence Index—on par with GPT-5.6 Sol at sub-maximum reasoning. Its weakness is not intelligence but depth. It is optimized for speed and cost, and at $0.58 per intelligence task, it is the cheapest model at its intelligence level. For research tasks that require genuine synthesis rather than extraction, that optimization shows.
Head-to-head results
Dimension | ChatGPT (GPT-6 Sol) | Claude (Opus 5.5) | Gemini (3.8 Flash) |
|---|---|---|---|
Email naturalness | 3rd | 1st | 2nd |
Email send-readiness | 1st | 2nd | 3rd |
Research synthesis | 1st | 2nd | 3rd |
Research brevity | 1st | 3rd | 2nd |
Cost per task | Mid | High | Low |
Speed | Fast | Slower | Fastest |
Best single use case | Structured research and professional work | Writing that sounds human | High-volume, low-cost tasks |
Price, speed, and limits
GPT-6 Sol | Claude Opus 5.5 | Gemini 3.8 Flash | |
|---|---|---|---|
Consumer plan | $20/month | $20/month | $19.99/month |
API input (1M tokens) | $2 | $4 | $0.75 |
API output (1M tokens) | $10 | $20 | $3.75 |
Context window | 1.05M tokens | 1M tokens | 1M tokens |
Output speed | Fast | 74–86 tokens/sec | ~300 tokens/sec |
Cost per AA Intelligence task | ~$1.20 | ~$2.80 (high effort) | ~$0.58 |
Claude Opus 5.5 costs roughly 2.3 times more per intelligence task than GPT-6 Sol and nearly 5 times more than Gemini 3.8 Flash. The question is whether the quality difference justifies the premium. For email writing that goes to clients, yes. For high-volume research extraction, no.
FAQ
Which model is best for writing emails that sound human?
Claude Opus 5.5. In multiple controlled tests, Claude produced the email that sounded the least AI-generated. It was the only model that consistently avoided corporate clichés and formatting that signals “AI wrote this”.
Which model is best for research tasks with multiple source documents?
GPT-6 Sol. It follows length constraints more precisely, cites sources accurately, and produces structured output that is immediately usable. Claude’s analysis is sharper but longer; Gemini’s extraction is accurate but shallow.
Is Gemini good enough for everyday work?
For routine email and basic research extraction, yes. Gemini 3.8 Flash is the cheapest model at its intelligence level and the fastest of the three. If your work involves high volume and low stakes, Gemini is the rational choice.
Should I use one model or two?
Use two. Start with Claude for anything you write that a human will read carefully—client emails, proposals, articles. Use GPT-6 Sol for research briefs, structured plans, and high-volume drafting. Use Gemini when cost and speed matter more than polish. The $40/month for Claude Pro plus ChatGPT Plus is worth it if writing and research are core to your work.
Do these models work with free plans?
The tests above used paid consumer plans. Free-tier models from Meta, Perplexity, Copilot, and Qwen were “competent but forgettable” in a separate 10-model email test—none broke the brief, but none produced copy worth shipping without a rewrite. If your output matters, pay for the model.
What about context windows?
All three models tested offer roughly 1 million tokens of context—enough to process a 200-page document in a single session. Context window size is no longer a differentiator at the premium tier. What matters now is what the model does with the context you give it.
Letters
No letters yet — be the first to write.