The Best AI for Writing in 2026

Coinmama
Ledger


The strangest thing about the current generation of AI models is that they have become dramatically more capable without necessarily becoming dramatically better writers.

That sounds contradictory until you spend much time with them. In 2026, frontier models can solve difficult mathematics problems, debug complex software, operate computers, search the web, reason across huge collections of documents and carry out multi-stage tasks that would have looked implausible only a couple of years ago. The improvement in coding and reasoning, in particular, has been astonishing.

Ask the same models to produce a genuinely memorable 1,500-word magazine article, though, and the progress is harder to see. The writing is usually competent, often very good and occasionally excellent, but the familiar habits remain: over-explaining, generic transitions, unnecessary summaries and a tendency to settle into a polished but unmistakably synthetic register.

That makes the current AI writing race more interesting than the old question of which chatbot produces the prettiest paragraph.

bybit

OpenAI now has GPT-5.6 Sol. Anthropic has Claude Opus 5. Google has Gemini 3.1 Pro Preview alongside its rapidly expanding Gemini 3.x family. On August 12, xAI added Grok 4.6 to the field.

All four are extremely capable systems, but the old shorthand — Claude for prose, ChatGPT for everything else, Gemini for Google users — no longer describes the market particularly well. Most of all, the idea that Claude should automatically be crowned the best writer now deserves much more scrutiny than it used to.

Writing quality is proving to be one of the more subjective and strangely nonlinear areas of AI progress, and that may tell us something important about how machine intelligence is developing.

The four AI writing empires

Here is the Gilded Age version of the market.

Platform Current frontier contender What it increasingly optimizes for Writing advantage Biggest question
Claude Claude Opus 5 Deep reasoning, agents, coding, knowledge work Editing, interpretation, nuanced rewriting Has Claude lost some of its old prose magic?
ChatGPT GPT-5.6 Sol General knowledge work and end-to-end workflows Research + structure + editing + production Is competence replacing personality?
Gemini Gemini 3.1 Pro Preview Multimodal reasoning and Google integration Source-heavy, grounded writing Can information advantage become distinctive prose?
Grok Grok 4.6 Real-time knowledge, coding and agents Live commentary and internet-native research Is immediacy enough to make it a serious writer?

The specifications themselves are reaching the point of absurdity. GPT-5.6 Sol operates with million-token-class context. Claude Opus 5 offers a one-million-token context window and up to 128,000 output tokens. Gemini 3.1 Pro accepts more than one million input tokens, while Grok 4.6 arrives with a still enormous 500,000-token window.

Those numbers matter, particularly for research-heavy work, but they are becoming less useful as a way of distinguishing the models. Once every major system can ingest books, reports, transcripts and large folders of documents, the more important question is what the model actually does with all that material.

That is where the companies begin to look quite different. Anthropic, OpenAI, Google and xAI are not simply building competing text generators anymore. They are building different versions of a general-purpose knowledge worker, and writing is increasingly just one component of a much larger system.

The spiky theory of AI

One useful way to understand the current situation is through the idea of the jagged frontier, a term popularized by Wharton professor Ethan Mollick and his research collaborators.

The basic observation is that AI capability does not progress in a smooth, predictable line. A model can perform one task at what looks like an expert level, then fail badly at another task that seems much simpler. The frontier is uneven.

chart showing spiky theory of ai

The image above traces an evolving debate over what the path to AGI might actually look like. Its intellectual roots go back to the “jagged technological frontier” described by researchers including Ethan Mollick and Fabrizio Dell’Acqua in 2023: AI systems can be astonishingly capable at one task while failing at another that seems equally simple. In late 2025, Tomas Pueyo visualized this idea as an expanding, irregular “blob” of AI capability, arguing that as models improve, the jagged edges will gradually fill in until AI becomes broadly competent across the full range of human tasks — the mainstream AGI thesis shown in the top row. Adam Hunt later adapted that graphic to illustrate a more provocative alternative: perhaps the frontier never smooths out. Instead, certain capabilities such as coding, mathematics or scientific reasoning could shoot far beyond human ability while other areas — common sense, social judgment or aspects of language — improve much more slowly or even stagnate. The result would not necessarily be a neat, human-like general intelligence, but an increasingly strange and uneven “spiky” intelligence: superhuman in a handful of domains while remaining surprisingly mediocre in others.

That idea feels even more relevant today. The strongest models are developing enormous spikes of capability in areas such as coding, mathematics, scientific reasoning and agentic computer use. Progress in those areas is fast, visible and relatively easy to measure. Writing is a different proposition.

The models have undoubtedly improved at factual accuracy, instruction following, long-context comprehension, editing and structural reasoning. They are better at handling a complicated brief, keeping track of a large document and following an editorial style guide than earlier generations were.

What is much harder to argue is that the best AI prose today feels proportionately more distinctive, stylish or human than the best AI prose from a year or two ago.

There is a fairly obvious reason for that. Code has tests. Mathematics has answers. Software can be executed, proofs can be checked and agents can either complete a task successfully or fail to complete it. Writing has no equivalent.

There is no unit test for a strong opening paragraph, and there is no universally accepted score for whether one sentence has better rhythm than another. More awkwardly, many of the qualities we value in good writing are difficult to reconcile with the optimization pressures placed on an assistant model. Great prose can be ambiguous, abrasive, funny, incomplete, emotionally strange or deliberately indirect. A model trained to be clear, safe, helpful and comprehensive can easily become the kind of writer who explains everything twice.

That may be one reason AI can improve dramatically at calculus without becoming proportionately funnier, or become a much better programmer without acquiring noticeably better taste.

The benchmark problem: AI can measure code better than prose

The model launches themselves offer some evidence for this. Look at what the AI labs choose to benchmark.

Anthropic’s Claude Opus 5 launch is full of Frontier-Bench, CursorBench, ARC-AGI, AutomationBench, OSWorld, scientific evaluations and tests of professional knowledge work. Anthropic has plenty of numbers showing that Opus is becoming a better software engineer, researcher and agent. What you will not find is an equally definitive public benchmark establishing that Opus 5 is a better novelist or essayist than Opus 4.8.

Google’s Gemini 3.1 Pro announcement follows a similar pattern. Google discusses creative work, but the hard claims are dominated by reasoning, coding and difficult problem-solving.

OpenAI has been more willing to talk explicitly about writing. When it launched GPT-5, for example, the company described it as its most capable writing collaborator yet and showed examples of poems and other prose. Even there, however, the quantitative evidence focused primarily on mathematics, coding, multimodal reasoning, health and knowledge work rather than presenting a single definitive measure of literary quality.

xAI is an interesting partial exception. When it launched Grok 4.1, the company highlighted performance on the independent Creative Writing v3 benchmark, which tests models across 32 writing prompts and compares outputs against detailed rubrics.

That is useful, but it immediately exposes another problem with writing benchmarks: Creative Writing v3 is itself judged by an LLM. In other words, one machine is being asked to decide whether another machine has taste.

Researchers are trying to improve on this. WritingBench covers more than 1,200 writing tasks across six broad domains and 100 subdomains, including persuasive, creative, informative and technical writing. EQ-Bench’s Longform Writing benchmark makes models plan, revise and sustain a story over multiple long-form turns. LitBench goes further by using thousands of human preference judgments, in part because evaluating literary quality with another general-purpose model remains unreliable.

The writing benchmarks worth knowing

Benchmark What it tries to measure Why it matters The catch
Creative Writing v3 Creative writing across 32 prompts xAI publicly used it to demonstrate Grok 4.1 LLM-judged
Longform Writing Planning, revision and sustained fiction Tests coherence beyond a clever paragraph Still uses an AI judge
WritingBench 1,200+ tasks across 100 writing subdomains Much broader than fiction alone Automated evaluation remains imperfect
LitBench Alignment with human creative-writing preferences Built around thousands of human-labelled comparisons More about reliable evaluation than a simple consumer leaderboard

None of this makes writing benchmarks useless. They are getting better and they give us far more information than intuition alone.

They simply cannot settle the question in the way a coding or mathematics benchmark can, because writing eventually runs into taste.

Cormac McCarthy is not great because he would score perfectly against a corporate style guide. A board memo and a punk-rock manifesto are not improved by applying exactly the same rubric to both. Much of what makes writing distinctive comes from knowing when to break the expected pattern rather than follow it.

That is an unusually awkward challenge for systems built around predicting what should most plausibly come next.

Claude: has the former writing champion lost its mojo?

For years, Claude was the default recommendation among a certain kind of serious AI writer, and not without reason.

Earlier Claude models often felt less mechanical than the competition. They could handle rhythm and implication well, preserve the character of source material during a rewrite and, at their best, resist the urge to explain every point to death.

That reputation has carried forward, but the response to recent models has become much more mixed.

Anthropic’s development priorities are also revealing. Claude Opus 5’s headline improvements revolve heavily around software engineering, autonomous work, tool use, knowledge work, computer operation and scientific research. Those are valuable improvements, but they do not necessarily tell us whether Opus is now a more enjoyable or distinctive prose stylist.

Among heavy Claude users, there has been a noticeable amount of disagreement about that question. Some say newer Opus models have become more verbose, argumentative or mannered, with a recognizable form of “Claude-speak.” Others still regard Claude as comfortably ahead for dialogue, editing and creative work.

Anecdotes from Reddit and X obviously do not constitute a scientific evaluation, and every major model release produces a predictable wave of people insisting that the previous version was better. Even so, the disagreement matters because it marks a change in perception.

A couple of years ago, “Claude writes best” was close to conventional wisdom among many power users. In 2026, it is much more obviously an opinion.

Claude’s writing verdict

Category Verdict
Natural prose Still potentially excellent, but inconsistent
Editing existing writing Excellent
Voice imitation Very strong with examples and instructions
Fiction Strong contender, no longer automatic winner
Long-form reasoning Excellent
Research workflow Strong and improving
Model mannerisms A growing complaint among some users
Overall The former champion now has to defend the title

Claude still belongs in any serious writer’s toolkit. It remains especially good when it has strong source material, a detailed editorial brief and examples of the voice it is supposed to preserve.

What I would no longer do is assume, before comparing outputs, that Claude must be the best writer simply because Claude used to be the best writer.

ChatGPT: the writer that is becoming a newsroom

OpenAI’s strategy is less about owning the crown for individual sentences and more about owning the entire process around them.

GPT-5.6 Sol is positioned as a frontier model for complex professional work, with OpenAI emphasizing coding, research, knowledge work, design and long-running workflows. The August update also specifically pushed the model toward more focused answers and less unnecessary elaboration.

That sounds like a minor behavioral adjustment until you consider how much AI writing suffers from excessive explanation. Many models know the facts, understand the argument and can organize the material, but still feel compelled to introduce every point, restate it and then offer a tidy conclusion.

For writers, learning when not to say something is a meaningful capability improvement.

ChatGPT’s larger advantage, however, is the environment OpenAI has built around the model. Projects can hold instructions, source material and previous work. Research can extend onto the web. Documents can be developed over multiple sessions. Connected tools increasingly allow the model to move from source gathering to analysis, drafting, editing and production without forcing the user to rebuild the context at every stage.

A 1,500-word article may require finding regulatory filings, checking earlier coverage, pulling figures out of PDFs, comparing several sources, validating quotes, deciding what the story actually is, structuring the argument and then rewriting the piece several times. The prose itself is only one layer. OpenAI appears increasingly determined to absorb the rest of that workflow.

ChatGPT’s writing verdict

Category Verdict
Natural prose Very good
Editing Excellent
Research-heavy journalism Excellent
Argument construction Excellent
Voice consistency Strong with sufficient examples
Workflow/tool integration Probably the strongest overall proposition
Main weakness Can still drift toward polished generic competence
Overall Best one-tool choice for many professional writers

That does not necessarily make ChatGPT the model most likely to produce the best single paragraph in a blind test.

It may make it the system in which the largest share of the actual writing job gets done.

Gemini: Google owns the library

Google’s strategic advantage is easy to understand. The company already owns much of the infrastructure through which professional information flows, and AI writing is increasingly inseparable from information retrieval.

Gemini 3.1 Pro supports more than one million input tokens and is designed for sophisticated reasoning across large collections of text, images, video, audio and PDFs. Google’s own positioning focuses heavily on reasoning, synthesis and complex problem-solving rather than literary style.

That fits the wider pattern. The frontier labs are investing enormous resources in reasoning and agency because those capabilities can be measured, productized and turned into useful software.

Gemini’s strength for writers comes from everything surrounding the model.

Google has Docs, Drive and Gmail, while NotebookLM has become one of the more compelling mainstream tools for source-grounded research. That gives Gemini a natural advantage whenever writing begins not with a blank page but with a messy pile of source material.

A novelist may care relatively little about that infrastructure. An analyst or journalist working from dozens of reports, email threads, transcripts and spreadsheets probably cares a great deal.

Gemini does not need to convince every writer that it is Shakespeare. It can be enormously useful simply by being unusually good at finding, organizing and connecting the right information before the writing begins.

Gemini’s writing verdict

Category Verdict
Raw prose Capable, but not its clearest differentiator
Research synthesis Excellent
Large source collections Excellent
Multimodal research Excellent
Google Workspace users Potentially unmatched integration
Distinctive voice Still something I would edit aggressively
Overall The researcher’s AI writer

For research-led work, that may be enough to make Gemini the best first stop even if another model handles the final prose pass.

Grok: the correspondent standing in the street

Grok is the hardest model in this comparison to judge because its latest version is barely out of the gate.

Grok 4.6 arrived on August 12, 2026, so any sweeping declaration about its long-term writing quality should be treated cautiously. xAI’s current positioning again emphasizes coding, software engineering, tool use and agentic work rather than presenting 4.6 primarily as a creative-writing model.

Where Grok genuinely differs from the others is its relationship with X.

For better and for worse, X remains one of the places where breaking news, market narratives, political arguments, memes, rumors and expert commentary emerge in real time. That makes Grok unusually useful for writers covering technology, crypto, markets, politics and internet culture.

The obvious danger is that immediacy and reliability are not the same thing.

Social media can tell you what people are talking about long before conventional sources catch up, but it can also amplify claims that are incomplete, misleading or simply false. The most useful role for Grok is therefore not to act as the final authority but to identify the live narrative, trace where claims are coming from and show you what deserves verification.

Used that way, it feels less like a copy editor and more like a reporter who happens to be standing in the middle of the crowd.

xAI has also shown more interest in writing benchmarks than some people realize. Its Grok 4.1 release specifically promoted performance on Creative Writing v3, even though the 4.6 launch has shifted the emphasis back toward coding and agentic work.

Again, the development path looks spiky rather than linear.

Grok’s writing verdict

Category Verdict
Raw prose 4.6 is too new for a definitive call
Real-time awareness Major strength
X-native research Unique advantage
Commentary and zeitgeist Potentially excellent
Creative writing pedigree More interesting than many assume
Deep source-grounded writing Requires verification discipline
Overall The wildcard — and the live news desk

And now the words themselves are getting watermarked

Just as the argument about writing quality becomes more complicated, another issue is moving rapidly into the foreground: provenance.

Anthropic has confirmed that supported new Claude models will embed an invisible, machine-readable watermark into generated text. The policy is linked to the European Union AI Act’s transparency requirements, and Anthropic says supported models launched in the EU from August 2, 2026 include the marking from launch. The company is also applying the system more broadly rather than restricting it only to European users.

According to Anthropic, the watermark is embedded into the generated text itself and can survive ordinary copying and pasting, although extensive rewriting, paraphrasing or translation may weaken or remove it.

The important detail is that the mark does not prove Claude wrote the underlying material.

Anthropic explicitly says Claude-generated markings can appear when Claude has proofread, translated, summarized or otherwise processed text that originally came from a human. (support.claude.com)

That distinction is going to matter enormously.

Consider a journalist who writes an entire article and asks Claude to remove typos and tighten a few sentences. Or a novelist who uses Claude to suggest cuts. Or a student who writes an essay unaided but asks an AI to check the grammar.

Those documents may be AI-assisted without being meaningfully AI-authored.

The technical signal may therefore tell us something useful about provenance while telling us much less about authorship.

Claude isn’t actually first

Anthropic is not starting from an empty field.

Google DeepMind already uses SynthID to embed invisible signals into AI-generated content, including text. In the Gemini app and web experience, Google says the system alters token probabilities in a way that leaves a statistical signature while preserving the meaning and apparent quality of the output. (deepmind.google)

OpenAI is moving in the same general regulatory direction, but its implementation is currently different. The company supports the EU Code of Practice on Transparency of AI-Generated Content and already uses provenance systems for supported generated images and audio.

As of August 13, 2026, however, OpenAI’s public provenance documentation does not describe an equivalent watermark for ordinary ChatGPT text. (openai.com)

xAI’s public documentation is more limited again. Grok-generated images and videos carry visible or machine-readable provenance information, but the company’s current consumer documentation does not describe an equivalent system for normal text responses. (docs.x.ai)

The AI text watermark situation

Platform Text watermarking status What writers should know
Claude Rolling out on supported new models Invisible text mark can indicate Claude processed the material, not necessarily authored it
Gemini Yes, via SynthID in Gemini app/web Google already embeds an imperceptible statistical signal into generated text
ChatGPT No equivalent public text implementation documented yet OpenAI supports EU transparency rules and already marks supported images/audio
Grok No public text watermark documented xAI currently documents watermarks for generated images and video

This table is almost certain to change, but the direction of travel is clear.

The watermark wars could change which AI writers choose

There is a strong case for provenance technology. The internet is filling rapidly with synthetic content, and platforms, publishers and researchers need better ways to understand where material came from.

Text, however, creates a particularly messy problem because writing is rarely a single act of generation.

A human can write 95% of a document and ask an AI to revise the other 5%. An AI can generate the first draft and a human can substantially rewrite it. A person can dictate an original argument and ask a model only to organize it. A newsroom can pass the same document through several editors, some human and some artificial.

At what point does the finished document become “AI-generated”?

That question sounds philosophical until institutions begin attaching consequences to the answer.

For journalists, authors, academics and students, the difference between AI-generated and AI-assisted is potentially huge. Yet watermarking technology may detect machine involvement without being able to explain what that involvement actually consisted of.

It could also create an unexpected competitive dynamic between the platforms. If one AI tool reliably marks everything it edits while another does not, writers who dislike having their workflow encoded into the finished text may begin taking that into account when choosing a model.

That does not make watermarking inherently bad. Provenance is likely to become increasingly important as synthetic media spreads.

It does mean that provenance policy could become a product feature in its own right, alongside context windows, prices and model quality.

There is also a limit to what these systems can tell us. A detectable Claude watermark cannot establish whether an argument was original, whether a factual claim is true or whether the ideas came from the human or the machine. Likewise, the absence of a watermark is not proof of human authorship.

Anthropic itself warns against making those inferences.

That caveat may become very important once schools, publishers, employers, search engines and social platforms begin building policies around these signals.

So which AI should a writer actually use?

Here is where I currently land.

If your main job is… Start with… Why
Research-heavy article ChatGPT Best combination of research, reasoning, drafting and editing
Huge source dossier Gemini Google + NotebookLM is a formidable research stack
Creative rewrite Claude Still capable of excellent editorial transformation
Fiction Claude, but test against ChatGPT and Grok Claude’s crown is no longer uncontested
Breaking tech/crypto commentary Grok + ChatGPT Grok finds the live narrative; ChatGPT disciplines it
Editing your own prose Claude or ChatGPT Both can be stronger editors than autonomous authors
Corporate documents ChatGPT Broad end-to-end production workflow
Google Workspace-heavy work Gemini The integration advantage is obvious
Finding what the internet thinks right now Grok X is its distinctive weapon
One subscription for a professional writer ChatGPT Best overall utility rather than necessarily best sentence-level prose

The table is useful, but it also hides an increasingly obvious reality: professional writers do not necessarily need to choose one model and remain loyal to it.

The better approach may be to use them as specialized tools and, where the work matters, make them check one another.

The four-model newsroom

Stage Best tool What it does
1. Source gathering Gemini / ChatGPT Find primary documents and assemble research
2. Live narrative check Grok Surface reactions, debates and emerging claims
3. Thesis development ChatGPT Turn information into an argument
4. First draft ChatGPT / Claude Produce long-form prose from the research
5. Voice pass Claude / ChatGPT Rewrite mechanical or generic passages
6. Adversarial edit A different model Attack assumptions and identify weak sections
7. Fact check ChatGPT / Gemini Re-open sources and verify claims
8. Final human edit You Remove everything that sounds like a machine wrote it

That last stage is still doing a lot of work. It may continue doing so for quite a while.

The benchmark nobody has solved: taste

AI companies understandably like benchmarks. They give researchers a way to measure progress and allow model developers to make specific claims rather than simply insisting that the new model “feels smarter.”

Writing does not fit neatly into that framework. The new generation of creative-writing benchmarks is useful precisely because researchers are trying to measure things such as style, coherence, originality, emotional impact and long-form consistency rather than relying entirely on anecdote. LitBench’s use of human preference data is especially interesting because it acknowledges the limits of asking one language model to judge another.

Even so, no benchmark can completely remove the subjective part of the equation. Taste depends heavily on context and audience. The prose that works brilliantly in a literary novel may be completely wrong for a financial report. An aggressive first-person column and a corporate board memo should not sound remotely alike, even though a generalized writing evaluator might reward many of the same qualities in both.

There is also a deeper technical tension here. Language models are exceptionally sophisticated systems for modelling what language is likely to come next. Great writers, by contrast, often become memorable precisely because they know when not to choose the expected phrase.

That does not mean AI will never become an exceptional writer. There is already enough evidence to rule out that kind of complacency.

It does suggest that the path from greater reasoning ability to greater style is not automatic. The ability to solve harder problems can improve rapidly while literary taste, humour and originality move at a much slower pace.

That may turn out to be one of the defining characteristics of the current generation of AI.

The Gilded Age verdict

The old hierarchy is breaking down. Claude still has a serious claim as one of the best models for editing, stylistic transformation and creative collaboration, but its status as the automatic prose champion is much less secure than it once appeared. Opus 5 can be significantly more capable in many measurable ways while being less appealing to some writers. There is no contradiction in that.

ChatGPT increasingly looks like the strongest overall writing system rather than necessarily the strongest pure prose model. Research, files, memory, editing, tool use and production are becoming part of the same environment, which matters enormously for professional work.

Gemini has the most obvious structural advantage when writing begins with large quantities of information. Google’s ability to connect AI with documents, email, search and NotebookLM means Gemini does not need to win every stylistic comparison to become indispensable for certain kinds of research-heavy work.

Grok remains the wildcard. Its proximity to X and the real-time internet gives it a distinctive role for news, markets, technology and cultural commentary, while xAI’s previous performance on dedicated creative-writing benchmarks suggests it should not be dismissed as simply the internet’s snarkiest chatbot.

None of the four has solved taste, and the arrival of text watermarking now adds another complication. Writers are going to have to think not only about what a model produces, but also about what traces it leaves behind and what institutions may eventually infer from those traces.

For years, it was easy to assume that beautiful prose would emerge naturally once models became sufficiently intelligent.

The reality looks more complicated. AI can become much better at mathematics without becoming funnier. It can make enormous advances in software engineering while developing conversational habits that some users find more irritating. It can improve across dozens of benchmarks without producing a correspondingly dramatic leap in prose style.

That does not suggest AI development is slowing down. If anything, it suggests we are finally getting a clearer look at the shape of machine intelligence.

Progress is uneven, and the skills humans tend to bundle together under the word “intelligence” do not necessarily improve at the same rate. Writing may turn out to be one of the best places to see that difference.



Source link

Changelly

Be the first to comment

Leave a Reply

Your email address will not be published.


*