If you make presentations, proposals and reports for a living, the last four months of AI announcements have been hard to follow and easy to misread. This is a plain-English guide: what shipped, what it means for document work, and which of it you can actually use in NextDocs today.
| Model | Vendor | Released | Why it matters for documents |
|---|---|---|---|
| Claude Opus 4.8 | Anthropic | 28 May 2026 | The last Opus before the 5 series; still a strong long-form writer |
| Claude Fable 5 and Mythos 5 | Anthropic | 9 Jun 2026 | New frontier tier. Mythos is the same model with safeguards lifted and restricted access |
| Claude Sonnet 5 | Anthropic | 30 Jun 2026 | Most agentic Sonnet; $2 in / $10 out per million tokens |
| GPT-5.6 (Sol, Terra, Luna) | OpenAI | 9 Jul 2026 | A three-tier family. Luna is the fast one, Sol the flagship |
| Kimi K3 | Moonshot | 16 Jul 2026 | 2.8 trillion parameter open-weights model, 1M context |
| Claude Opus 5 | Anthropic | 24 Jul 2026 | Near-Fable quality at half the price; $5 in / $25 out |
| DeepSeek V4-Pro | DeepSeek | 13 Aug 2026 | Agent-focused, 1M context, peak and off-peak pricing |
| Gemini Omni 1.1 Flash | 27 Aug 2026 | "Create anything from any input"; 4K upscale and scene extension | |
| Claude Fable 5.1 and Mythos 5.1 | Anthropic | 1 Sep 2026 | Most capable generally available Claude; $10 in / $50 out, cached reads cut to $0.25 |
| Gemini 3.8 Flash | 2 Sep 2026 | The fourth Flash generation in four months | |
| GPT-6 Astra | OpenAI | 4 Sep 2026 | Preview on 3 Sep, general availability the next day |
Two corrections to things you may have read: there is no Grok 5 and no Llama 5. Meta replaced Llama with Muse Spark in April and shipped Muse Spark 1.1 on 9 July. And Gemini 3.5 Pro, announced at I/O in May, had still not shipped by late August.
Long documents got more reliable, not just longer. The interesting number in these releases is not context length but how well a model holds structure across a 20-page document: consistent headings, numbered sections that stay numbered, tables that keep their columns, a conclusion that matches the introduction. That is what we test for, and it is where the new generation separates from the old one.
Prices went both ways. Fable 5.1 is a premium model at $10 and $50 per million tokens. Sonnet 5 at $2 and $10 is cheaper than the model it replaces. Gemini Flash keeps getting faster and cheaper each generation. For a document tool, this matters because a 60-page report is a lot of tokens; the choice of lane changes what a document costs to make.
Agents became the default framing. Every vendor now describes its model as an agent. For documents, the useful version of that is a model that can plan a document, write it, look at the result and fix it. NextDocs has run that create, verify, refine loop since v1.8; the new models make each step better.
NextDocs does not run one model. It runs a ladder, and the rung depends on your plan and the mode you pick.
Alongside the lanes, the model picker lets you choose explicitly: Gemini 3 Flash and Gemini 3.8 Flash, Gemini 3.1 Pro Preview, GPT-5.5, Claude Opus 4.8, Claude Sonnet 4.6, Kimi K2.6, GLM 5.3 and Qwen 3.7 Max. Some are on the free plan, the larger ones on Pro and above.
Why not simply the newest model everywhere? Because we test before we promote. Claude Sonnet 5 was tried on our 22-page document gate and did not pass it, so Sonnet 4.6 stayed on the premium lane. Gemini 3.8 Flash did pass, and it went live within days. The next post explains the gate and what we learned running five models on the same deck.
GPT-6 Astra and Claude Fable 5.1 are not in NextDocs yet. If you already pay for Claude Code or Codex, there is a way to use them for document work today: Shyne, our sister product, has a desktop app that detects a locally installed Claude Code, Codex or opencode and lets you run the Shyne agent on it. That means Claude Fable 5.1, Opus 5 and Sonnet 5 through Claude Code, or GPT-6 Astra and the GPT-5.6 family through Codex, billed to the subscription you already have and at no extra cost from Shyne.
Shyne's hosted models, if you would rather not bring your own, are the GPT-5.6 family with Terra as the default, plus Kimi K2.6, K2.7 Code and K3.
Model names will keep changing every few weeks. The question worth asking stays the same: does it produce a document you can send?
Try the model picker in NextDocs Β· Bring your own model on Shyne desktop
The NextDocs Team

Before a model gets a lane in NextDocs it has to finish a real 22-page document from production. Here is how the gate works, what GPT-5.6 Luna, Gemini 3.5 and 3.8 Flash, Claude Sonnet 4.6 and Claude Sonnet 5 did on it, and why the newest model is not always the one you want writing your report.
Read more
Generate up to 4 document variants simultaneously. Compare different structures, visual directions, and stories side by side. Layer themes on top. Pick your favorite. This changes everything about how you create.
Read more
We now make two products. This is the short, honest answer to which one fits what you're doing β plus what happens to your NextDocs account (nothing) and where our new work goes (Shyne).
Read more