Blogs
AI ModelsGPT-6ClaudeGeminiKimiNextDocs

The September 2026 model wave, explained for people who make documents

6 min
The September 2026 model wave, explained for people who make documents

The September 2026 model wave, explained for people who make documents

If you make presentations, proposals and reports for a living, the last four months of AI announcements have been hard to follow and easy to misread. This is a plain-English guide: what shipped, what it means for document work, and which of it you can actually use in NextDocs today.

What shipped, in order

Model Vendor Released Why it matters for documents
Claude Opus 4.8 Anthropic 28 May 2026 The last Opus before the 5 series; still a strong long-form writer
Claude Fable 5 and Mythos 5 Anthropic 9 Jun 2026 New frontier tier. Mythos is the same model with safeguards lifted and restricted access
Claude Sonnet 5 Anthropic 30 Jun 2026 Most agentic Sonnet; $2 in / $10 out per million tokens
GPT-5.6 (Sol, Terra, Luna) OpenAI 9 Jul 2026 A three-tier family. Luna is the fast one, Sol the flagship
Kimi K3 Moonshot 16 Jul 2026 2.8 trillion parameter open-weights model, 1M context
Claude Opus 5 Anthropic 24 Jul 2026 Near-Fable quality at half the price; $5 in / $25 out
DeepSeek V4-Pro DeepSeek 13 Aug 2026 Agent-focused, 1M context, peak and off-peak pricing
Gemini Omni 1.1 Flash Google 27 Aug 2026 "Create anything from any input"; 4K upscale and scene extension
Claude Fable 5.1 and Mythos 5.1 Anthropic 1 Sep 2026 Most capable generally available Claude; $10 in / $50 out, cached reads cut to $0.25
Gemini 3.8 Flash Google 2 Sep 2026 The fourth Flash generation in four months
GPT-6 Astra OpenAI 4 Sep 2026 Preview on 3 Sep, general availability the next day

Two corrections to things you may have read: there is no Grok 5 and no Llama 5. Meta replaced Llama with Muse Spark in April and shipped Muse Spark 1.1 on 9 July. And Gemini 3.5 Pro, announced at I/O in May, had still not shipped by late August.

What actually changed for document work

Long documents got more reliable, not just longer. The interesting number in these releases is not context length but how well a model holds structure across a 20-page document: consistent headings, numbered sections that stay numbered, tables that keep their columns, a conclusion that matches the introduction. That is what we test for, and it is where the new generation separates from the old one.

Prices went both ways. Fable 5.1 is a premium model at $10 and $50 per million tokens. Sonnet 5 at $2 and $10 is cheaper than the model it replaces. Gemini Flash keeps getting faster and cheaper each generation. For a document tool, this matters because a 60-page report is a lot of tokens; the choice of lane changes what a document costs to make.

Agents became the default framing. Every vendor now describes its model as an agent. For documents, the useful version of that is a model that can plan a document, write it, look at the result and fix it. NextDocs has run that create, verify, refine loop since v1.8; the new models make each step better.

Which models NextDocs runs today

NextDocs does not run one model. It runs a ladder, and the rung depends on your plan and the mode you pick.

  • Fast lane: GPT-5.6 Luna, running on Azure AI Foundry.
  • Quality lane: Gemini 3.8 Flash, promoted on 3 September after it passed our long-document gate.
  • Premium lane (Pro+ and Ultra): Claude Sonnet 4.6.

Alongside the lanes, the model picker lets you choose explicitly: Gemini 3 Flash and Gemini 3.8 Flash, Gemini 3.1 Pro Preview, GPT-5.5, Claude Opus 4.8, Claude Sonnet 4.6, Kimi K2.6, GLM 5.3 and Qwen 3.7 Max. Some are on the free plan, the larger ones on Pro and above.

Why not simply the newest model everywhere? Because we test before we promote. Claude Sonnet 5 was tried on our 22-page document gate and did not pass it, so Sonnet 4.6 stayed on the premium lane. Gemini 3.8 Flash did pass, and it went live within days. The next post explains the gate and what we learned running five models on the same deck.

How to use the very newest ones

GPT-6 Astra and Claude Fable 5.1 are not in NextDocs yet. If you already pay for Claude Code or Codex, there is a way to use them for document work today: Shyne, our sister product, has a desktop app that detects a locally installed Claude Code, Codex or opencode and lets you run the Shyne agent on it. That means Claude Fable 5.1, Opus 5 and Sonnet 5 through Claude Code, or GPT-6 Astra and the GPT-5.6 family through Codex, billed to the subscription you already have and at no extra cost from Shyne.

Shyne's hosted models, if you would rather not bring your own, are the GPT-5.6 family with Terra as the default, plus Kimi K2.6, K2.7 Code and K3.

A simple way to choose

  • Drafting a deck or a short document, and speed matters: the fast lane is fine.
  • A long report, a proposal with sections that must stay consistent: the quality lane.
  • Nuanced writing where tone matters, or a document you will send to a board: the premium lane, or pick Claude Opus 4.8 in the picker.
  • You want to compare: generate several versions with a different model each and keep the best one. This is the one thing a document tool can do that a chat window cannot.

Model names will keep changing every few weeks. The question worth asking stays the same: does it produce a document you can send?

Try the model picker in NextDocs Β· Bring your own model on Shyne desktop


The NextDocs Team