The newest image models are marketed on the two things they still get wrong
Nano Banana Pro, Seedream 4, Ideogram 3, FLUX Kontext — every launch leads with legible text and photorealism. The benchmarks say those failures just moved to where you are less likely to look.
A model launch in 2026 has a script. Show a photorealistic hero, put a crisp word or two of text inside the image, and claim the text problem is finally solved. Google’s Nano Banana Pro calls itself “the best model for creating images with correctly rendered and legible text”; Ideogram 3.0’s headline is literally “Photorealism. Legible text.”; Seedream 4.0 says it “broke through the bottlenecks of previous generative models in text processing”. The tell is that they are all advertising the same two capabilities — text and realism — which happen to be the two failure classes that made creative QA necessary in the first place. The claims are real improvements. The failures did not disappear; they moved.
Everyone is suddenly selling text
Two years ago, garbled lettering was the giveaway that an image was generated. Now text rendering is the headline feature of every flagship. Ideogram 3.0 published a human-preference ELO of 1132, ahead of Imagen 3 (1023), FLUX Pro (998), Recraft V3 (937) and DALL·E 3 (910), and sells on “Photorealism. Legible text.” Seedream 4.0 pushed resolution from 2K to 4K and claims it now renders text “correctly and clearly” alongside formulas and tables. Google’s Nano Banana Pro went further and simply named itself the best model in the world for in-image text.
The honest outlier is Black Forest Labs’ FLUX.1 Kontext, whose own announcement describes its typography as merely “competitive” — and even OpenAI’s gpt-image-1 makes the softer claim that it “reliably renders text within images,” pitched at storybooks and educational material rather than dense marketing layouts. When the whole industry’s marketing points at one feature, that feature is the one that used to be broken — and “better” is not the same as “done.”
The text problem moved; it did not leave
Independent benchmarks are less generous than the launch posts. On OCRGenBench, a visual-text benchmark of 19 state-of-the-art image models, only two — Nano Banana Pro (77.19) and Flux.2-dev (70.19) — clear 70 out of 100, and most score below 60. The best in-image text in the world lands at roughly 77%. That is fine for a three-word slogan and a coin-flip for a nutrition panel, a phone number or a paragraph of legal copy.
The STRICT stress test shows where the wall is. The strongest models, GPT-4o and Gemini, “maintain strong performance up to approximately 800 characters, beyond which accuracy begins to decrease”, while most diffusion models break down past about 200. And the failure is not always visible garble: “some models, especially FLUX 1.1 pro, tend not to follow instructions as the number of characters increases,” quietly dropping the requested text and painting a clean, wordless image instead. Non-Latin scripts are worst of all — the same study found “English generally yields the highest accuracy … while Chinese exhibits the lowest.” The cause is structural, not a maturity problem the next version outgrows: a locality bias in how diffusion models build an image limits their grip on long-range spatial structure like a line of text.
So “legible text” is true for a headline and false for a paragraph, a serial number, or a logotype in a second language — which is exactly the distinction a human skimming a batch of 300 product images will miss.
Hands and physics did not get a keynote
The failure classes nobody puts on a launch slide are still measurable. A 2026 anatomy benchmark scores current models with an Anatomical Error Rate and gives FLUX.1-dev 0.417 and SDXL 0.873 before correction; a comprehensive GPT-4o evaluation flags “abnormal human poses or anatomical structures” and “spatially implausible object overlaps” as recurring weaknesses. Physics is worse: on PhyBench the best text-to-image model scored only 1.71 out of 3 on mechanics, and DALL·E 3 “generates a balanced seesaw with an elephant and a mouse on either side.” A January 2026 benchmark of 21 state-of-the-art models concludes that “higher-order spatial reasoning remains a primary bottleneck,” and T2I-CompBench finds spatial relationships the single hardest sub-category, with counting still under ~0.60.
The most telling sign is that the evaluations themselves keep being rebuilt because the models saturate them. GenEval, once well-aligned with human judgment, has “drifted far” from it — an absolute error “as much as 17.7% for current models” — and its harder successor GenEval 2 puts the top model at just 35.8% prompt-level accuracy, collapsing toward zero as prompts get more compositional. The scores went up because the test got easier relative to the models, not because hands and physics got solved.
Editing models bring their own defects
The real 2025–26 shift is from generating to editing — instruction models like FLUX.1 Kontext and Nano Banana that change one thing in an existing image. That adds a new failure class. FLUX.1 Kontext’s own paper states that “excessive multi-turn editing can introduce visual artifacts that degrade image quality,” that the model “occasionally fails to follow instructions … ignoring specific prompt requirements,” and that even its distillation step “can introduce visual artifacts.” Google, launching Gemini 2.5 Flash Image, listed the open problems out loud: it is “actively working to improve long-form text rendering, even more reliable character consistency, and factual representation.”
And the provenance mark that ships with these models is not a quality signal. Every image from Gemini carries an invisible SynthID watermark so it can be identified as AI — which tells you the image is generated, not whether the hand has five fingers or the logo is spelled right. The failure classes in 2026 are the same ones from three model generations ago; they are just rarer and better hidden, which makes catching them by eye harder, not easier. You cannot audit the model, so you have to audit the output — one scan per creative before it ships.
What to do about it
- Read “legible text” as “legible short text.” The best model scores ~77/100 on a text benchmark and degrades past a few hundred characters — fine for a slogan, unreliable for a paragraph, a price, or a non-Latin logo.
- Keep checking hands, physics and spatial layout: 2026 benchmarks still measure anatomical errors and call spatial reasoning “a primary bottleneck,” even on the models topping the leaderboards.
- Treat multi-turn AI editing as a fresh risk, not a safe cleanup: the edit models’ own authors document accumulating artifacts and ignored instructions across turns.
- Do not read a SynthID or C2PA mark as a quality pass — it says “AI made this,” not “this is correct.”
- Automate the output check where volume is high: one scan returns text, anatomy, physics and artifact findings per creative in about two minutes, regardless of which model made it.
Sources
- Nano Banana Pro: Gemini 3 Pro Image model from Google DeepMind — Google
- Ideogram 3.0 — Realism, design, and consistent styles — Ideogram
- Seedream 4.0 Officially Released: Beyond Drawing, Into Imagination — ByteDance Seed
- Introducing FLUX.1 Kontext and the BFL Playground — Black Forest Labs
- Unveiling GPT-image-1 in Azure AI Foundry — Microsoft Azure Blog
- V7 Alpha — Midjourney
- Adobe Firefly: The next evolution of creative AI is here — Adobe
- STRICT: Stress Test of Rendering Images Containing Text — arXiv
- OCRGenBench: A Comprehensive Benchmark for Evaluating OCR Generative Capabilities — arXiv
- Towards Anatomically Plausible Human Image Generation (HAF-Bench / ASAP) — arXiv
- GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT-4o in Image Generation — arXiv
- PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models — arXiv
- Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models — arXiv (ICLR 2026)
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation — arXiv
- T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-Image Generation — arXiv
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing — Black Forest Labs (arXiv)
- Introducing Gemini 2.5 Flash Image — Google Developers Blog