Small Text Is Where AI Images Still Trip: What Designers Should Know About AI Infographics
Ask an AI image model for a poster with one big headline and you will usually get something usable. Ask it for an infographic with eighteen small captions, dates and numbers, and the picture changes. Some captions come out with wrong characters, some become unreadable, and some simply go missing.
For anyone who cares about type, this is the part of AI image generation worth watching closely. Small text, the kind that carries the actual information in a chart or a timeline, is where models differ the most.
What a controlled test found
In September 2026 the presentation tool Wonderslide published a test of 29 image models on infographics with small text. Every model got the same half-finished slide and the same brief: draw a horizontal timeline with an exact list of captions, in order, without touching the background or the title. There were seven slides, with 3, 10 or 18 blocks, in Latin, Cyrillic or mixed script. In total the seven slides carried 71 caption lines.
The study generated 250 images and scored them automatically with text recognition (OCR). Then a person checked the original images of three models line by line.
On that human review, two tiers of OpenAI’s GPT Image models, GPT Image 2 at medium quality and GPT Image 2.5 at low quality, had all 71 of 71 lines accepted in the first generation. Nano Banana 2 Fast had 43 of 71. The authors are careful to say this covers those three models on one slide template, not the whole market.
Where small text breaks
Density
The 18-block slide is the closest thing in the test to a real data-heavy infographic. On it, a person accepted 1 of 18 lines from Nano Banana 2 Fast. The same reviewer accepted all 18 lines from GPT Image 2 medium on that slide. When a model struggles with small text, more text makes it worse fast.
Script
The study ran the same ten events in English and in Russian. For Nano Banana 2 Fast, a person accepted 10 of 10 Latin lines and only 2 of 10 Cyrillic ones in the first generation. The two GPT tiers that were reviewed had all ten lines accepted in both scripts.
If you design for more than one language, test each script separately. A model that handles English captions well tells you little about how it handles Cyrillic, Greek or anything else.
Canvas size
Models did not return images at the same size. Frames in the study ranged from 1024×576 to 2816×1584, almost three times apart in width. Small type needs pixels. A caption that renders cleanly on a large frame can become a smudge on a small one, so check the output size before you judge the lettering.
A trap for type people: look-alike letters
One detail in the study’s limitations deserves attention from anyone who works with fonts. The automatic scorer treated look-alike Latin and Cyrillic letters as the same character. A Latin “c” inside a Russian word scored exactly like a Cyrillic “с”.
On screen these look identical. In practice they are different characters, which can affect search, copy and paste, screen readers and spell-checking. If an AI-generated graphic is going to be reused as text, or if you plan to retype or trace it, look for mixed alphabets. Most automated checks will not catch them.
Automated checks are not proof
The study’s most useful finding for designers may be about measurement itself. On 205 lines that a person could judge, the automatic score disagreed with the human verdict on 44 of them, about one in five. In 41 of those 44 cases, OCR marked a line as wrong that the person accepted.
In other words, a text recognizer is a rough filter, not a proofreader. It misreads small type the same way a tired eye does. If the numbers in an infographic matter, someone has to read the original image at full size.
Paying more does not always buy better lettering
Price did not track text quality neatly. On automatic scoring, the top tier of GPT Image 2 reproduced 58 lines against 56 for its medium tier, at 3.7 times the price. The newer GPT Image 2.5 low tier cost about a third of GPT Image 2 medium across the seven slides, and in three repeated runs the study could not separate the two on text quality.
Two models refused every slide outright because the brief exceeded their prompt-length limit. A long, precise brief, which is exactly what an infographic needs, can rule a model out before it draws anything.
A short checklist for AI infographics
- Test on your own content, at your real text density.
- Test each script you publish in, separately.
- Check the returned image size before judging the type.
- Read every caption at full size, especially numbers and dates.
- Watch for mixed alphabets if the text will be reused.
- Do not assume the most expensive tier writes the cleanest text.
The bottom line
AI image models have become good enough to draw small text, but only some of them, and only under some conditions. For designers, the practical skill is no longer choosing a font for a generated infographic. It is checking whether the model actually wrote the words you asked for.
Wonderslide ran this test because its own infographic mode builds timelines, charts and process diagrams from user data, so small-text accuracy is a product question for them. The full study publishes every prompt, score and generated image, which makes it a useful reference if you want to run a similar check on your own slides.