What the benchmark does

A paper posted to arXiv on 7 October introduces UltraText Bench, a test for whether an image generator can render dense visual text: long strings spread across several regions of the same picture, placed where the prompt says and still legible when it comes out.

The benchmark has 432 human-reviewed prompts across 24 real-world scene categories and three difficulty levels, split equally between English and Chinese. Each prompt supplies the exact strings for between four and twelve separate text regions, paired with a structured reference describing each region’s content, placement and visual attributes.

Generation is prompt-only. There is no reference image, no layout mask and no editing pass — the model is told what the words are and where they go, and has to produce the picture in one shot.

How it is scored

Scoring is automated. The authors use a vision-language model called Q-Judger to compare each generated image against the complete reference on four dimensions: text fidelity, text clarity, spatial quality and scene quality. Ten participants took part in a human evaluation of those automatic scores.

An open dictionary showing columns of printed words
Scenes with many separate pieces of text are where scores collapse. Illustrative image. Stefan G · pexels · Pexels License

Splitting fidelity from clarity turns out to matter. The authors report that Z-Image-Turbo gains 3.81 clarity points over Z-Image-Base while losing 14.76 fidelity points under the reported settings — in other words, it produces text that is sharper and more readable while getting more of the characters wrong. A single composite score would have shown an improvement.

Where the models fall over

The headline result is what happens as the text load goes up. Across 24 model configurations, performance tracks the number of regions and the number of characters rather than the difficulty of the scene. Qwen-Image-2512’s English composite falls from 86.50 at the easiest level to 42.86 at the hardest.

That is roughly half the score, from a model that is comfortably good at the easy end. Short-string rendering — a sign, a label, a word on a mug — has improved enough across the field that it no longer separates models. A menu, a poster or a form does.

Close-up of printed Arabic text on the page of a book
Dense text is what separates a picture of a document from a document. Illustrative image. Ahmad Shaufi · pexels · Pexels License

How to read the numbers

These are the authors’ own figures, reported in a preprint that has not been peer-reviewed, and the scorer is itself a model reading small text. The human evaluation involved ten people, which is a sanity check rather than a validation.

The benchmark covers English and Chinese only, which leaves out the scripts where rendering is hardest — Arabic shaping, Devanagari conjuncts, right-to-left layout. And as with any fixed prompt set, the scores will mean less once the prompts are in training data.

Why it matters

Dense text is the gap between a generated image that looks like a document and one that is a document. A model that can place twelve correct strings in the right places is a model that can produce a usable poster, slide, packaging mock-up or UI screen. On this evidence, none of the tested systems is there yet at the top difficulty level, and the ones that look closest are the ones measured at the bottom one.