Methodology notes: small-text infographic study Wonderslide, October 2026 Article: https://wonderslide.com/blog/infographic-small-text-image-models/ What this study did and did not measure, and why it matters for reading the numbers. 1. Channel and prices All 29 participants were called through a single API intermediary, on the catalog snapshot of September 17, 2026. Prices are computed from the intermediary's published price list and were not reconciled against an invoice: the account key is shared with a production system, so its balance could not serve as a meter. Prices, availability and behavior may differ through another channel, a vendor's own product or a later catalog. The catalog changes weekly; it grew by 57 entries in a single day during the study. 2. What a person reviewed The study owner compared the visible words and numbers in 30 original generated images with the reference text, 304 expected lines in all. 21 images were a stratified sample: all 7 slides, first generation only, for GPT Image 2 medium, GPT Image 2.5 low and Nano Banana 2 Fast (213 expected lines). 9 images were failure-selected diagnostic images (91 expected lines). Every line has a verdict in line-audit.csv. The review covers body text of these outputs. It is not a transcription, and it does not certify capitalization, punctuation, aesthetics or title preservation. The reviewer saw the reference text and enlarged crops, with original images available, so the judgments were not blind. Other models, later runs and the remaining outputs were not reviewed by a person. 3. How far automatic scoring can be trusted Excluding 8 unreadable lines, the stratified sample has 205 determinate judgments. Automatic exact-line scoring disagreed with the person on 44 of them, 21.46%. Of these, 41 are lines the person accepted and automatic scoring marked as not exact, and 3 are lines automatic scoring accepted and the person did not (all three are Nano Banana 2 Fast). The disagreement differs by model: 14.08% for GPT Image 2.5 low, 21.13% for GPT Image 2 medium and 30.16% for Nano Banana 2 Fast. These rates apply to the reviewed lines only; they are not correction factors, and other automatic scores must not be multiplied by them. The figure measures the whole automatic pipeline: recognition, geometric reconstruction, line assignment and normalization. It is not a recognizer-only character error rate, because no complete manual transcription exists for every deviation. The scorer reads, for example, K-1 as K-l and production line №1 with Ng1, and fragments of neighboring columns can be interleaved. The 9 diagnostic images (from Step1X Edit, Qwen Image Edit Plus, Seedream 4.0 and Ideogram v4) are reported separately and are not pooled into the rate: of their 91 lines, 75 were judged unreadable and 16 absent, and none was accepted. 4. Repeats The finalists were run up to three times per slide. That shows the spread between runs, but it is not enough for statistical significance. All repeated runs were scored automatically only. GPT Image 2 high, the most expensive participant, was run once per slide; its results are single measurements, no run-to-run spread is known for it, and they must not be compared with a median of repeats. 5. Test slides All seven slides are synthetic and use one template: a horizontal timeline with a year at the top of each column. The block reconstruction the scoring depends on relies on that layout; other infographic layouts would need a different method. The results describe this template and the two scripts it covers, not infographics in general. 6. Frames Returned frames range from 1024x576 to 2816x1584: 2.75 times in linear size, 7.6 times in area. A model with no size parameter, or one that ignores it, returns its own size. Small text suffers on a small canvas, so every accuracy number is a number at a frame, and two participants are comparable only when their frames are shown. One participant can return different frames at identical settings: GPT Image 2 returned 2048x1152 and 2560x1440 across repeats. The figures show such participants as a range, never as an average. 7. The comparison behind the recommendation is unequal GPT Image 2 medium returned 2048x1152 on 19 of its 21 generations; GPT Image 2.5 low returned 2560x1440 on all 21. The recommendation rests on price, which the frame does not affect. The automatic line gap between the two is not evidence of a quality difference: in the human review both had all 71 lines accepted. 8. The brief and the request asked for different sizes Every brief ends with "Return the edited slide at exactly 1280x720 pixels, matching the input canvas exactly", while the request asked for about 2560x1440 (or the nearest tier). None of the 250 generations came back at 1280x720. The returned sizes are real facts about each model, but the two instructions contradict each other, so a returned size that differs from the request cannot be read as disobedience. 9. What the automatic scorer cannot see Before comparing, the scorer folds twelve Latin letters into their Cyrillic look-alikes (o, a, e, c, p, x, y, k, m, t, b, h) on both sides, so a Latin "c" inside a Russian word scores the same as a Cyrillic "с". It lower-cases both sides, so letter case is not scored. It drops the standalone dash between a year and an event, which the recognizer often fails to see; a hyphen inside a token, such as the product code K-1, is kept. 10. A comparison of the scorer with a by-hand calculation A by-hand calculation on recognized text for all 24 recognized images of slide D10 matched the scorer's character-error rate within rounding on 21 images and differed on three (Flux 2 Pro, Nano Banana and Seedream 5.0 Lite). In all three the cell scored zero exact lines and the text was heavily garbled. This compared two methods on recognized text; it is not a comparison with a person's reading of the images. 11. One slide pair moves two variables The brief of slide D3 is 1,510 characters against 4,872 for D10, because the list of blocks makes up most of the brief. The step from D3 to D10 changes the block count and the brief length together, by a factor of 3.2. Density is therefore read on D10 against D18, whose briefs are 4,872 and 5,312 characters. 12. Coverage and refusals 27 of the 29 participants were measured. Qwen Image 3.0 and Qwen Image 3.0 Pro refused all seven slides on prompt length; they are absent from the results, not scored zero. Seven more covered fewer than seven slides: Ideogram v4 and both Kling variants one each, MAI Image 2.5 four, Luma Uni six (prompt-length limit); Nano Banana four and FIBO six (platform failures on the densest slides). Their totals are not placed beside seven-slide totals. Participants blocked by prompt length were not re-run on a shortened brief. A refusal says that the model refused, not that it draws text worse. 13. Reliability of the reassembly step 2.9% of ground-truth lines (47 of 1617) had their words scattered across unrelated reconstructed blocks, against a 10% threshold set before the count. In 11.0% of scored cells (18 of 163), some line's words landed inside their own block but not as one in-order run. Both numbers describe the reassembly step, not recognition accuracy. 14. Layout order Some models place events in a different order than the brief lists them. This does not affect the text score, because each event is matched to its own block, and layout order was not scored separately.