{"id":1400,"date":"2026-10-01T11:11:42","date_gmt":"2026-10-01T11:11:42","guid":{"rendered":"https:\/\/wonderslide.com\/blog\/?p=1400"},"modified":"2026-10-01T19:16:58","modified_gmt":"2026-10-01T19:16:58","slug":"infographic-small-text-image-models","status":"publish","type":"post","link":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/","title":{"rendered":"Which Image Model Can Draw an Infographic With Small Text? We Tested 29"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\"><em>Updated 2026-10-01.<\/em><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Wonderslide <a href=\"https:\/\/wonderslide.com\/features\/\">builds infographics from your data<\/a>: timelines, charts and processes with small text. Our infographic mode runs on GPT Image 2 at its medium quality tier, $0.11 an image. We wanted to know two things: whether a cheaper model is no worse on small text, and how other image editors compare.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The person who ran this study, its owner, checked original images by eye. In the human-reviewed first generation across the same seven cases, the owner accepted 71 of 71 lines for GPT Image 2 medium and GPT Image 2.5 low, versus 43 of 71 for Nano Banana 2 Fast. A line here is one caption on a timeline, such as &quot;2009 \u2014 Founded in Verano&quot;; the seven slides carry 71 of them. Image sizes (frames) returned: GPT Image 2 medium 2048&#215;1152, GPT Image 2.5 low 2560&#215;1440, Nano Banana 2 Fast 2741&#215;1530. This comparison applies to these reviewed outputs; it is not a market-wide ranking, an exact character-error estimate, or a guarantee for later generations.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">GPT Image 2.5 low costs a third as much as GPT Image 2 medium: $0.245 against $0.770 over the seven cases.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">In September 2026 Wonderslide tested 29 image models on seven synthetic timeline slides: 250 generations scored automatically (OCR, text recognition), plus a person&#8217;s review of original images: the first generations of three models, and nine diagnostic images from four others. Every prompt, score, review judgment and generated image (as previews) is published at the end of this article.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Key takeaways<\/h2>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>A person accepted every body-text line from GPT Image 2 medium and GPT Image 2.5 low.<\/strong> That is 71 of 71 in the first generation of each slide, against 43 of 71 for Nano Banana 2 Fast. The frames differed, and the review covers these three models only.<\/li>\n\n\n<li><strong>GPT Image 2.5 low costs a third of GPT Image 2 medium<\/strong> over the seven cases. Automatic scoring over three runs each cannot separate the two on text quality either.<\/li>\n\n\n<li><strong>Automatic scoring is not the final word.<\/strong> On the 205 first-generation lines a person could judge for the three reviewed models, it disagreed with the review on 44 (21.46%), and in 41 of those 44 it marked a line wrong that the person accepted. Every score not marked as checked by a person is automatic.<\/li>\n\n\n<li><strong>The top tier buys nothing measurable.<\/strong> On automatic scoring, GPT Image 2 high reproduced 58 lines against 56 for medium on one generation per case, at 3.7 times the price (frames differ; see the table).<\/li>\n\n\n<li><strong>Cyrillic held up when a person checked.<\/strong> All 10 Cyrillic lines from GPT Image 2.5 low and from GPT Image 2 medium were accepted, where automatic scoring counted 7 and 9. Over all runs, automatic scoring counted 18 Cyrillic misses against 4 Latin ones for the five GPT tiers, a gap the review did not confirm in the first generation of the two reviewed tiers; the other runs were not reviewed.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">GPT Image tiers compared: price and text accuracy<\/h2>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Tier<\/th><th>Price for the 7 cases<\/th><th>Checked by a person: exact lines (of 71)<\/th><th>Automatic OCR: exact lines, first generation (of 71)<\/th><th>Automatic OCR: character error rate, first generation<\/th><th>Frame returned, first generation<\/th><\/tr><\/thead><tbody><tr><td>GPT Image 2 low<\/td><td>$0.210<\/td><td>not reviewed<\/td><td>52<\/td><td>0.049<\/td><td>2048&#215;1152 on 4 cases, 2560&#215;1440 on 3<\/td><\/tr><tr><td>GPT Image 2 medium<\/td><td>$0.770<\/td><td>71<\/td><td>56<\/td><td>0.060<\/td><td>2048&#215;1152 on 7 cases<\/td><\/tr><tr><td>GPT Image 2 high<\/td><td>$2.870<\/td><td>not reviewed<\/td><td>58<\/td><td>0.072<\/td><td>2048&#215;1152 on 5 cases, 2560&#215;1440 on 2<\/td><\/tr><tr><td>GPT Image 2.5 low<\/td><td>$0.245<\/td><td>71<\/td><td>61<\/td><td>0.041<\/td><td>2560&#215;1440 on 7 cases<\/td><\/tr><tr><td>GPT Image 2.5 medium<\/td><td>$0.385<\/td><td>not reviewed<\/td><td>59<\/td><td>0.050<\/td><td>2560&#215;1440 on 7 cases<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">The first generation is one generation per case across all 7 cases (round 1 in the study). The automatic columns are the original OCR measurements, not human-verified accuracy. Frames differ between tiers, so read each line count together with its frame. Measured by Wonderslide, September 2026.<\/figcaption><\/figure>\n\n\n\n<h2 class=\"wp-block-heading\">How we tested<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Each of the 29 participants got the same half-finished 1280&#215;720 slide (a flat background and a title) and a written brief: draw a horizontal timeline with the exact lines listed, in order, without touching the background or the title. The participants were 24 third-party image editors and five GPT Image tiers: GPT Image 2 low, medium and high, and GPT Image 2.5 low and medium. Each tier counts as its own participant.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The seven slides differ in block count (3, 10 or 18), script (Cyrillic, Latin or mixed), brief length and illustration style (detailed instead of flat icons). Most pairs differ along one axis; the limitations below name the exception that matters here. In round 1, the first generation, every participant was given every slide once. The eight finalists (the five GPT Image tiers, Nano Banana 2 Fast, Seedream 5.0 Pro and Luma Uni) were then run up to two more times per slide, to see how much the result moves between identical runs. GPT Image 2 high, the most expensive tier, was not repeated.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Every score except the human review comes from automatic scoring (OCR, text recognition): a recognizer reads the text back from each image, and a script compares it with the reference lines. Every model was called through one API intermediary, on the catalog snapshot of 2026-09-17.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The owner of the study then compared the visible words and numbers in 30 original generated images with the reference text, 304 expected lines in all. 21 images were the main sample: all 7 slides, first generation, for GPT Image 2 medium, GPT Image 2.5 low and Nano Banana 2 Fast. The other 9 were diagnostic images, picked because earlier checks had flagged them as failures.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A model refusing to attempt a slide because the brief was too long is a different outcome from a model drawing the slide badly, and we kept the two separate throughout: of 203 first-round cells, 36 are refusals on prompt length, 163 were generated and scored, and 4 failed on the platform&#8217;s side without producing an image.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Both metrics below are automatic. We publish two text metrics, not one: lines reproduced exactly and characters reproduced wrongly. On a single generation per case the five GPT tiers sit between 61 and 52 exact lines out of 71, and between 0.041 and 0.072 wrong characters.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">GPT Image vs Nano Banana 2: what a person found in the original images<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">On the same 7 cases in the first generation, the owner accepted 71 of 71 lines for GPT Image 2 medium and 71 of 71 for GPT Image 2.5 low, versus 43 of 71 for Nano Banana 2 Fast. Nano Banana also had 19 readable deviations, 8 unreadable lines and 1 absent line. This is the human-reviewed comparison; the entire market and later runs have not been reviewed by eye.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Frames differ: Nano Banana returned 2741&#215;1530, GPT Image 2.5 low 2560&#215;1440, and GPT Image 2 medium 2048&#215;1152.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">These are human judgments about words and numbers, not an independent transcription or a certification of capitalization, punctuation, aesthetics or title preservation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The 9 failure-selected diagnostic images contain 91 expected lines: 75 were judged unreadable and 16 absent. None was accepted as exact. This is evidence about the specifically flagged outputs, not a random sample of recognizer performance.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">How reliable are the automatic scores?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The owner reviewed 30 original images and 304 expected lines. In the stratified sample, automatic exact-line scoring disagreed on 44 of 205 determinate human judgments, or 21.46%. This is measurement-pipeline disagreement, not recognizer-only character error.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The 8 unreadable stratified lines are excluded from that denominator. The separately selected diagnostic images are not pooled into the rate. No complete manual transcription exists for every deviation, so a pure recognizer character-error rate cannot be claimed.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Of the 44 disagreements, 41 were lines a person accepted that automatic scoring marked wrong, and 3 were lines it accepted that a person did not. The rate differs by model: 14.08% for GPT Image 2.5 low, 21.13% for GPT Image 2 medium and 30.16% for Nano Banana 2 Fast.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">These rates are conditional on the reviewed, readable-or-absent lines; they are not global correction factors. Do not multiply other automatic scores by them.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The source readings show short-word omissions and substitutions such as K-1 read as K-l, production line \u21161 read with Ng1, 1000-\u044f read with Cyrillic O letters, and fragments interleaved between neighboring columns.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">All 29 models: automatic scores and what a person checked<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">This table lists every participant. The automatic scores are the original measurements and have not been corrected. For the three models a person reviewed, they came out lower than the human count: 56 against 71 for GPT Image 2 medium, 61 against 71 for GPT Image 2.5 low and 30 against 43 for Nano Banana 2 Fast.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>Model<\/th><th>Group<\/th><th>Automatic OCR: exact lines, first generation<\/th><th>Checked by a person: exact lines<\/th><th>Frame returned<\/th><th>Price per generation<\/th><th>Slides covered (of 7)<\/th><\/tr><\/thead><tbody><tr><td>GPT Image 2 low<\/td><td>GPT Image<\/td><td>52 of 71<\/td><td>not reviewed<\/td><td>2048&#215;1152 (4), 2560&#215;1440 (3)<\/td><td>$0.030<\/td><td>7<\/td><\/tr><tr><td>GPT Image 2 medium<\/td><td>GPT Image<\/td><td>56 of 71<\/td><td>71 of 71<\/td><td>2048&#215;1152<\/td><td>$0.110<\/td><td>7<\/td><\/tr><tr><td>GPT Image 2 high<\/td><td>GPT Image<\/td><td>58 of 71<\/td><td>not reviewed<\/td><td>2048&#215;1152 (5), 2560&#215;1440 (2)<\/td><td>$0.410<\/td><td>7<\/td><\/tr><tr><td>GPT Image 2.5 low<\/td><td>GPT Image<\/td><td>61 of 71<\/td><td>71 of 71<\/td><td>2560&#215;1440<\/td><td>$0.035<\/td><td>7<\/td><\/tr><tr><td>GPT Image 2.5 medium<\/td><td>GPT Image<\/td><td>59 of 71<\/td><td>not reviewed<\/td><td>2560&#215;1440<\/td><td>$0.055<\/td><td>7<\/td><\/tr><tr><td>Boogu<\/td><td>Third-party editor<\/td><td>0 of 71<\/td><td>not reviewed<\/td><td>1024&#215;576<\/td><td>$0.060<\/td><td>7<\/td><\/tr><tr><td>FIBO<\/td><td>Third-party editor<\/td><td>0 of 61<\/td><td>not reviewed<\/td><td>1360&#215;768<\/td><td>$0.040<\/td><td>6<\/td><\/tr><tr><td>Flux 2 Dev<\/td><td>Third-party editor<\/td><td>0 of 71<\/td><td>not reviewed<\/td><td>2048&#215;1152<\/td><td>$0.024<\/td><td>7<\/td><\/tr><tr><td>Flux 2 Pro<\/td><td>Third-party editor<\/td><td>1 of 71<\/td><td>not reviewed<\/td><td>2048&#215;1152<\/td><td>$0.060<\/td><td>7<\/td><\/tr><tr><td>Flux 2 Turbo<\/td><td>Third-party editor<\/td><td>1 of 71<\/td><td>not reviewed<\/td><td>2048&#215;1152<\/td><td>$0.048<\/td><td>7<\/td><\/tr><tr><td>Grok Imagine Image 2.0<\/td><td>Third-party editor<\/td><td>0 of 71<\/td><td>not reviewed<\/td><td>2816&#215;1584<\/td><td>$0.060<\/td><td>7<\/td><\/tr><tr><td>HiDream O1<\/td><td>Third-party editor<\/td><td>0 of 71<\/td><td>not reviewed<\/td><td>2560&#215;1440<\/td><td>$0.040<\/td><td>7<\/td><\/tr><tr><td>Ideogram v4<\/td><td>Third-party editor<\/td><td>0 of 3<\/td><td>0 of 3 (diagnostic images only)<\/td><td>2720&#215;1536<\/td><td>$0.100<\/td><td>1<\/td><\/tr><tr><td>Kling Image O3<\/td><td>Third-party editor<\/td><td>0 of 3<\/td><td>not reviewed<\/td><td>2720&#215;1536<\/td><td>$0.028<\/td><td>1<\/td><\/tr><tr><td>Kling Image V3<\/td><td>Third-party editor<\/td><td>0 of 3<\/td><td>not reviewed<\/td><td>2720&#215;1536<\/td><td>$0.028<\/td><td>1<\/td><\/tr><tr><td>Luma Uni<\/td><td>Third-party editor<\/td><td>15 of 61<\/td><td>not reviewed<\/td><td>2784&#215;1504<\/td><td>$0.042<\/td><td>6<\/td><\/tr><tr><td>MAI Image 2.5<\/td><td>Third-party editor<\/td><td>5 of 33<\/td><td>not reviewed<\/td><td>1360&#215;768<\/td><td>by prompt length<\/td><td>4<\/td><\/tr><tr><td>Nano Banana<\/td><td>Third-party editor<\/td><td>1 of 33<\/td><td>not reviewed<\/td><td>1344&#215;768<\/td><td>$0.038<\/td><td>4<\/td><\/tr><tr><td>Nano Banana 2 Fast<\/td><td>Third-party editor<\/td><td>30 of 71<\/td><td>43 of 71<\/td><td>2741&#215;1530<\/td><td>$0.045<\/td><td>7<\/td><\/tr><tr><td>Nano Banana 2 Lite<\/td><td>Third-party editor<\/td><td>14 of 71<\/td><td>not reviewed<\/td><td>1376&#215;768<\/td><td>$0.040<\/td><td>7<\/td><\/tr><tr><td>Qwen Image 3.0<\/td><td>Refused every slide<\/td><td>n\/a<\/td><td>n\/a<\/td><td>n\/a<\/td><td>$0.033<\/td><td>0<\/td><\/tr><tr><td>Qwen Image 3.0 Pro<\/td><td>Refused every slide<\/td><td>n\/a<\/td><td>n\/a<\/td><td>n\/a<\/td><td>$0.075<\/td><td>0<\/td><\/tr><tr><td>Qwen Image Edit Plus<\/td><td>Third-party editor<\/td><td>0 of 71<\/td><td>0 of 30 (diagnostic images only)<\/td><td>1360&#215;768<\/td><td>$0.020<\/td><td>7<\/td><\/tr><tr><td>Seedream 4.0<\/td><td>Third-party editor<\/td><td>4 of 71<\/td><td>0 of 10 (diagnostic images only)<\/td><td>2560&#215;1440<\/td><td>$0.027<\/td><td>7<\/td><\/tr><tr><td>Seedream 4.5<\/td><td>Third-party editor<\/td><td>12 of 71<\/td><td>not reviewed<\/td><td>2560&#215;1440<\/td><td>$0.040<\/td><td>7<\/td><\/tr><tr><td>Seedream 5.0 Lite<\/td><td>Third-party editor<\/td><td>0 of 71<\/td><td>not reviewed<\/td><td>2560&#215;1440<\/td><td>$0.035<\/td><td>7<\/td><\/tr><tr><td>Seedream 5.0 Pro<\/td><td>Third-party editor<\/td><td>16 of 71<\/td><td>not reviewed<\/td><td>2730&#215;1536<\/td><td>$0.090<\/td><td>7<\/td><\/tr><tr><td>Step1X Edit<\/td><td>Third-party editor<\/td><td>0 of 71<\/td><td>0 of 48 (diagnostic images only)<\/td><td>1392&#215;752<\/td><td>$0.030<\/td><td>7<\/td><\/tr><tr><td>Wan 2.7<\/td><td>Third-party editor<\/td><td>8 of 71<\/td><td>not reviewed<\/td><td>2560&#215;1440<\/td><td>$0.030<\/td><td>7<\/td><\/tr><\/tbody><\/table><figcaption class=\"wp-element-caption\">First generation, one per slide. The automatic column is the original OCR measurement, not human-verified accuracy; a person counted lines across all slides only for the three models with a full count. Four others (Step1X Edit, Qwen Image Edit Plus, Seedream 4.0 and Ideogram v4) were seen only as diagnostic images picked because they had failed, so their counts are not a sample. Inside each group the order is alphabetical, not a ranking: third-party editors moved past each other between runs. A count out of fewer than 71 lines means the model covered fewer than 7 slides. Measured by Wonderslide, September 2026.<\/figcaption><\/figure>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"653\" src=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-price-vs-accuracy-automatic-ocr-1024x653.webp\" alt=\"Scatter plot of price against automatic OCR exact lines for 27 image models, round 1\" class=\"wp-image-1408\" srcset=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-price-vs-accuracy-automatic-ocr-1024x653.webp 1024w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-price-vs-accuracy-automatic-ocr-300x191.webp 300w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-price-vs-accuracy-automatic-ocr-768x490.webp 768w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-price-vs-accuracy-automatic-ocr-1536x980.webp 1536w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-price-vs-accuracy-automatic-ocr-1568x1001.webp 1568w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-price-vs-accuracy-automatic-ocr.webp 1600w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Automatic OCR (uncorrected), price vs. accuracy, round 1: 5 of 27 participants clear half the lines, at $0.21 to $2.87. A person counted lines for only three of these models.<\/em><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">GPT Image 2 vs 2.5: low, medium and high tiers on small text<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">If you already use GPT Image, the real choice is between versions and quality tiers. The pair that matters most to us is GPT Image 2 medium, the tier our product runs on, against GPT Image 2.5 low, the newer version&#8217;s cheapest tier. In the human review both had all 71 lines accepted in the first generation. The scores below are automatic; the tier comparisons use three runs each, except the top tier, which was run once.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On text quality the two tiers do not separate. Each was run over the seven cases three times. The tier in production returned 56, 62 and 62 exact lines out of 71; the cheaper tier returned 61, 62 and 61. The tier in production moves across six lines on its own, which is more than anything that separates the two.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The two tiers were not drawn on the same canvas either: the tier in production returned the smaller 2048&#215;1152 on 19 of its 21 generations, the cheaper tier 2560&#215;1440 on all 21.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Unequal frames are a condition to report, not a proven cause of the automatic line gap. Human review accepted every body-text line of both compared GPT tiers in the first run. The price gap is independent of frame size and recognition.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On the repeated tiers the two metrics part company, and we say so rather than picking the flattering one: on the median of three repeats GPT Image 2 medium has the lower character-error rate, 0.0225 against 0.0546, while GPT Image 2.5 low has one more exact line, 62 against 61 out of 71.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Paying above the tier in production buys nothing measurable: counted the same way as the medium tier, GPT Image 2&#8217;s top tier reproduced 58 lines against 56 out of 71 while costing 3.7 times as much. Those 58 and 56 are one generation per case, and the frames differ (see the first table).<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"609\" src=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-quality-ladder-automatic-ocr-1024x609.webp\" alt=\"Bar chart of automatic OCR exact lines for the five GPT Image tiers with the price of each, round 1\" class=\"wp-image-1409\" srcset=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-quality-ladder-automatic-ocr-1024x609.webp 1024w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-quality-ladder-automatic-ocr-300x179.webp 300w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-quality-ladder-automatic-ocr-768x457.webp 768w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-quality-ladder-automatic-ocr-1536x914.webp 1536w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-quality-ladder-automatic-ocr-1568x933.webp 1568w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-quality-ladder-automatic-ocr.webp 1600w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Automatic OCR, quality ladder, round 1: production tier is 3.1x recommended, top tier is 3.7x production. In plain terms, GPT Image 2 medium costs 3.1 times GPT Image 2.5 low, and GPT Image 2 high costs 3.7 times GPT Image 2 medium.<\/em><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What we recommend<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The run-to-run numbers in this section are automatic scores. The human review covers the first generation only.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For this job, move down a version rather than up a tier, and move for the price, not for the text. The newer version&#8217;s low tier costs a third of what the tier in production costs, $0.245 against $0.770 over the seven cases; its medium tier costs half, $0.385. Run three times each, the three tiers land between 56 and 62 exact lines out of 71 and this study cannot separate them on text quality. On the median of the three repeats the tier in production is the one with the cleanest characters (0.0225, against 0.0546 for the low tier and 0.0490 for the medium one), and that is the caveat the move carries.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The cheapest tier in the study is not the one we recommend. GPT Image 2 low costs $0.210 over the seven cases against GPT Image 2.5 low&#8217;s $0.245, and it is the one tier whose text this study can tell apart from the tier in production: over three runs of the seven cases it returned 52, 49 and 51 exact lines out of 71, where the tier in production returned 56, 62 and 62.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The price gap does not depend on text recognition: it is list-price arithmetic from September 2026. Before you switch, test the new tier on your own slides. The human review covers one generation of each slide, and the repeats were scored automatically.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Steadiness between identical runs<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The same prompt does not always give the same slide. These are automatic scores, and one check by a person, described at the end of this section, shows that they can be off. Here the newer version is GPT Image 2.5 and the older one is GPT Image 2.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The newer version&#8217;s medium tier is far steadier between identical runs than the older version&#8217;s: across three runs of the same slide it never moved by more than 1 exact line on any of its 7 cases, where the older version&#8217;s medium tier gave 8, 16 and 13 exact lines on the same 18-line slide, an 8-line swing, 44% of the slide.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Steadiness is a property of a tier, not of a version: the newer version&#8217;s <em>low<\/em> tier moved by 2 lines on 3 of its 7 cases, and the older version&#8217;s low tier by no more than 3 on any of its.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A person checked the first generation of the older version&#8217;s medium tier on the 18-line slide and accepted all 18 lines, where automatic scoring counted 8. So the 8-line swing is partly a scoring effect; the other two runs were not reviewed.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Can AI image models write Cyrillic?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">For GPT Image 2.5 low and GPT Image 2 medium, yes, on this template. It is Russian text on a timeline, so this tests those two tiers, not the alphabet in general. In the human review, each had all 10 lines accepted on the Cyrillic and on the Latin version of the same 10 events, where automatic scoring counted 7 and 9 of 10 on the Cyrillic version. For GPT Image 2.5 low, the Cyrillic gap in the automatic scores (7 of 10 against 10 of 10 on the Latin version) comes from the scoring, not from the images, in this first generation.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The automatic scores over all runs show a gap instead. &#8216;These models can&#8217;t do Cyrillic&#8217; does not survive measurement. But Cyrillic is not free either. On the same 10 events written in Russian and in English, the five GPT tiers missed 18 Cyrillic lines against 4 Latin ones, over the identical 13 runs each version of the slide received.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A third-party editor that struggles with small text struggles more with Cyrillic than with Latin on the exact same content: 5 of 10 lines against 8 of 10. Those two numbers are Nano Banana 2 Fast on the median of three repeats. In the human review of its first generation, a person accepted 2 of 10 Cyrillic lines and 10 of 10 Latin lines (on that same generation automatic scoring counted 1 of 10 and 8 of 10), so for this model the gap is in the images too.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"463\" src=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-cyrillic-vs-latin-automatic-ocr-1024x463.webp\" alt=\"Grouped bar chart of automatic OCR exact lines on the Cyrillic and the Latin version of the same 10-event slide, round 1\" class=\"wp-image-1410\" srcset=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-cyrillic-vs-latin-automatic-ocr-1024x463.webp 1024w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-cyrillic-vs-latin-automatic-ocr-300x136.webp 300w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-cyrillic-vs-latin-automatic-ocr-768x347.webp 768w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-cyrillic-vs-latin-automatic-ocr-1536x694.webp 1536w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-cyrillic-vs-latin-automatic-ocr-1568x709.webp 1568w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-cyrillic-vs-latin-automatic-ocr.webp 1600w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Automatic OCR, Cyrillic vs. Latin, same 10 events, round 1: no finalist scores higher on Cyrillic; gap 0 to 70 points.<\/em><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Does more text on a slide break the models?<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Short answer: in the one case a person checked, no. Automatic scores show one GPT tier falling off at 18 blocks, and the review shows that this drop was a scoring effect in the image checked.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">We compared the 10-block and the 18-block slides, whose briefs are close in length, so the block count is what changes. The scores in this paragraph and the next are automatic, and frames differ between tiers; see the table above.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Only one tier of the five falls off at 18 blocks, and it is the tier in production. Round 1, the share of exact lines on the densest case: GPT Image 2.5 low 89%, GPT Image 2.5 medium 83%, GPT Image 2 low 83%, GPT Image 2 high 67%, GPT Image 2 medium 44%. Four of the five land between 67% and 89%; the fifth is at 44%.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">That tier, GPT Image 2 medium, fell off on a slide where three identical runs returned 8, 16 and 13 exact lines of 18. Round 1 drew the 8, and round 1 is what the bars plot. On the median of the same three runs that cell is 13 of 18 (72%), back inside the range the other four tiers occupy, and nothing is left to call a break.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A person checked that first generation and accepted all 18 lines, where automatic scoring counted 8. In this image the drop is a scoring effect. For Nano Banana 2 Fast the person accepted 1 of 18 lines on the same slide, so there the drop is real.<\/p>\n\n\n\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"516\" src=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-text-density-automatic-ocr-1024x516.webp\" alt=\"Line charts of automatic OCR exact-line share against 3, 10 and 18 timeline blocks for the eight finalists, round 1\" class=\"wp-image-1411\" srcset=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-text-density-automatic-ocr-1024x516.webp 1024w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-text-density-automatic-ocr-300x151.webp 300w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-text-density-automatic-ocr-768x387.webp 768w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-text-density-automatic-ocr-1536x774.webp 1536w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-text-density-automatic-ocr-1568x790.webp 1568w, https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-text-density-automatic-ocr.webp 1600w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><em>Automatic OCR, text accuracy vs. block count, round 1: at 18 blocks every GPT tier holds at least 44% of the lines, the rest at most 17%. By eye, GPT Image 2 medium had 18 of 18 on the 18-block slide, where this chart shows 44%.<\/em><\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Limitations<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">What this study did not measure, and why it matters for reading the numbers:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>One API intermediary.<\/strong> Every model was called through one intermediary, on the catalog snapshot of 2026-09-17. Prices come from its price list, not from invoices. Results may differ through another channel, a vendor&#8217;s own product or a later catalog.<\/li>\n\n\n<li><strong>The human review is partial.<\/strong> A person reviewed body text in 30 original images: the first generation of three models, plus 9 flagged images. Other models, later runs, image aesthetics and titles were not reviewed. The reviewer saw the reference text and enlarged crops, so the review was not blind.<\/li>\n\n\n<li><strong>Automatic scoring is not corrected.<\/strong> It disagreed with the review on about one judged line in five, from 14.08% to 30.16% depending on the model. No pure recognizer error rate was measured, and every automatic score in this article is the original measurement.<\/li>\n\n\n<li><strong>Three runs show spread, not significance.<\/strong> The finalists were run up to three times per slide, and the repeats were scored automatically only. GPT Image 2 high was run once per slide.<\/li>\n\n\n<li><strong>One slide template.<\/strong> All seven slides are synthetic horizontal timelines with a year at the top of each column, and the 3-block slide also has a much shorter brief than the 10-block slide. The results are about this template, not about infographics in general.<\/li>\n\n\n<li><strong>Unequal frames and sizes.<\/strong> Returned frames ranged from 1024&#215;576 to 2816&#215;1584, 2.75 times in linear size. Small text suffers on a small canvas, so every line count is a count at a frame. The written brief asked for 1280&#215;720 and the request asked for about 2560&#215;1440. No generation came back at 1280&#215;720, and the two instructions contradict each other, so a returned size that differs from the request cannot be read as disobedience.<\/li>\n\n\n<li><strong>Not every model took every slide.<\/strong> 27 of 29 participants produced a measurable result. Two refused all seven slides because the brief exceeded their prompt-length limit, and seven more covered fewer than seven slides. A refusal does not say a model draws text worse; refused models were not re-run on a shortened brief.<\/li>\n\n\n<li><strong>What the automatic scorer cannot see.<\/strong> It treats look-alike Latin and Cyrillic letters as the same, ignores letter case and drops the standalone dash between a year and an event. A model that writes a Latin &quot;c&quot; inside a Russian word scores the same as one that writes the Cyrillic &quot;\u0441&quot;.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The reassembly step our scoring depends on is published with its reliability indicator rather than asserted: 2.9% of ground-truth lines (47 of 1617) had their words scattered across unrelated reconstructed blocks, against a 10% threshold fixed before the recount. This indicator describes the reassembly step, not recognition accuracy.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The full limitations text is in <a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-methodology-notes.txt\">methodology-notes.txt<\/a>. The human review is in <a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-manual-calibration.txt\">manual-calibration.txt<\/a> and <a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-line-audit.csv\">line-audit.csv<\/a>, and the check of the measurer is in <a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-ocr-calibration.txt\">ocr-calibration.txt<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Test prompts, results and raw data<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Everything behind the numbers above is published. Tables and text are released under CC BY 4.0; the images are for research and evaluation use only. The full terms are in the license file.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table class=\"has-fixed-layout\"><thead><tr><th>File<\/th><th>What is inside<\/th><th>License<\/th><\/tr><\/thead><tbody><tr><td><a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/09\/infographic-study-results.csv\">results.csv<\/a><\/td><td>One row per case, participant and run (294 rows: 250 with an image, 44 refusals or platform failures): model, tier, status, price, time, returned frame, both automatic text metrics, what the recognizer read, image file name<\/td><td>CC BY 4.0<\/td><\/tr><tr><td><a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-line-audit.csv\">line-audit.csv<\/a><\/td><td>The human review: 304 expected lines with the person&#8217;s verdict, the automatic result, what the recognizer read and the image file<\/td><td>CC BY 4.0<\/td><\/tr><tr><td><a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/09\/infographic-study-participants.csv\">participants.csv<\/a><\/td><td>The 29 participants: model endpoint, request parameters, price per call, prompt-length limit<\/td><td>CC BY 4.0<\/td><\/tr><tr><td><a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/09\/infographic-study-cases.csv\">cases.csv<\/a><\/td><td>The 7 test slides: what each one varies, the exact prompt, the reference title and lines<\/td><td>CC BY 4.0<\/td><\/tr><tr><td><a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/09\/infographic-study-ocr-detections.csv\">ocr-detections.csv<\/a><\/td><td>Everything the recognizer detected on the 250 images, one row per piece of text, with its position and confidence. An image with no text found has one row with empty text<\/td><td>CC BY 4.0<\/td><\/tr><tr><td><a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-manual-calibration.txt\">manual-calibration.txt<\/a><\/td><td>The human review: selection, results, disagreement with automatic scoring and its limits<\/td><td>CC BY 4.0<\/td><\/tr><tr><td><a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-methodology-notes.txt\">methodology-notes.txt<\/a><\/td><td>The full limitations text: what was not measured and why<\/td><td>CC BY 4.0<\/td><\/tr><tr><td><a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-ocr-calibration.txt\">ocr-calibration.txt<\/a><\/td><td>What was checked about the measurer&#8217;s reassembly step<\/td><td>CC BY 4.0<\/td><\/tr><tr><td><a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/09\/infographic-study-images.zip\">images.zip<\/a><\/td><td>The 7 input slides and the 250 generated slides (previews 1100 px wide), named by case, model and run<\/td><td>Research and evaluation only<\/td><\/tr><tr><td><a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-report.html\">report.html<\/a><\/td><td>All 250 generations on one page, next to their automatic scores and recognized text (20 MB)<\/td><td>Research and evaluation only<\/td><\/tr><tr><td><a href=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-license.txt\">LICENSE.txt<\/a><\/td><td>License terms<\/td><td><\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">How to cite: Wonderslide (2026). Small-text infographic generation: 29 image models on 7 timeline slides. https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/<\/p>\n\n\n\n<details class=\"wp-block-details is-layout-flow wp-block-details-is-layout-flow\"><summary>Columns in results.csv<\/summary>\n<ul class=\"wp-block-list\">\n<li><strong>case<\/strong>: the test slide, D3 to I10 (see cases.csv).<\/li>\n\n\n<li><strong>participant<\/strong>: the model and tier; participants.csv maps it to the endpoint and request parameters.<\/li>\n\n\n<li><strong>endpoint<\/strong>, <strong>tier<\/strong>: the model called and the parameters sent with the request.<\/li>\n\n\n<li><strong>repeat<\/strong>: 1 for round 1; 2 and 3 for the repeated runs of the finalists.<\/li>\n\n\n<li><strong>seed<\/strong>: the seed sent with the request, where the model accepts one.<\/li>\n\n\n<li><strong>status<\/strong>: completed, capability_refusal (the model declined, usually because the brief exceeded its prompt-length limit) or unavailable (the platform failed without an image).<\/li>\n\n\n<li><strong>width<\/strong>, <strong>height<\/strong>, <strong>frame<\/strong>: the size of the image the model returned.<\/li>\n\n\n<li><strong>seconds<\/strong>: generation time.<\/li>\n\n\n<li><strong>price_usd<\/strong>: the list price of this attempt at the API intermediary. Only the final attempt for each cell is in the table, so the column does not add up to the total spend on retries.<\/li>\n\n\n<li><strong>error<\/strong>: the platform&#8217;s message when a request failed.<\/li>\n\n\n<li><strong>lines_exact<\/strong>: reference lines the automatic scoring counted as read back with no character difference. This is not a human judgment; compare line-audit.csv for the images a person reviewed.<\/li>\n\n\n<li><strong>lines_total<\/strong>: reference lines on the slide.<\/li>\n\n\n<li><strong>char_error_rate<\/strong>: the automatic average, over reference lines, of edit distance divided by line length, capped at 1; a line with no match counts as 1.<\/li>\n\n\n<li><strong>blocks_found<\/strong>: reference lines where at least half of the words were found in the matched text.<\/li>\n\n\n<li><strong>invented_words<\/strong>: words the recognizer read that are not in the reference text (any number, or a word of five letters or more that is not a near-typo of an expected word).<\/li>\n\n\n<li><strong>title_preserved<\/strong>: 1 when the title is intact, lower as the title&#8217;s edit distance grows.<\/li>\n\n\n<li><strong>reordered_line<\/strong>: true when some line&#8217;s words all sit in its own block but not as one in-order run.<\/li>\n\n\n<li><strong>scattered_lines<\/strong>: reference lines whose words are all on the slide but never inside one reconstructed block.<\/li>\n\n\n<li><strong>blocks_detected<\/strong>: text blocks reconstructed from the image.<\/li>\n\n\n<li><strong>ocr_text<\/strong>, <strong>ocr_title_text<\/strong>: what the recognizer read in the body and in the title band.<\/li>\n\n\n<li><strong>image_file<\/strong>: the image&#8217;s path inside images.zip.<\/li>\n<\/ul>\n<\/details>\n\n\n\n<details class=\"wp-block-details is-layout-flow wp-block-details-is-layout-flow\"><summary>Columns in line-audit.csv<\/summary>\n<ul class=\"wp-block-list\">\n<li><strong>case<\/strong>, <strong>participant<\/strong>, <strong>repeat<\/strong>: the reviewed generation, the same keys as in results.csv. <strong>image_id<\/strong> joins them.<\/li>\n\n\n<li><strong>sample<\/strong>: stratified (the 21 images of the three compared models) or flagged (the 9 diagnostic images).<\/li>\n\n\n<li><strong>year<\/strong>: the year that identifies the reference line on the slide.<\/li>\n\n\n<li><strong>image_sha256<\/strong>: the hash of the original image the person reviewed.<\/li>\n\n\n<li><strong>expected<\/strong>: the reference line.<\/li>\n\n\n<li><strong>human_verdict<\/strong>: exact (accepted), different (readable but not the same), unreadable or missing (absent).<\/li>\n\n\n<li><strong>automatic_exact<\/strong>: whether automatic scoring counted the line as exact.<\/li>\n\n\n<li><strong>comparison<\/strong>: true_positive, false_negative (a person accepted, automatic scoring did not), false_positive (automatic scoring accepted, a person did not), true_negative, true_negative (neither a person nor automatic scoring accepted the line), or excluded_unreadable.<\/li>\n\n\n<li><strong>visible_text<\/strong>: what the person read, where recorded.<\/li>\n\n\n<li><strong>automatic_candidate<\/strong>, <strong>ocr_blocks<\/strong>: the text block the scorer matched to the line, and the blocks it reconstructed.<\/li>\n\n\n<li><strong>note<\/strong>: the reviewer&#8217;s note, where there is one.<\/li>\n\n\n<li><strong>image_file<\/strong>: the image&#8217;s path inside images.zip (a preview; the hash refers to the original).<\/li>\n<\/ul>\n<\/details>\n\n\n\n<script type=\"application\/ld+json\">{\"@context\": \"https:\/\/schema.org\", \"@type\": \"Dataset\", \"name\": \"Small-text infographic generation: 29 image models on 7 timeline slides\", \"description\": \"Raw data from a Wonderslide study of how image models reproduce small text on timeline infographic slides: 29 participants (24 third-party image editors and five GPT Image tiers), 7 synthetic test slides in Cyrillic and Latin script, 250 generations scored automatically, and a person's line-by-line review of 30 original images (304 lines). Includes one row per generation attempt with price, returned frame, automatic exact-line and character-error metrics and recognized text, the human verdict for every reviewed line, the test prompts and reference text, the recognizer output, and the generated images.\", \"url\": \"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/\", \"creator\": {\"@type\": \"Organization\", \"name\": \"Wonderslide\", \"url\": \"https:\/\/wonderslide.com\/\"}, \"dateModified\": \"2026-10-01\", \"temporalCoverage\": \"2026-09\/2026-10\", \"license\": \"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-license.txt\", \"isAccessibleForFree\": true, \"keywords\": [\"image generation\", \"text rendering\", \"infographics\", \"GPT Image 2\", \"GPT Image 2.5\", \"Nano Banana 2\", \"OCR\", \"Cyrillic\"], \"variableMeasured\": [\"lines reproduced exactly (automatic OCR scoring)\", \"character error rate (automatic OCR scoring)\", \"lines accepted by a person (human review of 30 images)\", \"price per generation\", \"returned frame size\"], \"distribution\": [{\"@type\": \"DataDownload\", \"name\": \"results.csv\", \"encodingFormat\": \"text\/csv\", \"contentUrl\": \"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/09\/infographic-study-results.csv\"}, {\"@type\": \"DataDownload\", \"name\": \"line-audit.csv\", \"encodingFormat\": \"text\/csv\", \"contentUrl\": \"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-study-line-audit.csv\"}, {\"@type\": \"DataDownload\", \"name\": \"participants.csv\", \"encodingFormat\": \"text\/csv\", \"contentUrl\": \"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/09\/infographic-study-participants.csv\"}, {\"@type\": \"DataDownload\", \"name\": \"cases.csv\", \"encodingFormat\": \"text\/csv\", \"contentUrl\": \"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/09\/infographic-study-cases.csv\"}, {\"@type\": \"DataDownload\", \"name\": \"ocr-detections.csv\", \"encodingFormat\": \"text\/csv\", \"contentUrl\": \"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/09\/infographic-study-ocr-detections.csv\"}, {\"@type\": \"DataDownload\", \"name\": \"images.zip\", \"encodingFormat\": \"application\/zip\", \"contentUrl\": \"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/09\/infographic-study-images.zip\"}]}<\/script>\n\n\n\n<h2 class=\"wp-block-heading\">FAQ<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\">Is GPT Image 2.5 low better for text than GPT Image 2 medium?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Not measurably. A person accepted all 71 lines from both in the first generation. Automatic scoring over three runs each cannot separate them either, nor GPT Image 2.5 medium: 56 to 62 exact lines out of 71. GPT Image 2.5 low costs a third as much.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">GPT Image 2 low vs medium vs high: which quality setting should you use for text?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">On automatic scoring, GPT Image 2 low is the one tier this study could tell apart from GPT Image 2 medium: over three runs it returned 52, 49 and 51 exact lines out of 71, against 56, 62 and 62. High bought nothing measurable over medium on one generation per case: 58 lines against 56, at 3.7 times the price.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">How does GPT Image 2 compare with Nano Banana 2 on small text?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">A person reviewed the first generation of both across the same seven slides. GPT Image 2 medium had 71 of 71 lines accepted and Nano Banana 2 Fast 43 of 71, at frames of 2048&#215;1152 and 2741&#215;1530. This covers these reviewed outputs only; it is not a ranking of the market or a guarantee for later generations.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Can AI image models write Cyrillic text?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">For GPT Image 2.5 low and GPT Image 2 medium, yes: a person accepted all 10 Cyrillic lines in the first generation, where automatic scoring counted 7 and 9. Nano Banana 2 Fast was weaker: 2 of 10 Cyrillic lines accepted against 10 of 10 Latin. Cyrillic is harder for some models, not impossible.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">How accurate is the automatic scoring of text in AI images?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">We compared it with a person&#8217;s review of 30 original images. On 205 judgeable lines it disagreed on 44, or 21.46%, mostly lines it marked wrong that the person accepted. That measures the whole scoring pipeline, not the recognizer alone, and it differs by model, so it is not a correction factor for other scores.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Build infographic slides from your own data<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The slides in this study were synthetic. Yours will not be. Wonderslide turns your document into a designed presentation, with timelines, charts and processes drawn from your data. Try it on your own deck. Upload a PDF, DOCX or PPTX and get a designed presentation in a few minutes. Free plan, no card required.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">If you are comparing presentation tools rather than image models, see <a href=\"https:\/\/wonderslide.com\/blog\/ai-presentation-maker-cost\/\">how much AI presentation makers cost<\/a> and <a href=\"https:\/\/wonderslide.com\/blog\/how-to-create-visuals-for-presentation-key-points\/\">how to create visuals for your presentation&#8217;s key points<\/a>.<\/p>\n\n\n\n<div class=\"wp-block-buttons is-layout-flex wp-block-buttons-is-layout-flex\">\n<div class=\"wp-block-button\"><a class=\"wp-block-button__link wp-element-button\" href=\"https:\/\/wonderslide.com\/?ref=blog-infographic-small-text-image-models\">Try Wonderslide free<\/a><\/div>\n<\/div>\n\n\n\n\n<style id=\"ws-blog-fix\">\n.entry-content .wp-block-table table{border-collapse:collapse;width:100%;margin:1em 0}\n.entry-content .wp-block-table th,.entry-content .wp-block-table td{border:1px solid rgba(128,128,128,.45);padding:.5em .75em;text-align:left;vertical-align:top}\n.entry-content .wp-block-table thead th{font-weight:600;border-bottom-width:2px}\n.entry-content .wp-block-table figcaption{margin-top:.5em;font-size:.875em;opacity:.8}\n.entry-content .wp-block-buttons{display:flex;flex-wrap:wrap;gap:.75em;margin:1.5em 0}\n.entry-content .wp-block-button__link{display:inline-block}\n<\/style>\n","protected":false},"excerpt":{"rendered":"<p>29 image models, 7 timeline slides. A person checked the original images of three models, and every GPT Image tier is priced. Includes the raw data.<\/p>\n","protected":false},"author":2,"featured_media":1422,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-1400","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-default","entry"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.5 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>GPT Image 2 vs 2.5 for Text: Low, Medium, High Tested - Wonderslide Blog<\/title>\n<meta name=\"description\" content=\"We tested 29 image models on small-text slides. In a human check of first-generation images, GPT Image 2 medium and 2.5 low got 71 of 71 lines, Nano Banana 2 Fast 43.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"GPT Image 2 vs 2.5 for Text: Low, Medium, High Tested - Wonderslide Blog\" \/>\n<meta property=\"og:description\" content=\"We tested 29 image models on small-text slides. In a human check of first-generation images, GPT Image 2 medium and 2.5 low got 71 of 71 lines, Nano Banana 2 Fast 43.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/\" \/>\n<meta property=\"og:site_name\" content=\"Wonderslide Blog\" \/>\n<meta property=\"article:published_time\" content=\"2026-10-01T11:11:42+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-10-01T19:16:58+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-small-text-image-models-feature.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"1200\" \/>\n\t<meta property=\"og:image:height\" content=\"628\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"Dmitrii Glazunov\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Dmitrii Glazunov\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"24 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/\"},\"author\":{\"name\":\"Dmitrii Glazunov\",\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/#\\\/schema\\\/person\\\/ce2fa06767be88fa92f1fa0fe08b0117\"},\"headline\":\"Which Image Model Can Draw an Infographic With Small Text? We Tested 29\",\"datePublished\":\"2026-10-01T11:11:42+00:00\",\"dateModified\":\"2026-10-01T19:16:58+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/\"},\"wordCount\":4616,\"publisher\":{\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/infographic-small-text-image-models-feature.webp\",\"articleSection\":[\"Default\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/\",\"url\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/\",\"name\":\"GPT Image 2 vs 2.5 for Text: Low, Medium, High Tested - Wonderslide Blog\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/infographic-small-text-image-models-feature.webp\",\"datePublished\":\"2026-10-01T11:11:42+00:00\",\"dateModified\":\"2026-10-01T19:16:58+00:00\",\"description\":\"We tested 29 image models on small-text slides. In a human check of first-generation images, GPT Image 2 medium and 2.5 low got 71 of 71 lines, Nano Banana 2 Fast 43.\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/#primaryimage\",\"url\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/infographic-small-text-image-models-feature.webp\",\"contentUrl\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/wp-content\\\/uploads\\\/2026\\\/10\\\/infographic-small-text-image-models-feature.webp\",\"width\":1200,\"height\":628},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/infographic-small-text-image-models\\\/#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Which Image Model Can Draw an Infographic With Small Text? We Tested 29\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/\",\"name\":\"Wonderslide Blog\",\"description\":\"The Wonderslide Blog: AI, Presentations &amp; Creative Impact\",\"publisher\":{\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/#organization\",\"name\":\"Wonderslide\",\"url\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/wp-content\\\/uploads\\\/2022\\\/12\\\/W-black-logo.png\",\"contentUrl\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/wp-content\\\/uploads\\\/2022\\\/12\\\/W-black-logo.png\",\"width\":1500,\"height\":182,\"caption\":\"Wonderslide\"},\"image\":{\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/#\\\/schema\\\/logo\\\/image\\\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/#\\\/schema\\\/person\\\/ce2fa06767be88fa92f1fa0fe08b0117\",\"name\":\"Dmitrii Glazunov\",\"url\":\"https:\\\/\\\/wonderslide.com\\\/blog\\\/author\\\/d-glazunov\\\/\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"GPT Image 2 vs 2.5 for Text: Low, Medium, High Tested - Wonderslide Blog","description":"We tested 29 image models on small-text slides. In a human check of first-generation images, GPT Image 2 medium and 2.5 low got 71 of 71 lines, Nano Banana 2 Fast 43.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/","og_locale":"en_US","og_type":"article","og_title":"GPT Image 2 vs 2.5 for Text: Low, Medium, High Tested - Wonderslide Blog","og_description":"We tested 29 image models on small-text slides. In a human check of first-generation images, GPT Image 2 medium and 2.5 low got 71 of 71 lines, Nano Banana 2 Fast 43.","og_url":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/","og_site_name":"Wonderslide Blog","article_published_time":"2026-10-01T11:11:42+00:00","article_modified_time":"2026-10-01T19:16:58+00:00","og_image":[{"width":1200,"height":628,"url":"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-small-text-image-models-feature.webp","type":"image\/webp"}],"author":"Dmitrii Glazunov","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Dmitrii Glazunov","Est. reading time":"24 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/#article","isPartOf":{"@id":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/"},"author":{"name":"Dmitrii Glazunov","@id":"https:\/\/wonderslide.com\/blog\/#\/schema\/person\/ce2fa06767be88fa92f1fa0fe08b0117"},"headline":"Which Image Model Can Draw an Infographic With Small Text? We Tested 29","datePublished":"2026-10-01T11:11:42+00:00","dateModified":"2026-10-01T19:16:58+00:00","mainEntityOfPage":{"@id":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/"},"wordCount":4616,"publisher":{"@id":"https:\/\/wonderslide.com\/blog\/#organization"},"image":{"@id":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/#primaryimage"},"thumbnailUrl":"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-small-text-image-models-feature.webp","articleSection":["Default"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/","url":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/","name":"GPT Image 2 vs 2.5 for Text: Low, Medium, High Tested - Wonderslide Blog","isPartOf":{"@id":"https:\/\/wonderslide.com\/blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/#primaryimage"},"image":{"@id":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/#primaryimage"},"thumbnailUrl":"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-small-text-image-models-feature.webp","datePublished":"2026-10-01T11:11:42+00:00","dateModified":"2026-10-01T19:16:58+00:00","description":"We tested 29 image models on small-text slides. In a human check of first-generation images, GPT Image 2 medium and 2.5 low got 71 of 71 lines, Nano Banana 2 Fast 43.","breadcrumb":{"@id":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/#primaryimage","url":"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-small-text-image-models-feature.webp","contentUrl":"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2026\/10\/infographic-small-text-image-models-feature.webp","width":1200,"height":628},{"@type":"BreadcrumbList","@id":"https:\/\/wonderslide.com\/blog\/infographic-small-text-image-models\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/wonderslide.com\/blog\/"},{"@type":"ListItem","position":2,"name":"Which Image Model Can Draw an Infographic With Small Text? We Tested 29"}]},{"@type":"WebSite","@id":"https:\/\/wonderslide.com\/blog\/#website","url":"https:\/\/wonderslide.com\/blog\/","name":"Wonderslide Blog","description":"The Wonderslide Blog: AI, Presentations &amp; Creative Impact","publisher":{"@id":"https:\/\/wonderslide.com\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/wonderslide.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/wonderslide.com\/blog\/#organization","name":"Wonderslide","url":"https:\/\/wonderslide.com\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/wonderslide.com\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2022\/12\/W-black-logo.png","contentUrl":"https:\/\/wonderslide.com\/blog\/wp-content\/uploads\/2022\/12\/W-black-logo.png","width":1500,"height":182,"caption":"Wonderslide"},"image":{"@id":"https:\/\/wonderslide.com\/blog\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/wonderslide.com\/blog\/#\/schema\/person\/ce2fa06767be88fa92f1fa0fe08b0117","name":"Dmitrii Glazunov","url":"https:\/\/wonderslide.com\/blog\/author\/d-glazunov\/"}]}},"_links":{"self":[{"href":"https:\/\/wonderslide.com\/blog\/wp-json\/wp\/v2\/posts\/1400","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wonderslide.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wonderslide.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wonderslide.com\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/wonderslide.com\/blog\/wp-json\/wp\/v2\/comments?post=1400"}],"version-history":[{"count":7,"href":"https:\/\/wonderslide.com\/blog\/wp-json\/wp\/v2\/posts\/1400\/revisions"}],"predecessor-version":[{"id":1423,"href":"https:\/\/wonderslide.com\/blog\/wp-json\/wp\/v2\/posts\/1400\/revisions\/1423"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/wonderslide.com\/blog\/wp-json\/wp\/v2\/media\/1422"}],"wp:attachment":[{"href":"https:\/\/wonderslide.com\/blog\/wp-json\/wp\/v2\/media?parent=1400"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wonderslide.com\/blog\/wp-json\/wp\/v2\/categories?post=1400"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wonderslide.com\/blog\/wp-json\/wp\/v2\/tags?post=1400"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}