Note for readers of this dataset: this document comes from the study's working repository, which is not public. File paths and commands in it refer to that repository. The data it describes is included here: results.csv, participants.csv, cases.csv, ocr-detections.csv and line-audit.csv. # Human calibration of automatic infographic scoring The study owner compared visible words and numbers in the original generated images against the reference text. The raw, unchanged export is [manual_review.json](manual_review.json); the reproducible calculation is [calibration-audit.json](human-calibration/calibration-audit.json), and every line judgment is in [line-audit.csv](human-calibration/line-audit.csv). This implements the human-check request in the publication review. It supersedes the old text-against-text sample as evidence for measurement calibration. It does not replace the original automatic results. Check: review_count(participant=all, sample=all, field=images, basis=human_round1) = 30 ## Selection and completion All 30 requested images were reviewed: 21 stratified images and 9 diagnostic images, totaling 304 expected lines. The stratified sample covers all 7 cases for Nano Banana 2 Fast, GPT Image 2.5 low and GPT Image 2 medium, on the first generation only. These are 213 expected lines. The diagnostic sample is exactly the previously flagged set, with 91 expected lines. There are no unanswered line judgments. Check: review_count(participant=all, sample=all, field=images, basis=human_round1) = 30 Check: review_count(participant=all, sample=stratified, field=images, basis=human_round1) = 21 Check: review_count(participant=all, sample=flagged, field=images, basis=human_round1) = 9 Check: review_count(participant=all, sample=all, field=lines, basis=human_round1) = 304 Check: review_count(participant=all, sample=stratified, field=lines, basis=human_round1) = 213 Check: review_count(participant=all, sample=flagged, field=lines, basis=human_round1) = 91 Check: cases_measured(participant=gpt25_low) = 7 Context: 2 = a model version name Context: 2.5 = a model version name ## Human-accepted text in the stratified sample | Participant | Human exact | Readable deviation | Unreadable | Absent | Original automatic exact | |---|---:|---:|---:|---:|---:| | GPT Image 2 medium | 71 | 0 | 0 | 0 | 56 | | GPT Image 2.5 low | 71 | 0 | 0 | 0 | 61 | | Nano Banana 2 Fast | 43 | 19 | 8 | 1 | 30 | Each participant was given 71 expected lines across the same cases. Frames were not equal: Nano Banana returned 2741x1530, GPT Image 2.5 low 2560x1440, and GPT Image 2 medium 2048x1152 in this reviewed sweep. These are human judgments about words and numbers, not an independent transcription or a certification of capitalization, punctuation, aesthetics or title preservation. Unreadable and absent lines do not count as human exact. Check: review_count(participant=gpt2_medium, sample=stratified, field=exact, basis=human_round1) = 71 Check: review_count(participant=gpt25_low, sample=stratified, field=exact, basis=human_round1) = 71 Check: review_count(participant=nano_banana_2_fast, sample=stratified, field=exact, basis=human_round1) = 43 Check: review_count(participant=nano_banana_2_fast, sample=stratified, field=different, basis=human_round1) = 19 Check: review_count(participant=nano_banana_2_fast, sample=stratified, field=unreadable, basis=human_round1) = 8 Check: review_count(participant=nano_banana_2_fast, sample=stratified, field=missing, basis=human_round1) = 1 Check: review_count(participant=gpt2_medium, sample=stratified, field=different, basis=human_round1) = 0 Check: review_count(participant=gpt2_medium, sample=stratified, field=automatic_exact, basis=human_round1) = 56 Check: review_count(participant=gpt25_low, sample=stratified, field=automatic_exact, basis=human_round1) = 61 Check: review_count(participant=nano_banana_2_fast, sample=stratified, field=automatic_exact, basis=human_round1) = 30 Check: review_count(participant=gpt2_medium, sample=stratified, field=lines, basis=human_round1) = 71 Check: review_count(participant=gpt25_low, sample=stratified, field=lines, basis=human_round1) = 71 Check: review_count(participant=nano_banana_2_fast, sample=stratified, field=lines, basis=human_round1) = 71 Check: frames(participant=nano_banana_2_fast, frame=2741x1530, scope=round1) = 7 Check: frames(participant=gpt25_low, frame=2560x1440, scope=round1) = 7 Check: frames(participant=gpt2_medium, frame=2048x1152, scope=round1) = 7 Context: 2 = a model version name Context: 2.5 = a model version name ## Measured automatic-scoring disagreement Exclude the 8 unreadable stratified lines. Among 205 determinate judgments (human exact, readable deviation or absent), the automatic exact-line decision disagrees with the human judgment on 44 lines: **21.46%**. There are 41 false negatives (human exact, automatic not exact) and 3 false positives (human different, automatic exact). The full confusion table is 144 true positives and 17 true negatives, alongside those disagreements. This is the observed error of the complete automatic exact-line measurement against the owner's visual judgments. Recognition, geometric reconstruction, line assignment and normalization can all contribute. It is not a recognizer-only character error rate or an estimate for the entire market. Check: review_count(participant=all, sample=stratified, field=unreadable, basis=human_round1) = 8 Check: review_count(participant=all, sample=stratified, field=determinate, basis=human_round1) = 205 Check: review_count(participant=all, sample=stratified, field=disagreements, basis=human_round1) = 44 Check: review_pct(participant=all, sample=stratified, field=disagreements, basis=human_round1) = 21.46 Check: review_count(participant=all, sample=stratified, field=false_negative, basis=human_round1) = 41 Check: review_count(participant=all, sample=stratified, field=false_positive, basis=human_round1) = 3 Check: review_count(participant=all, sample=stratified, field=true_positive, basis=human_round1) = 144 Check: review_count(participant=all, sample=stratified, field=true_negative, basis=human_round1) = 17 ## Disagreement differs across the tested participants | Participant | False negative | False positive | Determinate lines | Disagreement | |---|---:|---:|---:|---:| | GPT Image 2 medium | 15 | 0 | 71 | 21.13% | | GPT Image 2.5 low | 10 | 0 | 71 | 14.08% | | Nano Banana 2 Fast | 16 | 3 | 63 | 30.16% | These rates are conditional on the reviewed, readable-or-absent lines; they are not global correction factors. The raw automatic market-wide scores must not be multiplied by them. Check: review_count(participant=gpt2_medium, sample=stratified, field=false_negative, basis=human_round1) = 15 Check: review_count(participant=gpt25_low, sample=stratified, field=false_negative, basis=human_round1) = 10 Check: review_count(participant=nano_banana_2_fast, sample=stratified, field=false_negative, basis=human_round1) = 16 Check: review_count(participant=nano_banana_2_fast, sample=stratified, field=false_positive, basis=human_round1) = 3 Check: review_count(participant=gpt2_medium, sample=stratified, field=false_positive, basis=human_round1) = 0 Check: review_count(participant=gpt25_low, sample=stratified, field=false_positive, basis=human_round1) = 0 Check: review_count(participant=gpt2_medium, sample=stratified, field=determinate, basis=human_round1) = 71 Check: review_count(participant=gpt25_low, sample=stratified, field=determinate, basis=human_round1) = 71 Check: review_count(participant=nano_banana_2_fast, sample=stratified, field=determinate, basis=human_round1) = 63 Check: review_pct(participant=gpt2_medium, sample=stratified, field=disagreements, basis=human_round1) = 21.13 Check: review_pct(participant=gpt25_low, sample=stratified, field=disagreements, basis=human_round1) = 14.08 Check: review_pct(participant=nano_banana_2_fast, sample=stratified, field=disagreements, basis=human_round1) = 30.16 Context: 2 = a model version name Context: 2.5 = a model version name ## Mechanisms and unresolved positive disagreements The source readings show short-word omissions and substitutions such as `K-1` read as `K-l`, `production line №1` read with `Ng1`, `1000-я` read with Cyrillic O letters, and fragments interleaved between neighboring columns. The line audit preserves each original scorer candidate and reconstructed block beside the owner's verdict; no human judgment has been rewritten to fit the automatic result. All 3 positive disagreements belong to Nano Banana: D18 / 2016, I10 / 2013 and I10 / 2021. The owner reported a readable deviation without a full transcription on these lines. Their individual mechanisms remain unresolved. The scorer folds letter case, homoglyphs and punctuation, so the visual and normalized automatic criteria need not be identical. Check: review_count(participant=nano_banana_2_fast, sample=stratified, field=false_positive, basis=human_round1) = 3 Context: 2016 = the year identifying a reviewed reference line Context: 2013 = the year identifying a reviewed reference line Context: 2021 = the year identifying a reviewed reference line Context: 1 = a literal product-code suffix and production-line identifier in the reference Context: 1000 = the literal unit count in a reference string, not an observed rate ## Diagnostic images, reported separately The 9 failure-selected diagnostic images contain 91 expected lines: 75 were judged unreadable and 16 absent. None was accepted as exact. This is evidence about the specifically flagged outputs, not a random sample of recognizer performance. Do not pool it with the stratified error estimate. Check: review_count(participant=all, sample=flagged, field=images, basis=human_round1) = 9 Check: review_count(participant=all, sample=flagged, field=lines, basis=human_round1) = 91 Check: review_count(participant=all, sample=flagged, field=unreadable, basis=human_round1) = 75 Check: review_count(participant=all, sample=flagged, field=missing, basis=human_round1) = 16 ## Consequences for the article The old 2.9% scattered-block indicator is not recognition accuracy. The original 61-versus-56 automatic comparison does not establish a text quality advantage for the newer low tier. Both reviewed GPT tiers received 71 accepted lines out of 71; Nano Banana received 43 out of 71. The headline claiming a twofold market-wide gap is withdrawn. The human comparison is limited to this reviewed sweep and these participants. Tariff savings are independent of OCR and remain historical price-list arithmetic. Product migration still needs repeated or production validation. The remaining repeated-generation character scores are automatic and have not been manually calibrated by this review. Check: scattered_pct(scope=round1) = 2.9 Check: review_count(participant=gpt25_low, sample=stratified, field=automatic_exact, basis=human_round1) = 61 Check: review_count(participant=gpt2_medium, sample=stratified, field=automatic_exact, basis=human_round1) = 56 Check: review_count(participant=gpt25_low, sample=stratified, field=exact, basis=human_round1) = 71 Check: review_count(participant=gpt2_medium, sample=stratified, field=exact, basis=human_round1) = 71 Check: review_count(participant=nano_banana_2_fast, sample=stratified, field=exact, basis=human_round1) = 43 Context: 2 = the withdrawn twofold headline, not a measured ratio ## Limits and reproduction The reviewer saw the reference and reviewed enlarged crops, with original images available; the judgments were not blind. Crop localization used the original OCR geometry, which can bias which region is inspected. The original images were checked byte-for-byte before review and their hashes are preserved in the export. This calibrates body-text judgments only. Full manual transcriptions of all deviations are unavailable, so no recognizer-only character error rate is claimed. Later sweeps and other participants were not human-reviewed. No statistical confidence or general correction factor is inferred. Recompute the calibration and its published numbers from the repository: ```sh python3 -m study.calibrate --verify-assignment --write study/results/human-calibration python3 -m study.verify_claims python3 -m pytest tests/study ``` The original automatic `metrics.csv` and `runs.csv` remain unchanged.