专门评估生成图中文本视觉质量,让模型更懂人类对文字的挑剔。
TIQA: Human-Aligned Perceptual Text Quality Assessment in Generated Images
- 提出新任务TIQA,专注评估图像中文字区域的感知质量。
- 在120k文本裁片上训练,模型与人评分数相关性达0.942(PLCC)。
- 适合用于筛选高质量生成图,尤其关注文字清晰度的场景。
当前文本生成图像模型虽整体逼真度提升,但文字渲染仍是主要缺陷:图像整体可信,局部文字却常出现错字形、笔画断裂、间距不均等瑕疵,人类对此高度敏感。本文提出无参考式文本图像质量评估(TIQA),针对检测到的文字区域估算与人类感知一致的质量分数,并分离视觉质量与语义正确性。为此构建两个数据集:TIQA-Crops包含12万张来自3.6万幅AI生成图像的文本裁片(12个生成器),提供10,000条平均意见评分(MOS)及11万条代理标签用于预训练;TIQA-Images包含1,500张文字密集图像(10个近期生成器,含专有系统),附带整体质量与文字质量的主观评分。同时提出ANTIQA模型,具文本特定归纳偏置的轻量级预测器。在裁片与图像级评估中,ANTIQA均达到最佳人评对齐效果:在TIQA-Crops上PLCC/SROCC达0.942/0.935,在未见生成器的TIQA-Images文本质量MOS上为0.842/0.837。在五选一排名任务中,使用ANTIQA可使所选图像文字质量提升0.36 MOS(14%),证明其在基准测试、过滤与生成时选择中的实用价值。结果确立了感知文字质量作为现代文本生成图像的独立评估目标。
原文摘要 · Abstract (English)
Recent text-to-image models have improved global realism, but text rendering remains a persistent failure mode: images may look convincing overall, yet local typography often contains malformed glyphs, broken strokes, irregular spacing, and other artifacts that humans heavily penalize. We formulate Text-in-Image Quality Assessment (TIQA), a no-reference task that estimates a human-aligned perceptual quality score for detected text regions while disentangling visual text quality from semantic correctness. To support this setting, we introduce two datasets. TIQA-Crops contains 120k text crops from 36k AI-generated images produced by 12 generators, with 10k mean-opinion-score (MOS) labels and 110k proxy labels for pretraining. TIQA-Images contains 1,500 text-heavy images from 10 recent generators, including proprietary systems, with paired overall-quality and text-quality subjective scores. We also propose ANTIQA, a lightweight predictor with text-specific inductive biases. Across crop-level and image-level evaluations, ANTIQA achieves the best alignment with human judgments, reaching PLCC/SROCC of 0.942/0.935 on TIQA-Crops and 0.842/0.837 for text-quality MOS on unseen generators in TIQA-Images. In best-of-5 AI-generated image ranking, ANTIQA improves the text quality of the selected image by 0.36 MOS (14%), demonstrating utility for benchmarking, filtering, and generation-time selection. Together, these findings establish perceptual text quality as a distinct evaluation target for modern text-to-image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。