测试文生图中文字语义是否泄露到非文本区域,发现高精度不等于语义可控。
T2LSC-Bench: Benchmarking Localized Semantic Control in Text-to-Image Generation

- 设计分块测试集,分离语义关系与场景开放度,精准评估文字控制
- 实测语义泄露率最高达18.2%,远超文字识别准确率下降幅度
- 提出抗泄露提示法,降低泄露率同时保持生成质量,适合界面设计场景
近期文生图模型虽能准确生成指定文字,但在产品标签、标识和界面设计等场景中,目标文字需在指定区域内呈现,且不能改变主体身份或周围语义。本文将此类问题称为目标文字相关语义泄露,即目标语义通过非文本视觉内容表达。现有视觉-文本基准主要评估可读性、拼写和布局,未涵盖此类泄露。我们提出T2LSC-Bench,一个受控诊断基准,包含50个种子主体和每模型1,200个提示,共覆盖六模型7,160张图像。其因子化设计涵盖语义关系、场景开放度、提示模式与语言。采用双分支协议,结合OCR-VLM文本验证与结构化VLM语义判断,测量文本锚定准确率(TAA)、语义主体保留率(SSP)、语义泄露率(SLR)与条件语义泄露率(cSLR)。在压力测试下,SLR从1.2%升至18.1%,cSLR从1.3%升至18.2%,而TAA仅从91.4%降至90.9%。抗泄露提示使SLR由16.6%降至8.4%,不损害生成精度。420张图像的人工验证显示自动标注与人工判定高度一致。结果表明,文字生成准确并不意味着语义局部可控。
原文摘要 · Abstract (English)
Recent text-to-image models have become increasingly capable of rendering explicit text, but reliable localized text control requires more than generating the correct string. In applications such as product labeling, signage, and interface design, target text should be rendered within a designated text-bearing region without altering the predefined subject identity or surrounding scene semantics. We refer to violations of this requirement as target-text-associated semantic leakage, in which target-text semantics are expressed through non-textual visual content beyond the designated anchor. Existing visual-text benchmarks primarily evaluate readability, spelling accuracy, and layout, leaving this form of semantic leakage largely unexamined. We introduce T2LSC-Bench, a controlled diagnostic benchmark comprising 50 seed subjects and 1,200 prompt cases per model, yielding 7,160 evaluated images across six models. Its factorized design varies semantic relation, scene openness, prompt mode, and language. A dual-branch protocol combines OCR-VLM text verification with structured VLM semantic judgments to measure Text-at-Anchor Accuracy (TAA), Semantic Subject Preservation (SSP), Semantic Leakage Rate (SLR), and Conditional Semantic Leakage Rate (cSLR). Under stress-test conditions, SLR increases from 1.2% to 18.1% and cSLR from 1.3% to 18.2%, whereas TAA decreases only from 91.4% to 90.9%. Anti-leakage prompting reduces SLR from 16.6% to 8.4% without degrading rendering accuracy. Human validation on 420 images shows strong agreement between automatic and adjudicated annotations. These results show that accurate text rendering does not guarantee local containment of target-text semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。