arXiv:2510.11314cs.CL2025-10

用五种提示模板生成简化文本配图,提升认知障碍者理解力。

Template-Based Text-to-Image Alignment for Language Accessibility: A Study on Visualizing Text Simplifications

  • 设计五种视觉布局提示模板,约束物体数量与空间分布。
  • 基础对象聚焦模板语义对齐最佳,最小化视觉干扰更助理解。
  • 复古风格最易读,维基百科数据源效果最优,适合无障碍内容生成。

智力障碍者常难以理解复杂文本。现有文生图模型多关注美学而忽视可访问性,且未明确图像如何匹配文本简化(TS)。本文提出一种结构化视觉语言模型提示框架,基于400句来自OneStopEnglish、SimPA、Wikipedia和ASSET四个文本简化数据集的句子级简化内容,设计五种提示模板:基础对象聚焦、情境场景、教育版式、多层次细节与网格布局,均符合可访问性约束(如物体数量限制、空间分离、内容限制)。通过两阶段评估:第一阶段使用CLIPScore评估模板效果;第二阶段由四位无障碍专家对十种视觉风格生成的图像进行标注。结果表明,基础对象聚焦模板实现最高语义对齐,证明视觉极简有助于语言可访问性;专家评估确认复古风格最具可访问性,维基百科数据源最有效。不同维度评价一致性差异显著,文本简化度可靠性高,图像质量主观性强。本框架为无障碍内容生成提供实用指导,强调结构化提示在AI可访问性工具中的关键作用。

原文摘要 · Abstract (English)

Individuals with intellectual disabilities often have difficulties in comprehending complex texts. While many text-to-image models prioritize aesthetics over accessibility, it is not clear how visual illustrations relate to text simplifications (TS) generated from them. This paper presents a structured vision-language model (VLM) prompting framework for generating accessible images from simplified texts. We designed five prompt templates, i.e., Basic Object Focus, Contextual Scene, Educational Layout, Multi-Level Detail, and Grid Layout, each following distinct spatial arrangements while adhering to accessibility constraints such as object count limits, spatial separation, and content restrictions. Using 400 sentence-level simplifications from four established TS datasets (OneStopEnglish, SimPA, Wikipedia, and ASSET), we conducted a two-phase evaluation: Phase 1 assessed prompt template effectiveness with CLIPScores, and Phase 2 involved human annotation of generated images across ten visual styles by four accessibility experts. Results show that the Basic Object Focus prompt template achieved the highest semantic alignment, indicating that visual minimalism enhances language accessibility. Expert evaluation further identified Retro style as the most accessible and Wikipedia as the most effective data source. Inter-annotator agreement varied across dimensions, with Text Simplicity showing strong reliability and Image Quality proving more subjective. Overall, our framework offers practical guidelines for accessible content generation and underscores the importance of structured prompting in AI-generated visual accessibility tools.

可访问性文生图提示工程认知辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。