arXiv:2410.03430cs.CVcs.CL2024-10被引 3

测试生成模型能否为无障碍阅读文本快速定制合适图像。

Images Speak Volumes: User-Centric Assessment of Image Generation for Accessible Communication

  • 对比7个开源与闭源图像生成模型,评估其生成效果。
  • 用户研究发现部分模型表现良好,但需人工审核才能大规模使用。
  • 适合无障碍内容创作者和关注可访问性设计的研究者。

解释性图像在易读易懂(E2R)文本中起关键作用。然而,线上数据库中的图像往往不匹配对应文本,而定制图像成本高昂。本研究开展大规模实验,检验文本到图像生成模型是否能快速、低成本地提供可定制图像以弥补这一差距。我们评测了七个模型——四个开源、三个闭源,并对生成结果进行详尽评估。此外,我们针对E2R目标群体开展了用户研究,检验图像是否满足其需求。结果表明,部分模型表现突出,但尚无模型可在无需人工监督的情况下大规模应用。本研究为促进E2R内容创作者高效生成可访问信息,实现图像与目标群体需求精准匹配,迈出重要一步。

原文摘要 · Abstract (English)

Explanatory images play a pivotal role in accessible and easy-to-read (E2R) texts. However, the images available in online databases are not tailored toward the respective texts, and the creation of customized images is expensive. In this large-scale study, we investigated whether text-to-image generation models can close this gap by providing customizable images quickly and easily. We benchmarked seven, four open- and three closed-source, image generation models and provide an extensive evaluation of the resulting images. In addition, we performed a user study with people from the E2R target group to examine whether the images met their requirements. We find that some of the models show remarkable performance, but none of the models are ready to be used at a larger scale without human supervision. Our research is an important step toward facilitating the creation of accessible information for E2R creators and tailoring accessible images to the target group's needs.

图像生成可访问性用户研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。