arXiv:2607.21611cs.HCcs.AI2026-07中稿 · publication in Int…综述

系统梳理了提升科学图像可访问性的AI描述技术,助力视障者平等获取科技内容。

A Systematic Survey on Image Description Techniques for STEM Domains

论文配图:A Systematic Survey on Image Description Techniques for STEM Domains
图 1 · 摘自论文原文
  • 分析20篇论文,聚焦视障用户需求的智能图像描述方法
  • 发现现有技术仍存在事实错误和评估指标失真问题
  • 适合关注无障碍设计、人机交互与可解释AI的研究者

STEM领域视觉数据激增,但视障人士面临严重信息障碍。尽管人工智能在生成科技图像文字描述方面取得进展,研究分散且对真实用户的实际影响有限。本系统综述分析了20篇同行评审论文,基于PRISMA方法和ROBIS偏倚评估,考察了所针对的STEM图像类型、采用的AI与机器学习架构、使用的数据集与评价指标,以及描述信息的交互传递方式。结果表明,技术正从静态单次替代文本向融合对话接口、键盘导航及音频/触觉反馈的交互式多模态系统演进。然而,关键挑战依然存在:事实性错误与幻觉频发,缺乏以视障用户为中心设计的可用数据集,且过度依赖文本重叠度等自动指标,无法准确反映描述的实际有用性和可信度。文章最后提出未来研究方向,强调用户可控的描述冗余度、可解释可验证的AI流程,并推动无障碍描述工具融入主流科研与教学环境。

原文摘要 · Abstract (English)

The proliferation of visual data in Science, Technology, Engineering, and Mathematics (STEM) fields presents accessibility barrier for individuals with blindness or visual impairments. While recent advances in Artificial Intelligence (AI) offer new opportunities to generate textual descriptions of STEM images, the research landscape is fragmented and its impact on real users remains limited. This systematic survey examines 20 peer-reviewed studies on AI-based techniques for describing STEM visuals, with a specific focus on accessibility and human-computer interaction. Following the PRISMA methodology and a ROBIS-based risk-of-bias assessment, the review analyzes (i) the types of STEM visuals targeted, (ii) the AI and machine learning architectures employed, (iii) the datasets and evaluation metrics adopted, and (iv) the interaction modalities through which descriptions are delivered. The analysis reveals a shift from static, one-shot alt text toward interactive and multimodal systems that integrate conversational interfaces, keyboard navigation, and audio or haptic feedback. However, critical challenges persist, including factual inaccuracies and hallucinations, the scarcity of accessibility-first datasets co-designed with blind and low-vision users, and a heavy reliance on automatic text-overlap metrics that poorly capture perceived usefulness and trust. The survey concludes by outlining key research directions for HCI, emphasizing user-controlled verbosity, explainable and verifiable AI pipelines, and the integration of accessible description tools into mainstream STEM authoring and learning environments.

图像描述无障碍AI评估人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。