研究盲人与视力正常科学家如何用AI问答查阅科学论文中的图表。
Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists

- 通过访谈10位科学家,比较他们使用ChatGPT和Gemini查询图文内容的方式。
- 发现模糊或错误的图像描述会导致盲人和视力正常者都放弃AI辅助查询。
- 贡献了115条真实交互数据,支持未来无障碍科学问答系统研发。
视觉图表、图形和表格在科学论文中至关重要,传递文字无法涵盖的信息。传统上,盲人或低视力(BLV)科学家依赖静态替代文本访问图表,但人工智能(AI)的发展使交互式问答(QA)成为探索视觉内容的新范式;然而,人们对科学家如何实际使用视觉问答,以及如何提升其可访问性仍知之甚少。本研究采访了来自不同STEM领域的五位BLV科学家和五位视力正常科学家,探讨他们如何使用ChatGPT和Gemini两款AI工具查询多模态科学文档。研究结果揭示了科学家在阅读多模态内容时的现有实践(包括可访问性绕行方案),以及对AI生成回答适用性的反馈。进一步发现,图像描述模糊或不完整,以及更广泛的AI输出错误,都会导致BLV和视力正常科学家放弃使用AI工作流。为支持未来研究,我们还公开了115条参与者与AI工具交互的查询与回应数据集,涵盖其所在领域的论文。最后,文章讨论了对基于AI的科学问答系统的启示,强调跨能力与跨领域可访问性的重要性。
原文摘要 · Abstract (English)
Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text. While blind or low-vision (BLV) scientists have traditionally relied on static alternative text to access figures in papers, the rise of artificial intelligence (AI) has made interactive question-answering (QA) a feasible paradigm for visual exploration; yet little is known about how scientists use visual QA in practice or how to improve its accessibility. In this work, we interview five BLV and five sighted scientists across different STEM fields to understand how they use two AI tools, ChatGPT and Gemini, to query multimodal scientific documents. Our findings characterize how scientists review multimodal content, including existing practices (along with accessibility workarounds) for engaging with visuals, and feedback on the suitability of AI-generated responses to multimodal queries. We further find that vague or incomplete image descriptions, as well as incorrect AI outputs more broadly, can cause both BLV and sighted scientists to abandon AI workflows. To support future research, we additionally contribute a dataset of 115 queries and responses from our participants' interactions with the AI tools for papers in their field. We close by discussing implications for AI-powered scientific QA systems, emphasizing considerations for access across abilities and domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。