评测大模型对盲文的理解与表达能力,发现其远不如对英文的处理。
BrailleBench: Investigating Multi-Criteria Braille Comprehension in Large Language Models
- 构建盲文多标准评测基准,涵盖数学、常识等5570个题目
- 模型在盲文输入理解上表现差,二级盲文尤其脆弱
- 适合关注无障碍AI与残障群体技术公平的研究者
尽管大语言模型(LLMs)为知识获取和计算辅助提供支持,但其能力应惠及所有人群,包括视障与全盲用户。然而,现有AI系统是否能通过盲文实现同等功能尚不明确,因盲文的符号、缩写及数字表示带来独特理解挑战。为此,我们提出BrailleBench,一个用于评估LLMs在盲文理解上的多标准基准。该基准整合了来自五个数据集的5,570个实例,覆盖数学、常识与多跳问答任务,涉及英语及一级、二级盲文。设计不同配置以测试模型是否能理解盲文内容、用盲文作答,以及完成端到端盲文交互。为保证质量并避免评估偏差,基准通过专家审核的确定性流程构建,使用自研盲文工具包,未采用任何由大模型生成的数据。我们评估了六种代表性大模型,结果揭示印刷英文能力与盲文可访问性之间存在持续差距。盲文理解和表达呈非对称性,二级盲文在输入端尤为脆弱,全盲文请求进一步降低性能。实验观察为未来盲文AI系统开发提供了重要指导。所有相关资源均公开可用。
原文摘要 · Abstract (English)
Although Large language models (LLMs) mediate access to knowledge and computational assistance, their capabilities should benefit vulnerable groups in the same way. However, it is unclear whether existing AI systems are inclusive enough for blind and deafblind users to access the same functionality through Braille, whose indicators, contractions, and digital representations introduce distinct requirements for model comprehension. To this end, we introduce BrailleBench, a benchmark for evaluating LLMs in Braille comprehension from different Criteria. BrailleBench aligns 5,570 instances from five datasets, including mathematics, commonsense, and multi-hop question answering across English and Braille Grades 1 and 2. Different configurations are designed to understand whether the systems can comprehend Braille-authored content, express answers in Braille, and complete end-to-end Braille interaction. To ensure the quality and prevent evaluation bias, the benchmark is built through a deterministic, expert-reviewed pipeline via a self-created Braille Toolkit without using any data instances generated by LLMs. We evaluate six representative LLMs from various aspects. The results reveal a persistent gap between print-English capability and Braille accessibility. Braille understanding and expression are asymmetric, where Grade 2 is especially fragile on the input side compared to Grade 1, and fully Braille requests further reduce performance. The experimental observations provide valuable guidance for the development of future Braille AI systems. All related resources in BrailleBench are publicly available for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。