构建2020张数学图像描述数据集,助力无障碍教育
MIDAL: A Dataset of Math Image Descriptions for Accessible Learning

- 收集2020张跨教育阶段的数学图像及对应可访问描述
- 支持视觉语言模型训练,生成符合无障碍标准的图像描述
- 适用于提升数学推理与问答能力的语言模型微调
大量开放教育资源缺乏可访问性,尤其在深度图像描述方面。科学与数学等学科因内容复杂、术语多样,撰写图像描述尤为困难。为缓解这一问题,我们提出数学图像描述数据集MIDAL,包含2,020张覆盖多个教育阶段的数学图像及其可访问性最佳实践指导下的描述,旨在支持视觉语言模型训练,生成更准确的图像描述。该数据集不仅服务于数学图像描述生成,还可用于微调语言模型,提升其数学推理与答案生成能力,推动高等教育中STEM内容的可访问性发展。
原文摘要 · Abstract (English)
Many open educational resources are lacking in accessibility, especially in-depth image descriptions. In subjects like Science and Mathematics, however, it can be particularly difficult to write image descriptions since there can be many complicated expressions and names depending upon the course level. To help fill that gap in a small way, we introduce Math Image Descriptions for Accessible Learning (MIDAL), a math image-description dataset of 2,020 mathematical images spanning multiple educational levels, to aid in training vision language models to create image descriptions following accessibility best practices. We hope MIDAL is a valuable resource in enhancing the conversation and innovation regarding accessibility of STEM content in higher education. This dataset is however not just limited in math description generation but can also be used to fine-tune language models that can have improved mathematical reasoning and answers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。