用视障教师评估视觉描述,构建更适配盲人用户的图示数据集
Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions
- 让有视力者评估VLM生成的图示描述,而非自己撰写
- 数据集含5000张图、13.7万条样本,支持多种任务训练
- 适合做盲人用户适配的视觉描述模型研究
视障用户与标注者在需求和视觉能力上存在差异。为盲人及低视力(BLV)用户生成详尽的图示描述是其中一大挑战。虽然有视力者能轻松描述图像,但现有研究表明,由他们直接生成的内容成本高、易带偏见,且不符合BLV用户标准。本研究让有视力者评估由视觉语言模型(VLM)在多轮推理引导下生成的图示描述,结果表明该评估方式有效且对视障教育工作者极具价值。我们发布了Sightation数据集,涵盖5000张图、13.7万条样本,适用于补全、偏好判断、检索、问答和推理等任务的训练,并验证了其在多种下游任务中的微调潜力。
原文摘要 · Abstract (English)
Often, the needs and visual abilities differ between the annotator group and the end user group. Generating detailed diagram descriptions for blind and low-vision (BLV) users is one such challenging domain. Sighted annotators could describe visuals with ease, but existing studies have shown that direct generations by them are costly, bias-prone, and somewhat lacking by BLV standards. In this study, we ask sighted individuals to assess -- rather than produce -- diagram descriptions generated by vision-language models (VLM) that have been guided with latent supervision via a multi-pass inference. The sighted assessments prove effective and useful to professional educators who are themselves BLV and teach visually impaired learners. We release Sightation, a collection of diagram description datasets spanning 5k diagrams and 137k samples for completion, preference, retrieval, question answering, and reasoning training purposes and demonstrate their fine-tuning potential in various downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。