arXiv:2503.08221cs.CVcs.AI2025-03NeurIPS被引 24

首个面向视障者的第一人称视觉辅助数据集,用于评估大模型实际助盲能力。

EgoBlind: Towards Egocentric Visual Assistance for the Blind

  • 收集1392段视障者日常第一视角视频,5311个真实需求问题
  • 16个主流多模态大模型最高准确率仅60%,远低于人类87.4%
  • 揭示模型在场景理解与语义推理上的关键缺陷,适合助盲系统研究者

我们提出EgoBlind,首个从视障人士采集的第一人称视觉问答数据集,用于评估当代多模态大语言模型(MLLMs)的辅助能力。EgoBlind包含1,392段盲人及视力障碍者日常生活中的第一视角视频,以及5,311个由视障者直接提出或验证的问题,反映其实际视觉辅助需求。每个问题配有平均3个人工标注的参考答案以减少主观性。利用EgoBlind,我们全面评估了16个先进MLLMs,发现所有模型表现不佳,最佳模型准确率仅约60%,远低于人类87.4%的表现。为推动未来进步,我们总结现有MLLM在视障者第一人称视觉辅助中的主要局限,并探索改进的启发式方案。我们希望EgoBlind能成为提升视障人群独立性的有效AI助手研发基础。数据与代码已公开于https://github.com/doc-doc/EgoBlind。

原文摘要 · Abstract (English)

We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from the daily lives of blind and visually impaired individuals. It also features 5,311 questions directly posed or verified by the blind to reflect their in-situation needs for visual assistance. Each question has an average of 3 manually annotated reference answers to reduce subjectiveness. Using EgoBlind, we comprehensively evaluate 16 advanced MLLMs and find that all models struggle. The best performers achieve an accuracy near 60\%, which is far behind human performance of 87.4\%. To guide future advancements, we identify and summarize major limitations of existing MLLMs in egocentric visual assistance for the blind and explore heuristic solutions for improvement. With these efforts, we hope that EgoBlind will serve as a foundation for developing effective AI assistants to enhance the independence of the blind and visually impaired. Data and code are available at https://github.com/doc-doc/EgoBlind.

视觉辅助多模态模型视障科技数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。