arXiv:2505.04488cs.CVcs.AI2025-05被引 5

首个评估视频大模型助盲效果的研究,发现GPT-4o表现最佳

"I Can See Forever!": Evaluating Real-time VideoLLMs for Assisting Individuals with Visual Impairments

  • 构建日常任务基准集VisAssistDaily,评估视频大模型助盲能力
  • GPT-4o在任务成功率上最高,但对危险识别仍有不足
  • 提出SafeVid数据集并微调VITA-1.5,风险识别准确率提升至76%

视障人群在日常活动中面临重大挑战。尽管已有研究利用视觉语言模型提供帮助,但多数聚焦静态内容,难以应对复杂环境中的实时感知需求。近期的VideoLLMs实现了视觉与语音的实时交互,为辅助任务带来新可能。本文首次系统评估其在支持视障人士日常生活方面的有效性。我们通过面向视障用户的调研设计了基准测试集VisAssistDaily。基于该数据集,评估多个主流VideoLLMs,发现GPT-4o任务成功率最高。进一步用户研究揭示了对危险感知的担忧。为此,我们构建了环境感知数据集SafeVid,对VITA-1.5进行微调,将风险识别准确率从25.00%提升至76.00%。本工作为该领域未来研究提供了重要参考。

原文摘要 · Abstract (English)

The visually impaired population faces significant challenges in daily activities. While prior works employ vision language models for assistance, most focus on static content and cannot address real-time perception needs in complex environments. Recent VideoLLMs enable real-time vision and speech interaction, offering promising potential for assistive tasks. In this work, we conduct the first study evaluating their effectiveness in supporting daily life for visually impaired individuals. We first conducted a user survey with visually impaired participants to design the benchmark VisAssistDaily for daily life evaluation. Using VisAssistDaily, we evaluate popular VideoLLMs and find GPT-4o achieves the highest task success rate. We further conduct a user study to reveal concerns about hazard perception. To address this, we propose SafeVid, an environment-awareness dataset, and fine-tune VITA-1.5, improving risk recognition accuracy from 25.00% to 76.00%.We hope this work provides valuable insights and inspiration for future research in this field.

视频大模型助盲风险识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。