arXiv:2607.14660cs.CV2026-07

为视障人士设计真实视频评测集,检验多模态大模型辅助能力

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance

论文配图:VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
图 1 · 摘自论文原文
  • 用视障者自拍的第一视角视频构建评测集,贴近真实使用场景
  • 三类任务测试模型预判、问答与情境交互能力,现有模型在预判上表现差
  • 支持实时与离线评估,推动专用于视障辅助的模型研发

视障人士因无法获取视觉信息而面临日常挑战。尽管多模态大语言模型(MLLMs)在通用视觉与语言任务中表现优异,但在真实视障辅助场景中的实际效用仍鲜有研究。为此,我们提出VIABench,一个专为评估MLLM在视障辅助场景中表现而设计的综合性视频基准数据集,数据来自视障者自身录制或分享的第一人称视频。VIABench定义了三大核心任务:主动提醒(评估模型对即将发生的导航关键事件的预判与实时描述能力)、视觉问答(回答用户关于环境或物体的问题)、视觉引导交互(测试模型在情境中的意图交互推理能力)。为确保评估的严谨性,我们设计了支持在线(实时)与离线两种模式的评测流程。实验表明,当前MLLM在主动提醒任务中仍表现不足,尤其在准确预判与实时响应方面存在明显短板。我们希望VIABench能推动面向真实辅助需求的定制化MLLM研发,提升视障人士的出行与交互体验。代码与数据将公开于https://github.com/MCG-NJU/VIABench。

原文摘要 · Abstract (English)

Visually impaired individuals (VIIs) encounter significant daily challenges due to limited access to visual information. Although Multimodal Large Language Models (MLLMs) have achieved impressive results on general vision and language tasks, their practical utility in real-world blind assistance still remains largely underexplored. To fill this gap, we introduce VIABench, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves. VIABench defines three core tasks, each targeting a distinct requirement in visual assistance. Proactive Reminder: Assesses the model's ability to interpret ongoing video content while proactively anticipating and verbally describing upcoming navigation-critical events; Visual Question Answering (VQA): Evaluates the model's capacity to answer user-posed questions about the environment or objects within the video; Vision-Guided Interaction: Tests context-aware reasoning to accomplish intentional interactions between user and environment. To ensure a robust and fair evaluation, we propose a rigorous benchmarking pipeline that supports both online (real-time) and offline settings. Our experiments demonstrate that current MLLMs still struggle to deliver comprehensive support for VIIs, especially in the Proactive Reminder task, which demands accurate anticipation and real-time responsiveness. We hope VIABench will drive future research toward developing customized MLLMs for real-world assistance, ultimately improving navigation and interaction experiences for visually impaired individuals. Code and data will be released at https://github.com/MCG-NJU/VIABench.

视障辅助多模态评测基准第一人称视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。