arXiv:2409.14215cs.CV2024-09中稿 · WACV 2025, project…被引 5

构建首个面向视障人群的多任务视觉语言模型评测基准。

@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology

  • 基于视障用户调研设计五项核心任务评测框架。
  • 统一模型同时完成分割、测距、识别等五类辅助功能。
  • 适合研究无障碍技术与多模态通用模型的开发者。

随着视觉语言模型(VLMs)的发展,面向视障人士(PVIs)的人类中心辅助技术正向通用化演进,具备同时执行多项任务的能力。然而,针对此类技术的评测仍不充分。为此,我们首先构建了一个全新的辅助技术基准(@Bench)。该基准基于对视障用户的前期调研,涵盖五项最关键的任务:全景分割、深度估计、光学字符识别(OCR)、图像描述生成和视觉问答(VQA)。此外,我们提出一种新型辅助模型(@Model),可同步处理所有任务,并可扩展至更多辅助功能。该框架通过融合多模态信息,在各项任务上均表现优异,为视障人士提供更全面的支持。大量实验验证了该框架的有效性与泛化能力。

原文摘要 · Abstract (English)

As Vision-Language Models (VLMs) advance, human-centered Assistive Technologies (ATs) for helping People with Visual Impairments (PVIs) are evolving into generalists, capable of performing multiple tasks simultaneously. However, benchmarking VLMs for ATs remains under-explored. To bridge this gap, we first create a novel AT benchmark (@Bench). Guided by a pre-design user study with PVIs, our benchmark includes the five most crucial vision-language tasks: Panoptic Segmentation, Depth Estimation, Optical Character Recognition (OCR), Image Captioning, and Visual Question Answering (VQA). Besides, we propose a novel AT model (@Model) that addresses all tasks simultaneously and can be expanded to more assistive functions for helping PVIs. Our framework exhibits outstanding performance across tasks by integrating multi-modal information, and it offers PVIs a more comprehensive assistance. Extensive experiments prove the effectiveness and generalizability of our framework.

辅助技术多模态视障支持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。