arXiv:2608.28218cs.CV2026-08

让视觉语言模型按重要性排序描述内容,帮视障者更快抓住关键信息。

Focus Where It Counts: A Salience-Driven Vision-Language Model for Low Vision Assistance

论文配图:Focus Where It Counts: A Salience-Driven Vision-Language Model for Low Vision Assistance
图 1 · 摘自论文原文
  • 基于人类感知优先级设计新框架,按重要性顺序生成描述
  • 构建三个经视障用户验证的标注数据集,包含物体重要性标签
  • 开发可部署于助盲眼镜的系统,真实场景中验证实用性

视觉语言模型(VLMs)在辅助视障人士方面展现出巨大潜力,但现有模型主要面向通用图像描述,未显式建模人类感知优先级,导致难以突出场景中的关键信息。为此,本文提出一种基于显著性的图像描述框架,依据对视障用户的重要性排序场景元素。我们构建了三个显著性感知数据集:Salience COCO、Salience Flickr 与 Salience VizWiz,其中包含物体级显著性标注,反映不同环境下对低视力用户最相关的视觉信息。基于这些数据集,我们提出 Salience-LLaVA——一种融入显著性提示的视觉语言模型,能生成按重要性顺序排列的描述。本文的主要贡献包括:构建经低视力用户验证的显著性数据集;提出 Salience-LLaVA 模型;引入 SCMI 评估描述顺序准确性;并在助盲眼镜上部署系统,证明其实际可用性。代码与数据集已开源。

原文摘要 · Abstract (English)

Vision-language models (VLMs) are rapidly progressing and offer promising capabilities for assistive technologies supporting persons with blindness or low vision. However, existing VLMs are primarily designed for general-purpose captioning and do not explicitly model human perceptual priorities, thereby limiting their ability to emphasize the most relevant information in a scene. To address this gap, we propose a salience-driven captioning framework that prioritizes scene elements according to their importance for human-centered assistance. We curate three salience-aware datasets, namely, Salience COCO, Salience Flickr, and Salience VizWiz, with object-level salience annotations designed to reflect the visual information most relevant to low vision users across different environments. Building on these datasets, we introduce Salience-LLaVA, a salience-aware VLM that incorporates salience cues to generate captions in which important elements are mentioned in the order of importance. Our work makes four main contributions. We build salience-aware datasets verified by low vision participants, propose Salience-LLaVA to describe objects in the order of importance, introduce SCMI to evaluate ordering accuracy, and deploy the system on assistive glasses to demonstrate real-world practicality. Code and datasets are available at: https://github.com/topo-focus/Topofocus

视觉语言模型视障辅助显著性感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。