用视障者历史提问引导大模型生成更精准的图像描述
Guiding Multimodal Large Language Models with Blind and Low Vision People Visual Questions for Proactive Visual Interpretations
- 基于视障用户历史问题,匹配相似图像上下文
- 上下文感知描述准确回应用户问题达76.1%
- 适合无障碍视觉辅助系统开发者参考
多模态大语言模型(MLLM)因高精度和类人化描述能力,被用于支持视障和低视力(BLV)用户的视觉解释。然而现有应用常提供冗长、泛化的描述,与实际需求脱节。为此,我们提出一种系统:从VizWiz-LF数据集中识别与当前图像相似的历史视觉场景,并利用相关用户提问引导MLLM生成更贴合BLV用户需求的描述。在3名标注员对92组上下文感知与非上下文描述的评估中,上下文感知描述在76.1%(70/92)的情况下预判并回答了用户问题,且在54.4%(50/92)的对比中更受青睐。论文及数据已开源。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have been integrated into visual interpretation applications to support Blind and Low Vision (BLV) users because of their accuracy and ability to provide rich, human-like interpretations. However, these applications often default to comprehensive, lengthy descriptions regardless of context. This leads to inefficient exchanges, as users must go through irrelevant details rather than receiving the specific information they are likely to seek. To deliver more contextually-relevant information, we developed a system that draws on historical BLV users questions. When given an image, our system identifies similar past visual contexts from the VizWiz-LF dataset and uses the associated questions to guide the MLLM generate descriptions more relevant to BLV users. An evaluation with three human labelers who revised 92 context-aware and context-free descriptions showed that context-aware descriptions anticipated and answered users' questions in 76.1% of cases (70 out of 92) and were preferred in 54.4% of comparisons (50 out of 92). Our paper reviews, and data analysis are publicly available in a Github repository at https://github.com/rgonzalezp/guiding-multimodal-large-language-models-with-blind-and-low-vision-people-visual-questions .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。