减少视觉语言模型输出冗余,提升盲人助行系统实用性
Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants
- 用人类偏好奖励优化输出,提升简洁性与信息密度
- 引入环境感知判别器,降低无效提醒频率
- 适合关注无障碍辅助与高效交互的开发者
全球约2.83亿人患有视觉障碍,推动了利用视觉语言模型(VLMs)开发盲人助行系统的研究。然而现有模型输出常含大量冗余和无关细节,影响用户对环境的准确判断;且缺乏主动风险评估能力,导致提醒过于频繁,造成时间冗余。为此,我们提出WalkVLM-LR,通过在GRPO推理框架中引入四种基于人类偏好的定制奖励函数(简洁性、流畅性、关键词密度、准确性),优化输出以减少内容冗余。同时,设计共享视觉编码器的环境感知判别器,实现场景风险等级评估,动态触发提醒,降低计算冗余。实验表明,该方法在各项指标上均达最优,尤其在输出简洁性和时间冗余控制方面显著优于其他模型。
原文摘要 · Abstract (English)
Approximately 283 million people worldwide live with visual impairments, motivating increasing research into leveraging Visual Language Models (VLMs) to develop effective walking assistance systems for blind and low vision individuals. However, existing VLMs in walking assistant task often have outputs that contain considerable redundancy and extraneous details, adversely affecting users' ability to accurately assess their surroundings. Moreover, these models typically lack the capability to proactively assess environmental risks and adaptively trigger reminders based on the appropriate scene, leading to excessive temporal redundancy. To mitigate output and temporal redundancy, we propose WalkVLM-LR, a walking assistance model with less redundancy. To reduce output redundancy, we introduce four human-preference-based custom reward functions within the GRPO-based reasoning framework to optimize the output in terms of conciseness, fluency, keyword density, and accuracy, thereby producing more informative and streamlined outputs. To minimize temporal redundancy, we incorporate an environment awareness discriminator, which shares the visual encoder with the VLMs to reduce redundant computations and enhance discriminative efficiency, to make WalkVLM-LR assess scene risk levels and minimize unnecessary reminders. Experimental results demonstrate that our method achieves state-of-the-art performance across all evaluation metrics compared with other models, particularly in output conciseness and less temporal redundancy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。