用视觉语言模型生成室内导航指令,帮视障者精准到达目标位置。
Fine-Tuning Vision-Language Models for Visual Navigation Assistance
- 用LoRA微调BLIP-2模型,结合图像与自然语言生成导航指令。
- 改进的BERT F1指标更关注方向和顺序,评估效果提升显著。
- 适合辅助视障人群独立出行,尤其在无精确定位的室内场景。
我们研究基于视觉与语言的室内导航,帮助视障人士通过图像和自然语言指引抵达目标位置。传统导航系统因缺乏精确位置信息,在室内环境表现不佳。本文方法融合视觉与语言模型,生成分步导航指令,提升可访问性与独立性。我们在人工标注的室内导航数据集上,采用低秩适配(LoRA)对BLIP-2模型进行微调。提出一种改进的评估指标,通过强化方向性和序列性变量,更全面衡量导航性能。微调后模型在生成方向性指令方面显著优于原始BLIP-2,克服了其原有局限。
原文摘要 · Abstract (English)
We address vision-language-driven indoor navigation to assist visually impaired individuals in reaching a target location using images and natural language guidance. Traditional navigation systems are ineffective indoors due to the lack of precise location data. Our approach integrates vision and language models to generate step-by-step navigational instructions, enhancing accessibility and independence. We fine-tune the BLIP-2 model with Low Rank Adaptation (LoRA) on a manually annotated indoor navigation dataset. We propose an evaluation metric that refines the BERT F1 score by emphasizing directional and sequential variables, providing a more comprehensive measure of navigational performance. After applying LoRA, the model significantly improved in generating directional instructions, overcoming limitations in the original BLIP-2 model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。