用大模型模拟视障者反馈,自动优化导航指令生成。
LaF-GRPO: In-Situ Navigation Instruction Generation for the Visually Impaired via GRPO with LLM-as-Follower Reward
- 让大模型扮演视障用户,实时反馈指令优劣,指导视觉语言模型优化。
- 在27,000样本数据集上,指令生成准确率提升14%(BLEU),METEOR达0.542。
- 适合智能助盲系统研发者,无需真人测试即可迭代指令质量。
为视障人士(VI)生成精准、实时、分步的导航指令至关重要但研究较少。本文提出LaF-GRPO(LLM-as-Follower GRPO),利用大语言模型模拟视障用户对导航指令的反应,提供反馈奖励,用于引导视觉-语言模型(VLM)的后训练,从而提升指令准确性与实用性,同时减少昂贵的真实世界数据采集。为解决该领域缺乏专用基准的问题,我们构建了包含27,000个样本的开源数据集NIG4VI,涵盖多样化的导航场景与精确空间坐标,支持详细且开放式的实时指令生成。在NIG4VI上的实验表明,LaF-GRPO效果显著:零样本设置下(Zero-)BLEU提升14%;微调后(SFT+)METEOR达0.542,优于GPT-4o的0.323。定性分析进一步证实,本方法生成的指令更直观、更安全。
原文摘要 · Abstract (English)
Navigation instruction generation for visually impaired (VI) individuals (NIG-VI) is critical yet relatively underexplored. This study focuses on generating precise, in-situ, step-by-step navigation instructions that are practically usable for VI users. Specifically, we propose LaF-GRPO (LLM-as-Follower GRPO), where an LLM simulates VI user responses to navigation instructions, thereby providing feedback rewards to guide the post-training of a Vision-Language Model (VLM). This enhances instruction accuracy and usability while reducing costly real-world data collection needs. To address the scarcity of dedicated benchmarks in this field, we introduce NIG4VI, a 27k-sample open-source dataset to facilitate training and evaluation. It comprises diverse navigation scenarios with accurate spatial coordinates, supporting detailed and open-ended in-situ instruction generation. Experiments on NIG4VI demonstrate the effectiveness of LaF-GRPO through quantitative metrics (e.g., Zero-(LaF-GRPO) boosts BLEU 14\%; SFT+(LaF-GRPO) METEOR 0.542 vs. GPT-4o 0.323), and qualitative analysis further confirms that our method yields more intuitive and safer instructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。