提升视觉语言长序列任务规划能力,通过结构化偏好优化增强推理与决策。
Structured Preference Optimization for Vision-Language Long-Horizon Task Planning
- 基于任务相关性、视觉对齐和历史一致性评分,系统评估推理链质量。
- 在VirtualHome和Habitat 2.0上,推理准确率提升5.98%至3.30%,决策成功率提升4.68%至2.11%。
- 适合研究复杂环境长期任务规划的开发者与研究员参考。
现有视觉语言任务规划方法在短序列任务中表现优异,但在动态环境中的复杂长序列规划中仍存在不足。主要挑战在于难以有效训练模型生成高质量的长序列推理过程。为此,我们提出结构化偏好优化(SPO),通过结构化偏好评估与优化训练策略,提升长序列任务规划中的推理与动作选择能力。SPO包含:1)基于偏好的评分与优化,系统评估推理链的任务相关性、视觉对齐性和历史一致性;2)课程引导训练,模型从简单到复杂逐步适应,提升泛化能力与推理鲁棒性。为推动该领域研究,我们构建了ExtendaBench,一个涵盖1,509个任务的综合性基准,覆盖VirtualHome与Habitat 2.0,分为超短、短、中、长四类任务。实验表明,SPO显著提升推理质量与最终决策准确性,在长序列任务中优于已有方法。具体地,在VirtualHome上实现GCR+5.98%、SR+4.68%,在Habitat上实现GCR+3.30%、SR+2.11%。
原文摘要 · Abstract (English)
Existing methods for vision-language task planning excel in short-horizon tasks but often fall short in complex, long-horizon planning within dynamic environments. These challenges primarily arise from the difficulty of effectively training models to produce high-quality reasoning processes for long-horizon tasks. To address this, we propose Structured Preference Optimization (SPO), which aims to enhance reasoning and action selection in long-horizon task planning through structured preference evaluation and optimized training strategies. Specifically, SPO introduces: 1) Preference-Based Scoring and Optimization, which systematically evaluates reasoning chains based on task relevance, visual grounding, and historical consistency; and 2) Curriculum-Guided Training, where the model progressively adapts from simple to complex tasks, improving its generalization ability in long-horizon scenarios and enhancing reasoning robustness. To advance research in vision-language long-horizon task planning, we introduce ExtendaBench, a comprehensive benchmark covering 1,509 tasks across VirtualHome and Habitat 2.0, categorized into ultra-short, short, medium, and long tasks. Experimental results demonstrate that SPO significantly improves reasoning quality and final decision accuracy, outperforming prior methods on long-horizon tasks and underscoring the effectiveness of preference-driven optimization in vision-language task planning. Specifically, SPO achieves a +5.98% GCR and +4.68% SR improvement in VirtualHome and a +3.30% GCR and +2.11% SR improvement in Habitat over the best-performing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。