构建真实旅行规划基准,测试模型在复杂约束下的多模态决策能力
WorldTravel: A Realistic Multimodal Travel-Planning Benchmark with Tightly Coupled Constraints
- 设计150个跨5城的真实旅行场景,含15个以上强耦合时序逻辑约束
- 多模态环境含2000+渲染网页,模型需从视觉布局中提取参数
- 顶尖模型可行性仅32.67%(文本)和19.33%(多模态),暴露感知与推理断层
真实世界自主规划需协调紧密耦合的约束,单一决策决定后续所有行动的可行性。现有基准多为松耦合约束,可通过局部贪心决策解决,且依赖理想化数据,无法反映从动态网络环境提取参数的复杂性。我们提出 extbf{WorldTravel},包含5个城市共150个真实旅行场景,平均需处理15个以上相互依赖的时间与逻辑约束。为评估模型在真实部署中的表现,我们构建 extbf{WorldTravel-Webscape},一个包含2000多个渲染网页的多模态环境,要求智能体直接从视觉布局中感知约束参数以支持规划。对10个前沿模型的评估显示显著性能下降:即使最先进的GPT-5.2在纯文本设置下可行性也仅32.67%,在多模态环境中进一步降至19.33%。我们识别出关键的感知-动作鸿沟,并发现约10个约束处存在规划视野阈值,模型推理在此处持续失效,表明感知与推理仍是独立瓶颈。这些发现凸显下一代智能体需融合高保真视觉感知与长时程推理,才能应对脆弱的真实世界物流挑战。
原文摘要 · Abstract (English)
Real-world autonomous planning requires coordinating tightly coupled constraints where a single decision dictates the feasibility of all subsequent actions. However, existing benchmarks predominantly feature loosely coupled constraints solvable through local greedy decisions and rely on idealized data, failing to capture the complexity of extracting parameters from dynamic web environments. We introduce \textbf{WorldTravel}, a benchmark comprising 150 real-world travel scenarios across 5 cities that demand navigating an average of 15+ interdependent temporal and logical constraints. To evaluate agents in realistic deployments, we develop \textbf{WorldTravel-Webscape}, a multi-modal environment featuring over 2,000 rendered webpages where agents must perceive constraint parameters directly from visual layouts to inform their planning. Our evaluation of 10 frontier models reveals a significant performance collapse: even the state-of-the-art GPT-5.2 achieves only 32.67\% feasibility in text-only settings, which plummets to 19.33\% in multi-modal environments. We identify a critical Perception-Action Gap and a Planning Horizon threshold at approximately 10 constraints where model reasoning consistently fails, suggesting that perception and reasoning remain independent bottlenecks. These findings underscore the need for next-generation agents that unify high-fidelity visual perception with long-horizon reasoning to handle brittle real-world logistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。