首个军事场景长上下文推理基准,测试大模型跨源空间规划能力
Mil-SCORE: Benchmarking Long-Context Geospatial Reasoning and Planning in Large Language Models
- 构建多源异构军事场景数据集,支持多跳推理
- 现有模型在复杂场景下表现显著不足,平均得分仅32.1%
- 适合研究长上下文推理、空间决策与多模态融合的学者
随着大语言模型应用于更长、更复杂的任务,亟需能够真实反映长上下文需求的评估基准,要求模型具备选择性阅读和整合异构、多模态信息的能力。这一需求在大规模军事规划等地理空间规划问题中尤为突出,需要快速准确地分析地图、作战指令、情报报告等分散数据。为此,我们提出MilSCORE(军事情景上下文推理),据我们所知,这是首个由专家撰写、基于复杂模拟军事场景的场景级问答数据集,用于训练与评估。该基准旨在评估高风险决策与规划能力,检验模型在多源信息间进行战术与空间推理的能力,以及对长时程、地理信息丰富的上下文进行推断的能力。数据集包含七类问题,覆盖事实回忆、约束分析、策略推演与空间分析等。我们提供评估协议,并报告多种主流视觉-语言模型的基线结果。结果显示,当前系统在MilSCORE上仍有巨大提升空间,表明现有模型在真实场景级长上下文规划任务中表现不佳,凸显了该基准的挑战性。
原文摘要 · Abstract (English)
As large language models (LLMs) are applied to increasingly longer and more complex tasks, there is a growing need for realistic long-context benchmarks that require selective reading and integration of heterogeneous, multi-modal information sources. This need is especially acute for geospatial planning problems, such as those found in planning for large-scale military operations, which demand fast and accurate reasoning over maps, orders, intelligence reports, and other distributed data. To address this gap, we present MilSCORE (Military Scenario Contextual Reasoning), to our knowledge the first scenario-level dataset of expert-authored, multi-hop questions grounded in a complex, simulated military planning scenario used for training. MilSCORE is designed to evaluate high-stakes decision-making and planning, probing LLMs' ability to combine tactical and spatial reasoning across multiple sources and to reason over long-horizon, geospatially rich context. The benchmark includes a diverse set of question types across seven categories targeting both factual recall and multi-step reasoning about constraints, strategy, and spatial analysis. We provide an evaluation protocol and report baseline results for a range of contemporary vision-language models. Our findings highlight substantial headroom on MilSCORE, indicating that current systems struggle with realistic, scenario-level long-context planning, and positioning MilSCORE as a challenging testbed for future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。