arXiv:2511.10651cs.CLcs.AI2025-11被引 1

用大模型分步推理生成军事推演高质量分析报告

Data Analysis and Performance Evaluation of Simulation Deduction Based on LLMs

  • 将复杂推演分析拆解为多步子任务,结合自检与反思提升推理质量
  • 生成报告在结构化、逻辑性上显著优于基线方法,评分更高
  • 支持多种场景和数据类型,适合军事战略评估等高可靠性需求

模拟推演的数据分析与性能评估在现代战争中至关重要,可帮助军事人员洞察不同战略、战术及作战计划的潜在效果。传统人工分析耗时且易出错。为提高效率与准确性,利用具备强大分析与推理能力的大语言模型(LLMs)成为可能。然而,仅通过单一指令难以获得格式规范、内容高质量的分析报告。为此,我们提出一种方法:先将复杂任务分解为多个子任务,针对每个子任务设计有效的系统提示与用户提示;通过多轮交互并融入自我检查与反思机制,实现结构化数据提取及多步分析评估;同时定义并调用定制工具生成图表与计算指标;设计多种报告模板,适配不同应用场景与输入数据类型。大量评估结果表明,本方法生成的报告质量更高,得分显著优于基线方法。

原文摘要 · Abstract (English)

Data analysis and performance evaluation of simulation deduction plays a pivotal role in modern warfare, which enables military personnel to gain invaluable insights into the potential effectiveness of different strategies, tactics, and operational plans. Traditional manual analysis approach is time-consuming and limited by human errors. To enhance efficiency and accuracy, large language models (LLMs) with strong analytical and inferencing capabilities can be employed. However, high-quality analysis reports with well-structured formatting cannot be obtained through a single instruction input to the LLM. To tackle this issue, we propose a method that first decomposes the complex task into several sub-tasks and designs effective system prompts and user prompts for each sub-task. Multi-round interactions with the LLM incorporating self-check and reflection are then conducted to enable structured data extraction as well as multi-step analysis and evaluation. Furthermore, custom tools are defined and invoked to generate figures and compute metrics. We also design multiple report templates, each tailored to a specific application and input data type, ensuring their adaptability across a variety of scenarios. Extensive evaluation results demonstrate that the reports generated by our method exhibit higher quality, therefore obtaining higher scores than the baseline method.

大模型应用军事推演自动化分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。