arXiv:2601.14862cs.LGcs.CL2026-01

让AI推理更符合战略逻辑,提升长期预测准确性。

Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting

  • 融合多文档注意力与时间编码,加入教义一致性约束层
  • 在127个历史反事实中实现比基线模型更高的预测准确率
  • 适合军事、外交等需长期战略推演的场景

我们提出战略教义语言模型(sdLM),一种用于多文档战略推理的学习系统框架,包含教义一致性约束和校准不确定性。该方法结合多文档注意力、时间编码和教义一致性层,在长周期预测和计划合理性方面表现更优,显著减少严重教义违背。我们在三类基准上评估:47个专家评分的战略情景、336份教义文献(共12,847条陈述)的教义一致性,以及1945-2020年间127个历史反事实事件在12至60个月周期的地缘政治预测。sdLM在各项指标上均优于强通用大模型基线,且在长期判断上与人类专家相当。我们还报告了消融实验、缩放趋势及面向部署的性能/延迟特性,揭示各组件贡献并指导实际应用。

原文摘要 · Abstract (English)

We introduce Strategic Doctrine Language Models (sdLM), a learning-system framework for multi-document strategic reasoning with doctrinal consistency constraints and calibrated uncertainty. The approach combines multi-document attention, temporal encoding, and a doctrine-consistency layer to improve long-horizon forecasting and plan plausibility while reducing severe doctrinal violations. We evaluate sdLM using (i) expert-panel scoring of strategic scenarios (N=47), (ii) doctrine consistency on 336 doctrine publications (12,847 statements), and (iii) geopolitical forecasting on 127 historical counterfactuals (1945-2020) across 12-60 month horizons. Across these benchmarks, sdLM achieves higher strategic quality and better calibration than strong general-purpose LLM baselines, and remains competitive with human experts on long-horizon judgments. We further report ablations, scaling trends, and deployment-oriented performance/latency characteristics to clarify which components drive improvements and how they translate to operational settings.

战略推理地缘预测一致性约束

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。