让AI推理更符合战略逻辑,提升长期预测准确性。
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
- 融合多文档注意力与时间编码,加入教义一致性约束层
- 在127个历史反事实中实现比基线模型更高的预测准确率
- 适合军事、外交等需长期战略推演的场景
我们提出战略教义语言模型(sdLM),一种用于多文档战略推理的学习系统框架,包含教义一致性约束和校准不确定性。该方法结合多文档注意力、时间编码和教义一致性层,在长周期预测和计划合理性方面表现更优,显著减少严重教义违背。我们在三类基准上评估:47个专家评分的战略情景、336份教义文献(共12,847条陈述)的教义一致性,以及1945-2020年间127个历史反事实事件在12至60个月周期的地缘政治预测。sdLM在各项指标上均优于强通用大模型基线,且在长期判断上与人类专家相当。我们还报告了消融实验、缩放趋势及面向部署的性能/延迟特性,揭示各组件贡献并指导实际应用。
原文摘要 · Abstract (English)
We introduce Strategic Doctrine Language Models (sdLM), a learning-system framework for multi-document strategic reasoning with doctrinal consistency constraints and calibrated uncertainty. The approach combines multi-document attention, temporal encoding, and a doctrine-consistency layer to improve long-horizon forecasting and plan plausibility while reducing severe doctrinal violations. We evaluate sdLM using (i) expert-panel scoring of strategic scenarios (N=47), (ii) doctrine consistency on 336 doctrine publications (12,847 statements), and (iii) geopolitical forecasting on 127 historical counterfactuals (1945-2020) across 12-60 month horizons. Across these benchmarks, sdLM achieves higher strategic quality and better calibration than strong general-purpose LLM baselines, and remains competitive with human experts on long-horizon judgments. We further report ablations, scaling trends, and deployment-oriented performance/latency characteristics to clarify which components drive improvements and how they translate to operational settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。