用语言控制国际象棋策略,让电脑下棋更像人。
UniMaia: Steering Chess Policies with Language for Human-like Play

- 用文本编码器+控制网络调节预训练棋力模型,实现语义化控制
- 在提示词指令任务中表现超越现有方法,预测准确率领先
- 适合想定制棋风或研究人机交互的开发者和棋手
大型语言模型使自然语言成为控制复杂系统的新接口,但常需大规模多模态训练或牺牲领域先验。在国际象棋等结构化决策领域,专用策略网络性能强但缺乏语义可控性;而提示条件的语言模型虽灵活,却往往领域基础薄弱。本文提出UniMaia框架,通过参数高效文本编码器与类ControlNet的条件机制,调节冻结的Lc0棋力模型,实现开局选择、棋力强度等语义控制,同时保留预训练表示。进一步提出UniMaia-Aux,引入辅助时序条件与行为预测目标。构建大规模带元数据的Lichess数据集,开发半自动提示生成流程,并设立涵盖提示与元数据条件的基准测试。UniMaia在多个提示条件基准上达到最佳预期准确率,在通用指令跟随任务中保持竞争力,且在人类走法预测任务上优于专用元数据方法。UniMaia-Aux进一步提升预期准确率与行为建模能力,仅小幅降低最优走法准确率。结果表明,无需端到端多模态训练即可实现领域策略的提示控制,但需权衡可控性与预测性能。
原文摘要 · Abstract (English)
Recent advances in large language models have enabled natural language to serve as a flexible interface for controlling complex systems, but often at the cost of large-scale multimodal training or weakened domain-specific inductive biases. In structured decision-making domains such as chess, specialized policy networks achieve strong performance but lack semantic controllability, while prompt-conditioned language models are more flexible yet typically exhibit weaker domain grounding. We propose $\textbf{UniMaia}$, a framework for prompt-conditioned policy modulation that adapts a frozen Lc0-based chess policy network using a parameter-efficient text encoder and a ControlNet-style conditioning mechanism. UniMaia enables semantic control over gameplay, including opening selection and player strength, while preserving the pretrained policy representations. We further introduce $\textbf{UniMaia-Aux}$, which incorporates auxiliary temporal conditioning and behavioral prediction objectives. To support this work, we construct a large-scale metadata-augmented Lichess dataset, develop a semi-automated prompt-generation pipeline, and introduce benchmarks spanning both prompt-conditioned and metadata-conditioned settings. UniMaia achieves state-of-the-art expected accuracy on several prompt-conditioned benchmarks and competitive top-move accuracy on general instruction-following tasks, while remaining competitive with dedicated metadata-conditioned approaches on human move prediction benchmarks. UniMaia-Aux further improves expected accuracy and behavioral modeling across several evaluation settings, with modest trade-offs in top-move accuracy. Overall, our results demonstrate that prompt-conditioned control of domain-specific policy networks is feasible without end-to-end multimodal training, while highlighting trade-offs between controllability and predictive performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。