用专属嵌入和有限蒙特卡洛树搜索,让AI更像特定棋手。
Toward Modeling Player-Specific Chess Behaviors
- 为每位冠军设计专属嵌入,结合有限MCTS增强战术探索。
- 虽准确率下降,但16位世界冠军的风格相似度提升显著。
- 新提出基于JS散度的行为评估法,可区分不同棋手风格。
尽管人工智能在国际象棋上已达到超人水平,但建模人类棋手个体化决策风格仍是重大挑战。现有类人模型仅捕捉整体技能层级行为,无法复现具体历史冠军特征。标准走法准确率指标天然惩罚人类自然差异,忽视长期行为一致性,导致风格保真度评估不完整。为此,提出一种新架构:将统一的Maia-2模型适配至冠军专属嵌入,并融合有限蒙特卡洛树搜索(MCTS)以增强走法选择中的战术探索。为稳健评估,引入基于杰恩斯-申农散度(Jensen-Shannon divergence)的新行为度量。通过自编码器与均匀流形近似投影(UMAP)将高维棋盘表示压缩至潜在空间,将走法分布离散化于统一网格以比较行为相似性。16位历史世界冠军实验表明,虽引入MCTS后标准走法准确率下降,但新度量下平均JS散度显著降低,证明该方法能有效提升风格契合度,并成功区分个体棋手,为玩家与AI间行为对齐评估提供有力支持。
原文摘要 · Abstract (English)
While artificial intelligence has achieved superhuman performance in chess, developing models that accurately emulate the individualized decision-making styles of human players remains a significant challenge. Existing human-like chess models capture general population behaviors based on skill levels but fail to reproduce the behavioral characteristics of specific historical champions. Furthermore, the standard evaluation metric, move accuracy, inherently penalizes natural human variance and ignores long-term behavioral consistency, leading to an incomplete assessment of stylistic fidelity. To address these limitations, an architecture is proposed that adapts the unified Maia-2 model to champion-specific embeddings, further enhanced by the integration of a limited Monte Carlo Tree Search (MCTS) process to enrich tactical exploration during move selection. To robustly evaluate this approach, a novel behavioral metric based on the Jensen-Shannon divergence is introduced. By compressing high-dimensional board representations into a latent space using an AutoEncoder and Uniform Manifold Approximation and Projection (UMAP), move distributions are discretized on a common grid to compare behavioral similarities. Results across 16 historical world champions indicate that while integrating MCTS decreases standard move accuracy, it improves stylistic alignment according to the proposed metric, substantially reducing the average Jensen-Shannon divergence. Ultimately, the proposed metric successfully discriminates between individual players and provides promising evidence toward more comprehensive evaluations of behavioral alignment between players and AI models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。