CoupVisor优化玩家出牌与质疑决策,提升游戏胜率。
CoupVisor: Strategy Optimization by Round and Challenge Decision Support

- 基于角色概率与手牌数估算主张真实性
- 以最终胜利为目标的奖励机制表现最佳
- 适合策略游戏研究者与博弈算法开发者
本文提出CoupVisor,一个针对隐藏信息卡牌游戏Coup的决策支持系统。该系统解决两个核心问题:每回合应采取何种行动,以及何时应质疑对手的宣称。系统围绕统一的游戏事件描述构建,支持手动对战、回放记录、模拟、信念追踪、建议生成和基于学习的策略。CoupVisor通过结合各角色出现概率与宣称者剩余手牌数,估算主张为真的可能性,修正了早期游戏首次宣称即被标记为可疑的误判。我们在多种模拟对局中对比了规则导向顾问与多个学习型及启发式玩家,结果表明:奖励设计(短期收益或最终胜利)决定学习方法表现,以获胜为导向的奖励能生成超越所有基线的策略。
原文摘要 · Abstract (English)
This paper presents CoupVisor, a decision-support system for the hidden-information card game Coup. It addresses two questions: what a player should do on each turn, and when a player should challenge an opponent's claim. The system is built around a single description of game events, which is shared across manual play, replay of recorded games, simulation, belief tracking, advisor recommendations, and learning-based policies. CoupVisor estimates the chance that a claim is truthful by combining how likely each role is with how many cards the claimant still holds, which corrects a case where the very first claim of a game was flagged as suspicious despite no evidence. We compare a rule-following advisor and several learned and heuristic players across many simulated games and different opponent styles. Our main finding is that the choice of reward, whether it rewards short-term gains or ultimately winning the game, decides which learning approach performs best, and that a win-oriented reward produces a policy that outperforms all baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。