用机械手玩德州扑克,测试真实世界操作与决策能力。
DexHoldem: Playing Texas Hold'em with Dexterous Embodied System

- 构建机械手在桌面上执行14种扑克操作的基准测试
- 任务完成率最高61.2%,场景保持成功率47.5%
- 揭示视觉感知与完整状态恢复间的差距,适合机器人决策研究
评估具身系统在真实灵巧硬件上的表现,不仅需要单一基础技能,还要求智能体能感知动态桌面场景、选择合适动作、用灵巧手执行,并保持场景可用以支持后续决策。我们提出DexHoldem,一个基于ShadowHand的德州扑克灵巧操作真实世界系统级基准。该基准包含1,470次远程操控示范、标准化物理策略评测和一个具身感知评测,用于检验智能体能否恢复决策所需的结构化游戏状态。在基础操作上,$π_{0.5}$获得最高任务完成率(61.2%),而$π_{0.5}$与$π_0$在场景保持成功率上并列最高(47.5%)。在具身感知方面,Opus 4.7在严格问题级准确率上最佳(34.3%),而GPT 5.5在平均字段级准确率上领先(66.8%),暴露出孤立视觉子能力与完整路由相关状态恢复之间的差距。最后,通过三个案例研究,我们展示了等待、恢复调度、请求人类帮助及重复执行等行为如何在闭环部署中累积感知与策略错误。DexHoldem因此在共享物理环境中评估灵巧桌面操作、具身感知与具身决策路由。
原文摘要 · Abstract (English)
Evaluating embodied systems on real dexterous hardware requires more than isolated primitive skills: an agent must perceive a changing tabletop scene, choose a context-appropriate action, execute it with a dexterous hand, and leave the scene usable for later decisions. We introduce DexHoldem, a real-world system-level benchmark built around Texas Hold'em dexterous manipulation with a ShadowHand. DexHoldem provides 1,470 teleoperated demonstrations across 14 Texas Hold'em manipulation primitives, a standardized physical policy benchmark, and an agentic perception benchmark that tests whether agents can recover the structured game state needed for embodied decision making. On primitive execution, $π_{0.5}$ obtains the highest task completion rate ($61.2\%$), while $π_{0.5}$ and $π_0$ tie on scene-preserving success rate ($47.5\%$). On agentic perception, Opus 4.7 obtains the best strict problem-level accuracy ($34.3\%$), while GPT 5.5 obtains the best average field-wise accuracy ($66.8\%$), exposing a gap between isolated visual sub-capabilities and complete routing-relevant state recovery. Finally, we instantiate the full embodied-agent loop in three case studies, where waiting, recovery dispatches, human-help requests, and repeated primitive execution reveal how perception and policy errors accumulate during closed-loop deployment. DexHoldem therefore evaluates dexterous tabletop execution, agentic perception, and embodied decision routing in a shared physical setting. Project page: https://dexholdem.github.io/Dexholdem/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。