让机器人实时判断何时该干预,精准助人而不多管闲事。
MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention

- 用信念表+触发器实现动态环境下的持续推理与决策
- 36.63% 精准干预率,远超基准模型的12.05%
- 适合需要智能助手的医疗、养老等场景
心智理论(ToM)使具身智能体能够理解人类信念、目标和意图,但现有评估多基于离线问答或场景级动作预测。MindPower虽推进了从感知到行动的机器人中心式推理,却未检验智能体在动态环境中持续交互并仅在需要时干预的能力。为此,我们提出MindHelper挑战,将具身ToM评估拓展至实时闭环精准干预。智能体需持续观察环境、维护个体化信念、识别人类是否需要帮助、生成可执行动作,并在无需干预时保持沉默。我们进一步提出MindClaw框架——一种简单有效的爪式结构,整合个体化信念表、具身认知技能与基于触发器的认知调度器。实验表明,MindClaw实现36.63%精准干预率与14.36%任务准确率,显著优于直接使用视觉语言模型(VLM)的基线,后者分别低于12.05%与3.80%。
原文摘要 · Abstract (English)
Theory-of-Mind (ToM) reasoning enables embodied agents to understand human beliefs, goals, and intentions, but existing benchmarks mainly evaluate this ability through offline question answering or scenario-level action prediction. MindPower advances embodied ToM by introducing robot-centric reasoning from perception to action; however, it does not evaluate whether an agent can continuously interact with a changing environment and intervene only when assistance is needed. Building on MindPower, we introduce the MindHelper Challenge, which extends embodied ToM evaluation to real-time closed-loop precision intervention. An agent must continuously observe the environment, maintain actor-specific beliefs, identify when a human requires assistance, generate executable actions, and remain silent when intervention is unnecessary. We further propose MindClaw, a simple yet effective Claw-style framework that integrates an actor-specific Belief Table, embodied cognitive skills, and a Trigger-based cognitive dispatcher. Experiments show that MindClaw achieves 36.63% precise intervention rate and 14.36\% task accuracy, substantially outperforming direct VLM baselines, whose corresponding results remain below 12.05% and 3.80%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。