用虚拟人模拟心理评估,训练对话时机与情感节奏。
Learning When to Ask: Simulation-Trained Humanoids for Mental-Health Diagnosis
- 构建276个带表情动作的虚拟患者,模拟真实访谈场景。
- 自研TD3模型在对话完整性和节奏上优于其他算法。
- 适合临床心理评估机器人研发者参考使用。
测试人形机器人与用户互动耗时长、易损耗,限制迭代与多样性。但筛查代理需掌握对话时机、语调、回应提示及面部与语音关注点,以诊断抑郁和创伤后应激障碍。多数模拟器忽略非语言行为策略学习;许多控制器仅追求任务准确率,忽视信任、节奏与关系建立。本文将人形机器人作为对话代理,在无硬件负担下进行训练。采用代理中心、仿真优先的流程,将访谈数据转化为276个同步语音、眼神、面部与头身姿态的Unreal Engine MetaHuman患者,并集成PHQ-8与PCL-C评估流。感知融合策略循环决定何时说话、何时回应及如何避免打断,受安全屏障保护。训练使用反事实回放(有限非语言扰动)与不确定性感知轮次管理器,主动探查以减少诊断模糊性。结果为纯仿真,人形机器人即为迁移目标。对比三种控制器,定制TD3优于PPO与CEM,实现接近天花板的覆盖率,节奏更稳定且奖励相当。决策质量分析显示几乎无话语重叠、切口时间对齐、澄清提问少、等待时间短。性能在模态缺失与渲染器更换下仍稳定,持留患者集上排名一致。贡献包括:(1) 以代理为中心的模拟器,生成276个具边界非语言反事实的交互患者;(2) 将时机与关系纳入第一类控制变量的安全学习环;(3) 对比研究(TD3 vs PPO/CEM),显著提升完整度与社交时机;(4) 消融与鲁棒性分析揭示优势来源,支持临床监督下的机器人试运行。
原文摘要 · Abstract (English)
Testing humanoid robots with users is slow, causes wear, and limits iteration and diversity. Yet screening agents must master conversational timing, prosody, backchannels, and what to attend to in faces and speech for Depression and PTSD. Most simulators omit policy learning with nonverbal dynamics; many controllers chase task accuracy while underweighting trust, pacing, and rapport. We virtualise the humanoid as a conversational agent to train without hardware burden. Our agent-centred, simulation-first pipeline turns interview data into 276 Unreal Engine MetaHuman patients with synchronised speech, gaze/face, and head-torso poses, plus PHQ-8 and PCL-C flows. A perception-fusion-policy loop decides what and when to speak, when to backchannel, and how to avoid interruptions, under a safety shield. Training uses counterfactual replay (bounded nonverbal perturbations) and an uncertainty-aware turn manager that probes to reduce diagnostic ambiguity. Results are simulation-only; the humanoid is the transfer target. In comparing three controllers, a custom TD3 (Twin Delayed DDPG) outperformed PPO and CEM, achieving near-ceiling coverage with steadier pace at comparable rewards. Decision-quality analyses show negligible turn overlap, aligned cut timing, fewer clarification prompts, and shorter waits. Performance stays stable under modality dropout and a renderer swap, and rankings hold on a held-out patient split. Contributions: (1) an agent-centred simulator that turns interviews into 276 interactive patients with bounded nonverbal counterfactuals; (2) a safe learning loop that treats timing and rapport as first-class control variables; (3) a comparative study (TD3 vs PPO/CEM) with clear gains in completeness and social timing; and (4) ablations and robustness analyses explaining the gains and enabling clinician-supervised humanoid pilots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。