用强化学习构建可模拟多动物行为的动态仿真器,支持真实轨迹复现与假设推演。
Data-driven simulator of multi-animal behavior with unknown dynamics via offline and online reinforcement learning
- 基于强化学习估计未知运动动力学,将不完整模型转为动作空间优化
- 在果蝇、蝾螈等物种上实现更高行为复现率和更强奖励获取能力
- 支持新实验场景下的反事实行为预测,适合复杂群体行为研究
动物运动仿真在行为研究中具有重要价值。尽管机器人模仿学习进展推动了人类与动物运动的再现,但生物领域多动物仿真仍面临真实运动转换模型未知与模拟版本脱节的核心挑战。由于运动动力学通常未知,仅依赖数学模型难以满足需求;如何在复现真实轨迹的同时支持奖励驱动优化仍是开放问题。本文提出一种基于深度强化学习与反事实仿真的数据驱动多动物行为仿真器。通过将不完整转换模型中的运动变量作为强化学习框架内的动作进行估计,缓解高自由度带来的病态问题。同时引入基于距离的伪奖励,对齐数字空间与物理空间状态。在人工智能代理、果蝇、蝾螈和家蚕的验证中,该方法在物种特异性行为复现度和奖励获取方面优于标准模仿学习与强化学习方法。此外,能实现新实验条件下的反事实行为预测,并支持多个体建模,灵活生成假设轨迹,展现出模拟与解析复杂多动物行为的潜力。
原文摘要 · Abstract (English)
Simulators of animal movements play a valuable role in studying behavior. Advances in imitation learning for robotics have expanded possibilities for reproducing human and animal movements. A key challenge for realistic multi-animal simulation in biology is bridging the gap between unknown real-world transition models and their simulated counterparts. Because locomotion dynamics are seldom known, relying solely on mathematical models is insufficient; constructing a simulator that both reproduces real trajectories and supports reward-driven optimization remains an open problem. We introduce a data-driven simulator for multi-animal behavior based on deep reinforcement learning and counterfactual simulation. We address the ill-posed nature of the problem caused by high degrees of freedom in locomotion by estimating movement variables of an incomplete transition model as actions within an RL framework. We also employ a distance-based pseudo-reward to align and compare states between cyber and physical spaces. Validated on artificial agents, flies, newts, and silkmoth, our approach achieves higher reproducibility of species-specific behaviors and improved reward acquisition compared with standard imitation and RL methods. Moreover, it enables counterfactual behavior prediction in novel experimental settings and supports multi-individual modeling for flexible what-if trajectory generation, suggesting its potential to simulate and elucidate complex multi-animal behaviors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。