用机器人鱼与活鱼互动,量化评估行为模型的准确性。
Robots that learn to evaluate models of collective behavior

- 用强化学习训练机器人鱼在闭环中与活鱼互动
- 神经网络模型在目标达成等指标上差距最小
- 适合研究动物行为建模与仿生机器人的人
理解与建模动物行为对群体运动、决策及生物启发式机器人研究至关重要。然而,当前行为模型的评估仍多依赖于离线静态轨迹统计对比。本文提出一种基于强化学习的框架,利用仿生机器人鱼(RoboFish)通过闭环交互评估计算模型的活鱼行为。我们在仿真中训练了四种鱼模型的策略:一个简单常量跟随基线、两个规则模型,以及一个基于生物机制的卷积神经网络模型,并将这些策略迁移到真实机器人鱼系统中,与真实鱼群互动。策略训练目标是引导模拟鱼到达指定目标位置,从而量化真实鱼与模拟鱼响应的差异。我们通过计算模拟与真实行为度量分布间的Wasserstein距离来评估模型,包括目标达成率、个体间距离、墙接触频率和对齐程度等。结果显示,神经网络模型在目标达成和其他多数指标上的模拟到现实差距最小,表明其行为保真度高于传统规则模型。更重要的是,该方法能在相同闭环条件下定量区分候选模型。本工作展示了基于学习的机器人实验如何揭示行为模型的缺陷,并提供了一个通过具身交互评估动物行为模型的通用框架。
原文摘要 · Abstract (English)
Understanding and modeling animal behavior is essential for studying collective motion, decision-making, and bio-inspired robotics. Yet, evaluating the accuracy of behavioral models still often relies on offline comparisons to static trajectory statistics. Here we introduce a reinforcement-learning-based framework that uses a biomimetic robotic fish (RoboFish) to evaluate computational models of live fish behavior through closed-loop interaction. We trained policies in simulation using four distinct fish models-a simple constant-follow baseline, two rule-based models, and a biologically grounded convolutional neural network model-and transferred these policies to the real RoboFish setup, where they interacted with live fish. Policies were trained to guide a simulated fish to goal locations, enabling us to quantify how the response of real fish differs from the simulated fish's response. We evaluate the fish models by quantifying the sim-to-real gaps, defined as the Wasserstein distance between simulated and real distributions of behavioral metrics such as goal-reaching performance, inter-individual distances, wall interactions, and alignment. The neural network-based fish model exhibited the smallest gap across goal-reaching performance and most other metrics, indicating higher behavioral fidelity than conventional rule-based models under this benchmark. More importantly, this separation shows that the proposed evaluation can quantitatively distinguish candidate models under matched closed-loop conditions. Our work demonstrates how learning-based robotic experiments can uncover deficiencies in behavioral models and provides a general framework for evaluating animal behavior models through embodied interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。