arXiv:2608.28409cs.RO2026-08

让机器人团队按价值分配风险,用利他策略减少重复探索。

Cooperative Risk-Aware Exploration in Heterogeneous Multi-Robot Systems Using Algorithmic Altruism

论文配图:Cooperative Risk-Aware Exploration in Heterogeneous Multi-Robot Systems Using Algorithmic Altruism
图 1 · 摘自论文原文
  • 基于生态利他行为设计博弈框架,按成员价值分配风险。
  • 仿真与实验证明能降低重复探索,提升机器人间距与任务效率。
  • 适合有危险环境探索需求的异构机器人团队使用。

多机器人系统在危险环境中执行探索任务具有优势,但有效部署需决定信息采集位置及风险在异构成员间的分布。本文提出一种基于生态启发式利他行为的博弈论框架,用于合作式风险感知探索。每个机器人选择有限时域轨迹,以最大化信息增益,同时惩罚冗余探索和预期危害暴露。通过代理特异性价值参数引入异质性,利他耦合由受汉密尔顿法则启发的相关性权重建模。提出一种轨迹规划的博弈结构,定义了社会纳什均衡,根据代理相关性调整行动效用。该效用塑造使代理内化自身轨迹选择对队友的影响,促使低价值机器人承担风险以利于高价值成员并提升团队表现。定义探索效用函数,奖励区域覆盖与不确定性降低,同时惩罚冗余与风险,支持基于投影梯度的递推时域规划器进行航点优化。仿真显示,利他规划减少了冗余探索,改善了机器人间分离度,并根据代理价值重新分配风险,同时保持相当的地图覆盖率。进一步在硬件实验中验证,规划航点由轮式机器人通过单积分器控制器与屏障证书实现跟踪。

原文摘要 · Abstract (English)

Multi-robot systems are well-positioned for exploration in hazardous environments, but effective deployment requires deciding not only where robots should gather information, but also how risk should be distributed across heterogeneous team members. This paper develops a game-theoretic framework for cooperative risk-aware exploration based on ecologically inspired altruistic behavior. Each robot selects a finite-horizon trajectory to maximize information gain while penalizing redundant exploration and expected hazard exposure. Heterogeneity is introduced through agent-specific value parameters for encoding altruistic coupling, which is modeled through relatedness weights inspired by Hamilton's rule. We introduce a game-theoretic structure for trajectory planning that defines a Social Nash Equilibrium, which modifies the utility of agent actions according to agent relatedness. This utility shaping causes agents to internalize the effect of their trajectory choices on teammates, encouraging lower-valued robots to accept risk when doing so benefits higher-valued agents and improves team performance. We define an exploration utility for agents that rewards area coverage and uncertainty reduction, while also penalizing redundancy and risk, enabling projected gradient-based waypoint optimization in a receding-horizon planner. Simulations show that altruistic planning reduces redundant exploration, improves inter-robot separation, and reallocates risk according to agent value while maintaining comparable map coverage. We further demonstrate the approach in hardware experiments, where planned waypoints are tracked by wheeled robots using single-integrator controllers and barrier certificates.

多机器人风险分配利他行为探索算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。