arXiv:2604.03905cs.ROcs.AI2026-04被引 1

让异构机器人团队在不改策略的情况下自适应传感器差异。

DC-Ada: Reward-Only Decentralized Sensor Adaptation for Heterogeneous Multi-Robot Teams

论文配图:DC-Ada: Reward-Only Decentralized Sensor Adaptation for Heterogeneous Multi-Robot Teams
图 1 · 摘自论文原文
  • 用随机搜索调整每台机器人的感知映射,不更新主策略。
  • 在严重覆盖任务中提升完成度,仅需团队回报信号。
  • 无需梯度、无持续通信,适合部署时快速适配。

异构是实际多机器人团队的典型特征:平台在感知模态、范围、视场和故障模式上常有差异。在理想感知下训练的控制器,一旦部署到缺失或错配传感器的机器人上,性能会急剧下降,即使任务和动作接口不变。本文提出DC-Ada,一种仅依赖奖励的去中心化适应方法:保持预训练共享策略不变,仅通过紧凑的每机器人观测变换,将异构感知映射到固定推理接口。该方法无梯度、通信极少,采用预算化的接受/拒绝随机搜索,结合短周期公共随机数回放,在严格步数预算下运行。我们在确定性2D多机器人模拟器中评估了四种异构场景(H0–H3)及五次种子,每轮使用20万联合环境步数的匹配预算。结果表明,异构性显著降低冻结策略性能,且无单一缓解方法在所有任务和指标上占优。观测归一化在仓储物流中对奖励鲁棒性最强,搜救任务中表现也优异;而冻结策略在协同地图构建中奖励最高。DC-Ada提供互补优势:在严重覆盖任务中明显提升完成度,仅需标量团队回报,无需策略微调或持久通信。这些结果使DC-Ada成为异构团队实用的部署期适应方法。

原文摘要 · Abstract (English)

Heterogeneity is a defining feature of deployed multi-robot teams: platforms often differ in sensing modalities, ranges, fields of view, and failure patterns. Controllers trained under nominal sensing can degrade sharply when deployed on robots with missing or mismatched sensors, even when the task and action interface are unchanged. We present DC-Ada, a reward-only decentralized adaptation method that keeps a pretrained shared policy frozen and instead adapts compact per-robot observation transforms to map heterogeneous sensing into a fixed inference interface. DC-Ada is gradient-free and communication-minimal: it uses budgeted accept/reject random search with short common-random-number rollouts under a strict step budget. We evaluate DC-Ada against four baselines in a deterministic 2D multi-robot simulator covering warehouse logistics, search and rescue, and collaborative mapping, across four heterogeneity regimes (H0--H3) and five seeds with a matched budget of $200{,}000$ joint environment steps per run. Results show that heterogeneity can substantially degrade a frozen shared policy and that no single mitigation dominates across all tasks and metrics. Observation normalization is strongest for reward robustness in warehouse logistics and competitive in search and rescue, while the frozen shared policy is strongest for reward in collaborative mapping. DC-Ada offers a useful complementary operating point: it improves completion most clearly in severe coverage-based mapping while requiring only scalar team returns and no policy fine-tuning or persistent communication. These results position DC-Ada as a practical deploy-time adaptation method for heterogeneous teams.

多机器人异构系统自适应去中心化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。