用采样控制框架实现大规模蜂群的少步精准引导
Learning Sampled-data Control for Swarms via MeanFlow
- 直接在控制空间学习分段最优控制系数,适配有限更新场景
- 通过桥接信号导出可梯度化的回归目标,高效训练控制策略
- 保持线性动力学约束,适合通信受限的大规模群体控制
由于通信或计算限制,大规模蜂群常需以有限控制更新进行引导,但现有基于学习的方法多建模瞬时速度场,忽略了实际中控制为有限窗口量的事实。为此,本文将近期的MeanFlow机器学习框架推广至一般线性动态系统,提出一种直接作用于控制空间的采样数据学习新框架,用于蜂群引导。该方法学习每个时间区间内最小能量控制所对应的有限时域系数,并推导出其与局部桥接诱导监督信号之间的微分恒等式,从而得到一个简单的停梯度回归目标,可从桥接样本中高效学习区间系数场。所学策略通过采样数据更新部署,确保控制器严格满足预设的线性时不变动力学与执行通道约束。该方法可在少量控制步骤下实现大规模蜂群引导,且与底层控制系统的有限窗口作用结构一致。
原文摘要 · Abstract (English)
Steering large-scale swarms with only limited control updates is often needed due to communication or computational constraints, yet most learning-based approaches do not account for this and instead model instantaneous velocity fields. As a result, the natural object for decision making is a finite-window control quantity rather than an infinitesimal one. To address this gap, we consider the recent machine learning framework MeanFlow and generalize it to the setting with general linear dynamic systems. This results in a new sampled-data learning framework that operates directly in control space and that can be applied for swarm steering. To this end, we learn the finite-horizon coefficient that parameterizes the minimum-energy control applied over each interval, and derive a differential identity that connects this quantity to a local bridge-induced supervision signal. This identity leads to a simple stop-gradient regression objective, allowing the interval coefficient field to be learned efficiently from bridge samples. The learned policy is deployed through sampled-data updates, guaranteeing that the resulting controller exactly respects the prescribed linear time-invariant dynamics and actuation channel. The resulting method enables few-step swarm steering at scale, while remaining consistent with the finite-window actuation structure of the underlying control system.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。