无人机雷达搜索中动态环境下的智能策略选择
SEArch: Optimistic Policy Selection Between Scene Noise and Drift for UAV Radar Search

- 基于多策略在线选择,自适应环境变化
- 实测最多降低30%搜索误差,提升鲁棒性
- 适合资源受限的实时无人机目标探测场景
搭载雷达的无人机在复杂环境中执行目标搜索任务时,目标特征(如人体呼吸微动)可穿透遮挡被检测。但随着无人机移动,雷达统计特性随环境动态变化,固定处理策略难以奏效;而感知与适应必须在资源受限的机载节点上实时完成。由于单一检测器无法应对所有情况,本文采用多策略范式,将搜索问题建模为在线策略选择,以累计损失与最优策略之差(即遗憾)衡量性能。该设置同时包含场景内随机噪声和跨场景漂移。现有方法仅能处理其中一种情形,本文提出广义随机对抗(SEA)框架,无需场景动态先验知识即可同时建模两类变化。为适配机载计算约束,设计轻量级乐观正则化领袖跟踪(OFTRL)选择器 extsc{SEArch},实现 $O(\barσ_T \sqrt{T} + \sqrt{J})$ 的遗憾界,其中 $\barσ_T$ 表示雷达测量噪声,$J$ 为任务周期 $T$ 内的场景切换次数。为进一步应对频繁场景变化,提出窗口化版本 extsc{W-SEArch},每 $w$ 轮重启一次,单窗口内最多一次切换时遗憾为 $O(\barσ_I \sqrt{w})$。实验表明,在多种非平稳环境下,相比非自适应基线,最多减少30%遗憾。
原文摘要 · Abstract (English)
Unmanned Aerial Vehicles (UAVs) equipped with radar sensors are deployed for target search missions in diverse environments, where targets exhibit characteristic signatures (e.g., respiration micro-motion in human search) detectable through occlusions. A fundamental challenge arises from shifts in radar statistics as the UAV moves through a dynamic and potentially non-stationary environment, rendering any fixed signal-processing strategy suboptimal; yet perception and adaptation must run onboard a resource-constrained aerial node in real time. Since no single detector performs well across all conditions, we adopt a multi-policy paradigm and formulate UAV target search as an online policy selection problem over a library of specialized detectors, with performance measured by regret, the cumulative loss gap relative to the best policy in each scene. The setting couples in-scene stochastic noise with inter-scene shifts. Whereas prior methods capture only one regime, we account for both through the Stochastically Extended Adversary (SEA) framework, without requiring oracle knowledge of scene dynamics. Because adaptation must run at the UAV, we instantiate SEA through \textsc{SEArch}, a lightweight optimistic Follow the Regularized Leader (OFTRL) selector with an adaptive learning rate, achieving regret $O(\barσ_T \sqrt{T} + \sqrt{J})$, where $\barσ_T$ captures radar measurement noise and $J$ is the number of scene transitions over the mission horizon $T$. To enable rapid adaptation under frequent scene changes, we further introduce \textsc{W-SEArch}, a windowed variant that restarts every $w$ rounds and achieves regret $O(\barσ_I \sqrt{w})$ under at most one transition per window. Experiments show up to 30\% regret reduction compared to non-adaptive baselines across a range of non-stationary settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。