提出新方法提升视觉控制系统在动态干扰下的鲁棒性
Agent-Centric Observation Adaptation for Robust Visual Control under Dynamic Perturbations

- 基于信息瓶颈原理,只保留干净前景信息,避免干扰信息污染表征
- 在多种动态退化场景下,恢复95.3%原始性能,优于传统重建方法
- 无需标注、无参考图像,可零样本泛化到未见干扰类型
现实世界视觉系统面临天气变化、传感器噪声、压缩伪影和背景干扰等时变扰动。现有图像修复方法多针对固定退化类型,优化像素级保真度,但未解决非平稳退化切换下的表现问题,也未验证像素保真是否保留下游模型所需任务信息。为此,我们引入视觉退化控制基准(VDCS),在渲染场景中注入马尔可夫切换的物理退化。研究发现,基于重建的表征存在根本缺陷:忠实还原退化观测会迫使隐状态编码特定干扰信息,污染下游模型。从信息瓶颈视角看,将表征锚定于干净前景可消除此污染。据此,我们提出冻结的即插即用模块ACO-MoE,结合路由式修复专家与前景掩码分支。ACO-MoE 在合成渲染数据上离线预训练,使用自动生成的退化对与仿真生成的前景掩码,无需人工标注。推理时仅需输入退化RGB图像,无需退化标签、参考帧或前景掩码。在VDCS、DMC-GB和RoboSuite上,无论模型无关还是模型依赖的后端,均显著提升下游控制性能,在挑战性马尔可夫切换退化下恢复95.3%的纯净输入性能,并零样本泛化至预训练未涵盖的视觉扰动。
原文摘要 · Abstract (English)
Real-world visual systems face time-varying perturbations, including weather, sensor noise, compression artifacts, and background distractions. Existing image restoration methods are typically designed for fixed corruption types and optimized for pixel-level fidelity, leaving open two questions: how restoration behaves under non-stationary corruption switching, and whether pixel-level fidelity preserves the task-relevant information needed by downstream models. To study this setting, we introduce the Visual Degraded Control Suite (VDCS), a benchmark that injects Markov-switching physical degradations into rendered scenes. We further identify a fundamental failure mode of reconstruction-based representations: faithfully reconstructing corrupted observations forces the latent state to encode corruption-specific nuisance information, thereby contaminating downstream models. From an information-bottleneck perspective, anchoring the representation to the clean foreground eliminates this contamination. Motivated by this analysis, we propose \emph{Agent-Centric Observations with Mixture-of-Experts} (ACO-MoE), a frozen, plug-and-play observation adapter that combines a routed bank of restoration experts with a foreground-mask branch. ACO-MoE is pretrained entirely offline on synthetic rendered data with automatically generated degradation pairs and simulation-derived foreground masks, requiring no manual annotation. At inference time, it takes only corrupted RGB as input without corruption labels, clean reference frames, or foreground masks. Across VDCS, DMC-GB, and RoboSuite, ACO-MoE consistently improves downstream control with both model-free and model-based backbones, recovering 95.3\% of clean-input performance under challenging Markov-switching corruptions. It also generalizes zero-shot to unseen visual perturbations excluded from adapter pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。