arXiv:2604.01669cs.CVcs.AI2026-04中稿 · ICME2026

无需先验领域信息,实现机器人感知在动态环境中的持续适应。

Robust Embodied Perception in Dynamic Environments via Disentangled Weight Fusion

  • 解耦表征消除环境风格干扰,聚焦跨场景语义特征。
  • 动态权重融合避免遗忘旧知识,不存储历史数据仍保持高精度。
  • 适用于未知交互场景,对鲁棒性要求高的机器人系统特别有用。

具身感知系统在开放物理空间中持续交互时面临动态环境分布漂移的严峻挑战。现有领域增量感知方法通常依赖测试阶段预先获取的领域标识,限制了其在未知交互场景中的实用性。同时,模型易过拟合于特定上下文感知噪声,导致泛化能力不足和灾难性遗忘。为此,我们提出一种无需领域标识和样本的增量学习框架,旨在实现具身多媒体系统的鲁棒连续环境适应。该方法设计了解耦表征机制,消除非必要环境风格干扰,引导模型提取跨场景共享的语义内在特征,从而减少感知不确定性并提升泛化能力。进一步采用权重融合策略,在参数空间中动态整合新旧环境知识,确保模型在不存储历史数据的前提下适应新分布,并最大限度保留旧环境的判别能力。多组标准基准数据集上的大量实验表明,所提方法在完全无样本、无领域标识设置下显著降低灾难性遗忘,性能优于现有最先进方法。

原文摘要 · Abstract (English)

Embodied perception systems face severe challenges of dynamic environment distribution drift when they continuously interact in open physical spaces. However, the existing domain incremental awareness methods often rely on the domain id obtained in advance during the testing phase, which limits their practicability in unknown interaction scenarios. At the same time, the model often overfits to the context-specific perceptual noise, which leads to insufficient generalization ability and catastrophic forgetting. To address these limitations, we propose a domain-id and exemplar-free incremental learning framework for embodied multimedia systems, which aims to achieve robust continuous environment adaptation. This method designs a disentangled representation mechanism to remove non-essential environmental style interference, and guide the model to focus on extracting semantic intrinsic features shared across scenes, thereby eliminating perceptual uncertainty and improving generalization. We further use the weight fusion strategy to dynamically integrate the old and new environment knowledge in the parameter space, so as to ensure that the model adapts to the new distribution without storing historical data and maximally retains the discrimination ability of the old environment. Extensive experiments on multiple standard benchmark datasets show that the proposed method significantly reduces catastrophic forgetting in a completely exemplar-free and domain-id free setting, and its accuracy is better than the existing state-of-the-art methods.

具身智能增量学习鲁棒感知解耦表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。