arXiv:2412.09439cs.CV2024-12

提升视觉模型在开放环境下的公平性与鲁棒性,解决多视角、持续学习等挑战。

Towards Robust and Fair Vision Learning in Open-World Environments

  • 提出公平域适应与跨视角几何自适应方法,减少数据依赖和视图偏差。
  • 构建开放世界公平持续学习框架,实现动态场景下模型性能稳定提升。
  • 针对大规模视频多模态数据,设计基于Transformer的鲁棒特征学习方案。

论文围绕开放世界环境下视觉学习的公平性与鲁棒性,提出四项关键贡献:首先,基于双射最大似然与公平适应学习,提出新型公平域适应方法,缓解大规模数据需求问题;其次,构建开放世界公平持续学习框架,融合公平持续学习与开放世界持续学习研究路径;第三,针对多摄像头视图数据,提出基于几何的跨视图自适应框架,学习跨视图不变特征;最后,面向大规模视频与多模态数据,提出基于Transformer的鲁棒特征表示方法,并设计新型领域泛化方法以增强视觉基础模型的鲁棒性。理论分析与实验结果表明,所提方法显著优于已有方法,推动了机器视觉学习在公平性与鲁棒性方面的进展。

原文摘要 · Abstract (English)

The dissertation presents four key contributions toward fairness and robustness in vision learning. First, to address the problem of large-scale data requirements, the dissertation presents a novel Fairness Domain Adaptation approach derived from two major novel research findings of Bijective Maximum Likelihood and Fairness Adaptation Learning. Second, to enable the capability of open-world modeling of vision learning, this dissertation presents a novel Open-world Fairness Continual Learning Framework. The success of this research direction is the result of two research lines, i.e., Fairness Continual Learning and Open-world Continual Learning. Third, since visual data are often captured from multiple camera views, robust vision learning methods should be capable of modeling invariant features across views. To achieve this desired goal, the research in this thesis will present a novel Geometry-based Cross-view Adaptation framework to learn robust feature representations across views. Finally, with the recent increase in large-scale videos and multimodal data, understanding the feature representations and improving the robustness of large-scale visual foundation models is critical. Therefore, this thesis will present novel Transformer-based approaches to improve the robust feature representations against multimodal and temporal data. Then, a novel Domain Generalization Approach will be presented to improve the robustness of visual foundation models. The research's theoretical analysis and experimental results have shown the effectiveness of the proposed approaches, demonstrating their superior performance compared to prior studies. The contributions in this dissertation have advanced the fairness and robustness of machine vision learning.

视觉学习公平性鲁棒性持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。