用物体-概念-关系三元组提升模型在开放世界中的泛化能力
OCRT: Boosting Foundation Models in the Open World with Object-Concept-Relation Triad
- 通过无监督解耦与迭代优化,从视觉输入中提取高阶概念与关系结构
- 在多个下游任务上显著提升SAM和CLIP的泛化性与鲁棒性
- 适合需要跨域适应与抗干扰的视觉基础模型应用
尽管基础模型(FMs)声称具备强大能力,但在面对分布偏移、弱监督或恶意攻击时,其泛化性能显著下降。现有领域泛化或对抗微调方法多为任务或模型特定,忽视了实际应用中的通用性及模型间的可迁移性。本文研究基础模型在域外数据上的泛化问题,提出新颖框架OCRT(物体-概念-关系三元组),使模型能从原始视觉输入中提取稀疏的高层概念与复杂关系结构。核心思想是通过无监督解耦与迭代精炼,将视觉场景中的物体与一组以物体为中心的表征绑定。具体而言,将物体中心表征投影至模型可理解的语义概念空间,并估计其重要性以过滤无关元素;随后构建具有灵活度的概念图,融合概念及其重要性,从而从信息丰富的概念中提取高阶因素,并促进概念间的关系推理。大量实验表明,OCRT可在多个下游任务中显著增强SAM和CLIP的泛化能力与鲁棒性。
原文摘要 · Abstract (English)
Although foundation models (FMs) claim to be powerful, their generalization ability significantly decreases when faced with distribution shifts, weak supervision, or malicious attacks in the open world. On the other hand, most domain generalization or adversarial fine-tuning methods are task-related or model-specific, ignoring the universality in practical applications and the transferability between FMs. This paper delves into the problem of generalizing FMs to the out-of-domain data. We propose a novel framework, the Object-Concept-Relation Triad (OCRT), that enables FMs to extract sparse, high-level concepts and intricate relational structures from raw visual inputs. The key idea is to bind objects in visual scenes and a set of object-centric representations through unsupervised decoupling and iterative refinement. To be specific, we project the object-centric representations onto a semantic concept space that the model can readily interpret and estimate their importance to filter out irrelevant elements. Then, a concept-based graph, which has a flexible degree, is constructed to incorporate the set of concepts and their corresponding importance, enabling the extraction of high-order factors from informative concepts and facilitating relational reasoning among these concepts. Extensive experiments demonstrate that OCRT can substantially boost the generalizability and robustness of SAM and CLIP across multiple downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。