弱监督下自动学习视觉概念,提升对象中心推理鲁棒性
Weakly supervised concept Bottleneck Learning for Robust Two stage Object centric visual reasoning

- 动态分配概念维度,自动优化训练参数
- 低标签预算下保持高概念对齐与推理准确率
- 适合小样本视觉推理任务,可替代端到端模型
两阶段神经符号架构通过分离感知与关系推理,为视觉问题求解提供优雅范式。但高阶谓词的视觉锚定通常需要昂贵标注。本文提出动态正交概念瓶颈(D-OCB),一种面向对象的槽-变分自编码器框架,在极弱监督下提取人类对齐的符号谓词。D-OCB通过动态学习最优超参数分配,避免人工调参;结合重建自监督与跨概念子空间相关性惩罚,引入概念独立性先验。关键在于,为应对极低监督下的不稳定性,D-OCB采用动态维度分配机制:表现良好的概念主动释放潜空间维度给落后概念,有效防止表示坍缩,显著提升整体概念准确性。大量实验证明,该框架在极小标签预算下实现高概念对齐与下游视觉推理准确率,媲美或超越端到端模型。
原文摘要 · Abstract (English)
Two-stage neuro-symbolic architectures provide an elegant paradigm for visual problem solving by cleanly separating connectionist perception of predefined symbols from possibly later defined relational reasoning thereon. However, anchoring high-level predicates into visual frames typically necessitates annotations that are expensive to acquire. In this work, we introduce the Dynamic Orthogonal Concept Bottleneck (D-OCB), an object-centric slot- VAE framework designed to extract human-aligned symbolic predicates under extremely weak supervision. D-OCB eliminates the arduous manual tuning of loss-balancing coef- ficients by dynamically learning optimal hyperparameter allocations during training. To infuse prior knowledge on independence of concept categories, in addition to standard re- construction self-supervision we penalize correlation across concept subspaces. Crucially, to combat the instability of very low supervision regimes, D-OCB incorporates a dynamic di- mensionality allocation mechanism; this adaptive formulation allows well-represented con- cepts to yield latent dimensions to underperforming concepts that are lagging behind, effectively preventing representation collapse and significantly improving overall concept accuracy. Through an extensive empirical evaluation, we demonstrate that our framework achieves high concept alignment and downstream visual reasoning accuracy using minimal label budgets, matching or outperforming end-to-end paradigms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。