解决视觉模型中罕见语义概念学习不公的问题
SemCovNet: Towards Fair and Semantic Coverage-Aware Learning for Underrepresented Visual Concepts
- 设计动态加权机制,平衡视觉与语义特征学习
- 引入覆盖率差异指数,量化并降低语义公平性偏差
- 适合关注模型公平性与可解释性的研究者
现代视觉模型依赖丰富的语义表示,超越类别标签,涵盖描述性概念和上下文属性。然而,现有数据集存在语义覆盖不平衡(SCI),一种此前被忽视的长尾语义偏差。与类别不平衡不同,SCI发生于语义层面,影响模型对稀有但有意义语义的学习与推理。为此,我们提出语义覆盖感知网络(SemCovNet),显式纠正语义覆盖差异。该模型包含语义描述符图(SDM)用于学习语义表示,描述符注意力调制(DAM)模块动态加权视觉与概念特征,并引入描述符-视觉对齐(DVA)损失,使视觉特征与描述语义对齐。我们通过覆盖率差异指数(CDI)量化语义公平性,衡量覆盖与误差间的匹配度。多数据集实验证明,SemCovNet显著提升模型可靠性,大幅降低CDI,实现更公平、更均衡的表现。本工作将SCI确立为可测量、可修正的偏差,为推进语义公平性和可解释视觉学习奠定基础。
原文摘要 · Abstract (English)
Modern vision models increasingly rely on rich semantic representations that extend beyond class labels to include descriptive concepts and contextual attributes. However, existing datasets exhibit Semantic Coverage Imbalance (SCI), a previously overlooked bias arising from the long-tailed semantic representations. Unlike class imbalance, SCI occurs at the semantic level, affecting how models learn and reason about rare yet meaningful semantics. To mitigate SCI, we propose Semantic Coverage-Aware Network (SemCovNet), a novel model that explicitly learns to correct semantic coverage disparities. SemCovNet integrates a Semantic Descriptor Map (SDM) for learning semantic representations, a Descriptor Attention Modulation (DAM) module that dynamically weights visual and concept features, and a Descriptor-Visual Alignment (DVA) loss that aligns visual features with descriptor semantics. We quantify semantic fairness using a Coverage Disparity Index (CDI), which measures the alignment between coverage and error. Extensive experiments across multiple datasets demonstrate that SemCovNet enhances model reliability and substantially reduces CDI, achieving fairer and more equitable performance. This work establishes SCI as a measurable and correctable bias, providing a foundation for advancing semantic fairness and interpretable vision learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。