提出上下文属性密度模型,显著提升细粒度计数准确率
Exploring Contextual Attribute Density in Referring Expression Counting
- 设计U形密度估计算法,融合文本与多尺度视觉特征
- 实现计数误差降低30%,定位准确率提升10%
- 适合关注视觉-语言对齐与细粒度理解的研究者
指称表达计数(REC)算法旨在实现多样化细粒度文本描述下的灵活交互式计数。然而,现有方法因难以精准对齐属性信息与视觉模式而受限。鉴于视觉密度的重要性,我们推测当前方法的不足源于对‘上下文属性密度’(CAD)探索不足。在REC任务中,我们将CAD定义为特定细粒度属性在视觉区域中的信息密度。为此,我们提出一种U形CAD估计算法,使指称表达与GroundingDINO提取的多尺度视觉特征相互作用,并引入密度监督以有效编码CAD。随后通过一种基于CAD优化查询的新注意力机制解码。该框架显著优于现有SOTA方法,在计数指标上实现30%误差降低,在定位准确率上提升10%。结果揭示了上下文属性密度在REC中的关键作用。代码将发布于github.com/Xu3XiWang/CAD-GD。
原文摘要 · Abstract (English)
Referring expression counting (REC) algorithms are for more flexible and interactive counting ability across varied fine-grained text expressions. However, the requirement for fine-grained attribute understanding poses challenges for prior arts, as they struggle to accurately align attribute information with correct visual patterns. Given the proven importance of ''visual density'', it is presumed that the limitations of current REC approaches stem from an under-exploration of ''contextual attribute density'' (CAD). In the scope of REC, we define CAD as the measure of the information intensity of one certain fine-grained attribute in visual regions. To model the CAD, we propose a U-shape CAD estimator in which referring expression and multi-scale visual features from GroundingDINO can interact with each other. With additional density supervision, we can effectively encode CAD, which is subsequently decoded via a novel attention procedure with CAD-refined queries. Integrating all these contributions, our framework significantly outperforms state-of-the-art REC methods, achieves $30\%$ error reduction in counting metrics and a $10\%$ improvement in localization accuracy. The surprising results shed light on the significance of contextual attribute density for REC. Code will be at github.com/Xu3XiWang/CAD-GD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。