用图结构显式建模语义关系,让数据筛选更透明精准
Mapping the Concept Landscape: Structural Perception of Global Distributions for Transparent Data Pruning

- 将图文对转为实体-事件-属性图,构建全局语义结构
- 通过概念稀有度评估,优先保留罕见语义样本
- 可解释性强,适合需要透明数据筛选的场景
现有数据剪枝方法主要依赖高维特征嵌入衡量样本重要性,但压缩向量常掩盖细粒度语义交互,导致稀有语义概念覆盖不足。本文提出映射概念景观(MCL)框架,不使用抽象嵌入,而是将每个图像-文本对表示为包含实体、事件和属性的显式样本级图。通过将这些个体图整合为整体数据集级图,刻画语义概念的全局分布并量化其在全语料中的稀有程度。基于此结构化感知,设计贪心概念覆盖率最大化算法,迭代选择能最大提升高价值、低频概念覆盖率的样本。在多个基准上的实验表明,该方法不仅优于当前最优剪枝方法,且提供可解释的选择审计路径。
原文摘要 · Abstract (English)
Existing data pruning methods predominantly rely on high-dimensional feature embeddings to measure sample importance. However, these compressed vectors often obscure fine-grained semantic interactions, leading to suboptimal coverage of rare semantic concepts in the pruned subsets. In this paper, we propose Mapping the Concept Landscape (MCL), a novel structural perception framework for transparent data pruning. Instead of abstract embeddings, we represent each image-caption pair as an explicit sample-level graph comprising entities, events, and attributes. By integrating these individual graphs into a comprehensive dataset-level graph, we characterize the global distribution of semantic concepts and quantify their rarity across the entire corpus. Based on this structured perception, we develop a greedy concept-coverage maximization algorithm that iteratively selects samples to maximize the marginal gain of high-value, under-represented concepts. Experimental results on various benchmarks demonstrate that our method not only achieves superior pruning efficiency compared to state-of-the-art methods but also provides a transparent and interpretable audit trail for the selection process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。