arXiv:2605.24779cs.LGcs.AI2026-05

提出兼顾选中与未选数据结构的新型信息度量,提升数据选择的平衡性与鲁棒性。

Complement Submodular Information Measures for Balanced and Robust Data Selection

论文配图:Complement Submodular Information Measures for Balanced and Robust Data Selection
图 1 · 摘自论文原文
  • 设计互补感知的子模目标函数,显式建模选中集与剩余集间的结构关联。
  • 在隐藏切片任务中显著提升稀有语义结构保留能力,同时抑制噪声异常点。
  • 适用于需均衡划分的数据筛选场景,如训练/验证集构建与鲁棒子集选择。

子模优化已成为数据选择、检索、摘要和表示学习的基础范式,因其能建模覆盖度、多样性与代表性。然而,经典子模目标仅优化选中子集,未显式保留选中集与剩余数据间的结构信息。在现代机器学习应用中,如训练/验证/测试集划分、基准构建和鲁棒子集选择,选择质量关键取决于选中集与补集之间的结构平衡。本文提出互补子模信息(CSI),一类新的互补感知子模目标,用于量化子集与其补集间的共享结构信息。该框架衍生出多个经典子模函数的互补感知变体,包括设施选址、图割、LogDet、饱和覆盖、集合覆盖、概率集合覆盖及基于特征的函数。我们分析了CSI目标的理论性质,证明在曲率有界条件下具备近似单调性,从而获得接近(1−1/e)的贪心近似保证。实验表明,CSI目标在鲁棒的隐藏切片感知子集选择任务中持续优于标准子模目标,显著提升了稀有/尾部语义结构的保留能力,同时有效抑制噪声与孤立异常点,大幅改善下游预测性能。合成实验进一步说明不同CSI实例可捕捉互补的代表性、多样性、连通性及均衡邻域保持等概念。

原文摘要 · Abstract (English)

Submodular optimization has become a fundamental paradigm for data selection, retrieval, summarization, and representation learning due to its ability to model coverage, diversity, and representativeness. However, classical submodular objectives optimize only the selected subset and do not explicitly preserve structural information between the selected subset and the remaining data. In many modern machine learning applications, including train/validation/test splitting, benchmark construction, and robust subset selection, the quality of a selection depends critically on preserving balanced structure across both the selected subset and its complement. In this work, we introduce Complement Submodular Information (CSI), a new class of complement-aware submodular objectives that quantify shared structural information between a subset and its complement. Our framework induces complement-aware variants of several classical submodular functions including Facility Location, Graph Cut, LogDet, Saturated Coverage, Set Cover, Probabilistic Set Cover, and Feature Based Functions. We analyze the theoretical properties of CSI objectives and show that they exhibit approximate monotonicity under bounded curvature conditions, leading to near-$(1-1/e)$ greedy approximation guarantees. Empirically, CSI objectives consistently outperform standard submodular objectives on robust hidden-slice-aware subset selection. In particular, CSI objectives significantly improve preservation of coherent rare/tail semantic structure while simultaneously suppressing noisy and isolated outliers, leading to substantially improved downstream predictive performance. Synthetic experiments further illustrate how different CSI instantiations capture complementary notions of representativeness, diversity, connectivity, and balanced neighborhood preservation.

数据选择子模优化鲁棒性结构平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。