用多模态方法识别并去除数据中的冗余信息,提升模型效率。
ML-driven detection and reduction of ballast information in multi-modal datasets
- 融合熵、互信息等多指标,构建跨模态去冗余框架。
- 在稀疏数据中可删减超70%特征,分类性能反而提升。
- 适合需要压缩数据、加速训练的机器学习实践者。
现代数据集常包含冗余或低效信息(即‘球载’),增加维度、存储与计算成本却无实质分析价值。本文提出一种通用多模态框架,适用于结构化、半结构化、非结构化及稀疏数据,综合运用熵、互信息、Lasso、SHAP、PCA、主题建模与嵌入分析,识别并剔除冗余特征。提出新型‘球载评分’,将多源信号整合为统一的跨模态剪枝策略。实验表明,在稀疏或半结构化数据中,特征空间可削减超过70%,分类性能保持甚至提升,同时显著降低训练时间与内存占用。框架揭示了统计、语义、基础设施等不同类型的球载信息,并为构建更轻量高效的机器学习流程提供实用指导。
原文摘要 · Abstract (English)
Modern datasets often contain ballast as redundant or low-utility information that increases dimensionality, storage requirements, and computational cost without contributing meaningful analytical value. This study introduces a generalized, multimodal framework for ballast detection and reduction across structured, semi-structured, unstructured, and sparse data types. Using diverse datasets, entropy, mutual information, Lasso, SHAP, PCA, topic modelling, and embedding analysis are applied to identify and eliminate ballast features. A novel Ballast Score is proposed to integrate these signals into a unified, cross-modal pruning strategy. Experimental results demonstrate that significant portions of the feature space as often exceeding 70% in sparse or semi-structured data, can be pruned with minimal or even improved classification performance, along with substantial reductions in training time and memory footprint. The framework reveals distinct ballast typologies (e.g. statistical, semantic, infrastructural), and offers practical guidance for leaner, more efficient machine learning pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。