arXiv:2510.12383cs.LG2025-10被引 1

跨模态数据错误检测新挑战:表格+图像协同识别更准

Towards Cross-Modal Error Detection with Tables and Images

  • 用多模态数据联合检测表格与图像中的不一致错误
  • 最佳组合在四个数据集上达最高F1分数,但重尾数据仍难处理
  • 适合关注电商、医疗等多模态数据质量的从业者

大规模数据质量保障仍是大型组织的持续挑战。尽管已有进展,但在涉及表格、图像和文本等多种数据模态的场景中,保持数据准确性和一致性依然复杂。传统错误检测方法通常仅针对单一模态(如表格),常忽略跨模态错误,而这类错误在电子商务和医疗等领域尤为常见。为此,我们首次针对表格数据中的跨模态错误检测开展基准测试,评估了四个数据集上的五种基线方法。结果表明,结合强大AutoML框架时,Cleanlab(标签错误检测框架)与DataScope(数据估值方法)表现最优,达到最高F1值。然而,当前方法在重尾分布的真实世界数据上仍受限,亟需进一步研究。

原文摘要 · Abstract (English)

Ensuring data quality at scale remains a persistent challenge for large organizations. Despite recent advances, maintaining accurate and consistent data is still complex, especially when dealing with multiple data modalities. Traditional error detection and correction methods tend to focus on a single modality, typically a table, and often miss cross-modal errors that are common in domains like e-Commerce and healthcare, where image, tabular, and text data co-exist. To address this gap, we take an initial step towards cross-modal error detection in tabular data, by benchmarking several methods. Our evaluation spans four datasets and five baseline approaches. Among them, Cleanlab, a label error detection framework, and DataScope, a data valuation method, perform the best when paired with a strong AutoML framework, achieving the highest F1 scores. Our findings indicate that current methods remain limited, particularly when applied to heavy-tailed real-world data, motivating further research in this area.

跨模态数据质量错误检测表格

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。