通过属性语义分析提升数据质量检测准确率,发现81处遗漏值。
Attribute-Based Semantic Type Detection and Data Quality Assessment
- 基于属性名语义与规则库分类23种数据类型
- 在922个属性中识别出81个缺失值,优于现有工具
- 适合数据清洗、质量评估人员使用
跨领域数据驱动决策依赖高质量数据,但现有方法未能充分利用属性标签中的语义信息,导致数据质量问题持续存在。本文提出一种基于属性语义类型检测与数据质量评估的新方法,结合属性名语义、规则分析及格式缩写词典,构建包含约23种类型的实用语义分类体系,涵盖非负数、类别、ID、姓名、字符串、地理、时间、URL、IP、邮箱、二进制及年龄、百分比等有界数值类型。与当前先进系统Sherlock对比,本方法在分类鲁棒性和质量评估适用性上表现更优。对来自UCI机器学习资源库的50个数据集进行详尽分析,验证了该方法在识别潜在数据质量问题方面的有效性。相比YData Profiling,本方法在922个属性中检测到81个缺失值,而后者仅发现1个。
原文摘要 · Abstract (English)
The reliance on data-driven decision-making across sectors highlights the critical need for high-quality data; despite advancements, data quality issues persist, significantly impacting business strategies and scientific research. Current data quality methods fail to leverage the semantic richness embedded in words inside attribute labels (or column names/headers in tables) across diverse datasets and domains, leaving a crucial gap in comprehensive data quality evaluation. This research addresses this gap by introducing an innovative methodology centered around Attribute-Based Semantic Type Detection and Data Quality Assessment. By leveraging semantic information within attribute labels, combined with rule-based analysis and comprehensive Formats and Abbreviations dictionaries, our approach introduces a practical semantic type classification system comprising approximately 23 types, including numerical non-negative, categorical, ID, names, strings, geographical, temporal, and complex formats like URLs, IP addresses, email, and binary values plus several numerical bounded types, such as age and percentage. A comparative analysis with Sherlock, a state-of-the-art Semantic Type Detection system, shows the advantages of our approach in terms of classification robustness and applicability to data quality assessment tasks. Our research focuses on well-known data quality issues and their corresponding data quality dimension violations, grounding our methodology in a robust academic framework. Detailed analysis of fifty distinct datasets from the UCI Machine Learning Repository showcases our method's proficiency in identifying potential data quality issues. Compared to established tools like YData Profiling, our method exhibits superior accuracy, detecting 81 missing values across 922 attributes where YData identified only one.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。