自动识别数据中的日期格式,提升可视化前处理效率
Automating Date Format Detection for Data Visualization
- 基于最小熵与自然语言建模,自动推断日期格式
- 在大规模数据集上准确率超90%,最小熵方法响应极快
- 适合集成到可视化工具和数据库中,降低人工干预
数据分析中的日期解析是工作流中的主要瓶颈。为此,我们提出两种算法:一种基于最小熵,另一种基于自然语言建模,可从字符串数据中自动推断日期格式。这两种算法在大规模数据列语料库上准确率超过90%,显著简化了可视化环境中的数据准备流程。其中最小熵方法计算速度快,支持交互式反馈。所提方法易于集成至数据可视化工具和数据库系统,有效减少手动处理需求。
原文摘要 · Abstract (English)
Data preparation, specifically date parsing, is a significant bottleneck in analytic workflows. To address this, we present two algorithms, one based on minimum entropy and the other on natural language modeling that automatically derive date formats from string data. These algorithms achieve over 90% accuracy on a large corpus of data columns, streamlining the data preparation process within visualization environments. The minimal entropy approach is particularly fast, providing interactive feedback. Our methods simplify date format extraction, making them suitable for integration into data visualization tools and databases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。