无需干净参考数据,用智能体自动清洗数据并权衡修复与安全性的实验研究。
Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs

- 构建包含推理、证据检索与保守修复的智能体框架,实现无参考数据清洗。
- 最优配置在合成数据上检测F1达0.421,但无方案在所有指标上均最优。
- 强调修复安全性与可追溯性,适合高可靠性场景如金融与医疗数据处理。
在缺乏可信干净参考数据的情况下进行数据清洗极具挑战,异常值可能既是错误也是有效观测。本文研究不同智能体能力对无参考数据清洗的影响,提出一个基于证据的框架,融合结构化上下文、数据画像、大模型推理、可执行检查、受控证据检索、源排序、引用对齐、保守修复、可逆脚本和溯源日志。在金融、临床与环境监测数据集上,通过受控合成污染和原始数据描述性分析,完成126次实验。评估包含两个基线及逐步加入可执行工具、证据检索、证据控制与保守修复的LLM序列。合成评估中,确定性画像基线取得最高检测F1为0.561;在基于LLM的配置中,完整保守配置获得最高F1 0.421,但无配置在所有指标上表现最佳。源排序配置支持率最低,决策级引用对齐仍较弱。完整保守配置未产生任何不安全或不必要的修改,尽管这些率在添加保守策略前已为零,且其未执行直接修复。结果表明,新增能力带来检测、修复、证据依附、保守行为、可复现性与操作成本之间的权衡,而非一致提升。研究提供了结构化框架与实证方法,用于评估无参考智能体清洗中的权衡。
原文摘要 · Abstract (English)
Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observations. This paper studies how different agent capabilities affect reference-free data cleaning and proposes an evidence-grounded framework that combines structured context, profiling, LLM reasoning, executable checks, controlled evidence retrieval, source ranking, citation alignment, conservative repair, reversible scripts, and provenance logging. Seven configurations are evaluated across financial, clinical, and environmental-monitoring datasets using controlled synthetic corruption and original-data descriptive analysis, resulting in 126 completed runs. The evaluation includes two comparison baselines and a progressive LLM-based sequence that adds executable tools, evidence retrieval, evidence controls, and conservative repair. In the synthetic evaluation, the deterministic profiling baseline achieved the highest detection F1-score of 0.561. Among the LLM-based configurations, the full conservative configuration achieved the highest F1-score of 0.421, but no configuration performed best across all evaluation criteria. The source-ranked configurations achieved the lowest unsupported-rule rates, while decision-level citation alignment remained weak. The full conservative configuration produced no unsafe or unnecessary modifications, although these rates were already zero before the conservative policy was added, and it performed no direct repairs. Overall, the results show that additional capabilities introduce trade-offs among detection, repair, evidence grounding, conservative behaviour, reproducibility, and operational cost rather than producing consistent improvements. The study provides a structured framework and empirical methodology for evaluating these trade-offs in reference-free agentic data cleaning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。