用AI代理自动整合多源数据,提升临床数据标准化效率。
Interactive Data Harmonization with LLM Agents: Opportunities and Challenges
- 构建基于大模型的交互式系统,自动设计数据融合流程
- 在临床数据场景中实现标准格式转换,支持可复用管道生成
- 适合数据科学家与领域专家协作,降低数据整合门槛
数据整合是将来自不同来源的数据进行统一的重要任务。尽管该领域已有多年研究,但仍因模式不匹配、术语差异和采集方法不同而耗时且困难。本文提出以智能体方式推进数据整合,展示一种名为Harmonia的系统,结合大模型推理、交互界面和数据整合原语库,自动合成数据整合流水线。我们在临床数据整合场景中验证该系统,能够帮助用户交互式创建可复用的标准化转换管道。最后,我们讨论当前挑战与开放问题,并提出未来研究方向。
原文摘要 · Abstract (English)
Data harmonization is an essential task that entails integrating datasets from diverse sources. Despite years of research in this area, it remains a time-consuming and challenging task due to schema mismatches, varying terminologies, and differences in data collection methodologies. This paper presents the case for agentic data harmonization as a means to both empower experts to harmonize their data and to streamline the process. We introduce Harmonia, a system that combines LLM-based reasoning, an interactive user interface, and a library of data harmonization primitives to automate the synthesis of data harmonization pipelines. We demonstrate Harmonia in a clinical data harmonization scenario, where it helps to interactively create reusable pipelines that map datasets to a standard format. Finally, we discuss challenges and open problems, and suggest research directions for advancing our vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。