让领域专家用LLM轻松处理数据质量问题,提升真实场景下机器学习落地效率。
Towards Human-Guided, Data-Centric LLM Co-Pilots
- 构建多智能体系统,结合规划协调与精准执行,实现动态数据处理
- 在医疗数据集上验证,显著优于现有共谋工具的数据清洗能力
- 支持人类介入,适合医疗、金融等需专业判断的领域应用
机器学习有望变革多个领域,但其应用常受限于领域专家需求与可落地模型之间的脱节。尽管现有基于大语言模型(LLM)的共谋系统已降低技术门槛,仍以模型为中心,忽视了真实世界中普遍存在的数据问题,如缺失值、标签噪声和领域特异性。为此,本文提出CliMB-DC:一个以人为本、以数据为中心的LLM共谋框架,融合先进数据处理工具与LLM推理能力,实现上下文感知的稳健数据处理。核心是一个多智能体推理系统,包含战略协调者用于动态规划与适应,以及专用执行者完成精确操作。通过人机协同方式系统性融入领域知识。我们构建了一个数据核心挑战分类体系,据此集成前沿数据工具到可扩展开源架构中,便于社区持续扩充。在真实医疗数据集上的实证表明,CliMB-DC能将未清理数据转化为机器学习可用格式,显著优于现有共谋基线。该框架可赋能医疗、金融、社会科学等领域专家主动推动实际应用。
原文摘要 · Abstract (English)
Machine learning (ML) has the potential to revolutionize various domains, but its adoption is often hindered by the disconnect between the needs of domain experts and translating these needs into robust and valid ML tools. Despite recent advances in LLM-based co-pilots to democratize ML for non-technical domain experts, these systems remain predominantly focused on model-centric aspects while overlooking critical data-centric challenges. This limitation is problematic in complex real-world settings where raw data often contains complex issues, such as missing values, label noise, and domain-specific nuances requiring tailored handling. To address this we introduce CliMB-DC, a human-guided, data-centric framework for LLM co-pilots that combines advanced data-centric tools with LLM-driven reasoning to enable robust, context-aware data processing. At its core, CliMB-DC introduces a novel, multi-agent reasoning system that combines a strategic coordinator for dynamic planning and adaptation with a specialized worker agent for precise execution. Domain expertise is then systematically incorporated to guide the reasoning process using a human-in-the-loop approach. To guide development, we formalize a taxonomy of key data-centric challenges that co-pilots must address. Thereafter, to address the dimensions of the taxonomy, we integrate state-of-the-art data-centric tools into an extensible, open-source architecture, facilitating the addition of new tools from the research community. Empirically, using real-world healthcare datasets we demonstrate CliMB-DC's ability to transform uncurated datasets into ML-ready formats, significantly outperforming existing co-pilot baselines for handling data-centric challenges. CliMB-DC promises to empower domain experts from diverse domains -- healthcare, finance, social sciences and more -- to actively participate in driving real-world impact using ML.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。