arXiv:2502.03368cs.AIcs.DB2025-02被引 12

用聊天方式轻松构建复杂AI数据流水线,让非程序员也能做数据分析。

PalimpChat: Declarative and Interactive AI analytics

  • 通过自然语言对话生成由大模型驱动的AI处理流程
  • 支持在生物医学数据上完成提取与分析,无需编程基础
  • 适合科研、法律和地产等领域的非技术用户快速探索数据

得益于生成式架构和大语言模型的发展,数据科学家现在可以编写机器学习操作流水线来处理大规模非结构化数据。近年来,声明式AI框架(如Palimpzest、Lotus和DocETL)兴起,能够构建高效且日益复杂的流水线,但这些系统通常仅对专业程序员开放。本文展示的PalimpChat是一个基于聊天的Palimpzest接口,通过自然语言即可创建并运行复杂AI流水线,弥合了这一差距。系统整合了基于ReAct的推理代理Archytas与Palimpzest的关联型及大模型操作符,展示了聊天界面如何使声明式AI框架真正面向非专家用户。演示系统已公开上线。在SIGMOD'25上,参与者可体验三个真实场景——科学发现、法律发现和房地产搜索——或使用PalimpChat处理自己的数据集。本文重点阐述在Palimpzest优化器支持下,PalimpChat如何简化如生物医学数据提取与分析等复杂工作流。

原文摘要 · Abstract (English)

Thanks to the advances in generative architectures and large language models, data scientists can now code pipelines of machine-learning operations to process large collections of unstructured data. Recent progress has seen the rise of declarative AI frameworks (e.g., Palimpzest, Lotus, and DocETL) to build optimized and increasingly complex pipelines, but these systems often remain accessible only to expert programmers. In this demonstration, we present PalimpChat, a chat-based interface to Palimpzest that bridges this gap by letting users create and run sophisticated AI pipelines through natural language alone. By integrating Archytas, a ReAct-based reasoning agent, and Palimpzest's suite of relational and LLM-based operators, PalimpChat provides a practical illustration of how a chat interface can make declarative AI frameworks truly accessible to non-experts. Our demo system is publicly available online. At SIGMOD'25, participants can explore three real-world scenarios--scientific discovery, legal discovery, and real estate search--or apply PalimpChat to their own datasets. In this paper, we focus on how PalimpChat, supported by the Palimpzest optimizer, simplifies complex AI workflows such as extracting and analyzing biomedical data.

AI流水线自然语言数据挖掘大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。