用多智能体框架让大模型从乱说变可追溯的分析伙伴
DataPuzzle: Breaking Free from the Hallucinated Promise of LLMs in Data Analysis
- 拆解复杂问题,分步调用专精智能体提取信息
- 输出表格图谱等结构化结果,支持逐层验证
- 适合需要可审计分析的科研与工业场景
大型语言模型(LLMs)在多模态数据分析中应用日益广泛,因其能提供流畅灵活的接口来解释复杂输入。然而,这种流畅性常掩盖深层结构性缺陷:主流的‘提示-答案’范式将LLM视为黑箱分析师,把证据、推理和结论合并为单一不可见的回答,导致结果脆弱、不可验证且常具误导性。我们主张根本性转变:从生成转向结构化提取,从单体提示转向模块化、基于代理的工作流。LLM不应作为预言家,而应作为协作工具——在透明流程中承担信息提取、翻译、关联等专项任务。我们提出DataPuzzle,一个概念性的多智能体框架,可分解复杂问题,将信息组织为可解释的形式(如表格、图谱),并协调智能体角色以支持透明且可验证的分析。该框架为重建大模型驱动分析中的可见性与控制力提供了愿景蓝图——将模糊的答案转化为可追溯的流程,将脆弱的流畅性转变为可问责的洞见。这不是微小改进,而是重思大模型时代可信、可审计分析系统构建方式的呼吁。结构不是束缚,而是清晰之路。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly applied to multi-modal data analysis -- not necessarily because they offer the most precise answers, but because they provide fluent, flexible interfaces for interpreting complex inputs. Yet this fluency often conceals a deeper structural failure: the prevailing ``Prompt-to-Answer'' paradigm treats LLMs as black-box analysts, collapsing evidence, reasoning, and conclusions into a single, opaque response. The result is brittle, unverifiable, and frequently misleading. We argue for a fundamental shift: from generation to structured extraction, from monolithic prompts to modular, agent-based workflows. LLMs should not serve as oracles, but as collaborators -- specialized in tasks like extraction, translation, and linkage -- embedded within transparent workflows that enable step-by-step reasoning and verification. We propose DataPuzzle, a conceptual multi-agent framework that decomposes complex questions, structures information into interpretable forms (e.g. tables, graphs), and coordinates agent roles to support transparent and verifiable analysis. This framework serves as an aspirational blueprint for restoring visibility and control in LLM-driven analytics -- transforming opaque answers into traceable processes, and brittle fluency into accountable insight. This is not a marginal refinement; it is a call to reimagine how we build trustworthy, auditable analytic systems in the era of large language models. Structure is not a constraint -- it is the path to clarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。