让AI同时用表格和文本做深度研究,报告更准更全。
Towards Knowledgeable Deep Research: Framework and Benchmark
- 设计多智能体框架HKA,融合结构化与非结构化知识
- 在9个领域41个专家问题上超越现有模型,尤其擅长图表分析
- 适合需要精准数据支持的科研与决策类任务
深度研究(DR)要求大模型智能体自主完成多步信息检索、处理与推理,生成综合性报告。现有研究多聚焦非结构化网页内容,而更具挑战性的任务应结合结构化知识,提供数据基础,支持定量计算,实现深入分析。本文提出新型任务——有知深度研究(KDR),要求智能体在报告中同时使用结构化与非结构化知识。为此,我们构建了混合知识分析框架(HKA),一种多智能体架构,可对两类知识进行推理,并将文本、图表整合为连贯的多模态报告。关键设计是结构化知识分析器,利用代码模型与视觉-语言模型生成图表及对应洞察。为支持系统评估,我们构建了KDR-Bench,涵盖9个领域,包含41个专家级问题,整合大量结构化知识资源(如1,252张表格)。我们还为每题标注主要结论与关键点,并提出三类评估指标:通用型、知识导向型与视觉增强型。实验表明,HKA在通用与知识导向指标上持续优于多数现有DR智能体,甚至在视觉增强指标上超越Gemini DR智能体,凸显其在结构化知识深度分析上的有效性。本工作有望成为未来结构化知识分析与多模态深度研究的基石。
原文摘要 · Abstract (English)
Deep Research (DR) requires LLM agents to autonomously perform multi-step information seeking, processing, and reasoning to generate comprehensive reports. In contrast to existing studies that mainly focus on unstructured web content, a more challenging DR task should additionally utilize structured knowledge to provide a solid data foundation, facilitate quantitative computation, and lead to in-depth analyses. In this paper, we refer to this novel task as Knowledgeable Deep Research (KDR), which requires DR agents to generate reports with both structured and unstructured knowledge. Furthermore, we propose the Hybrid Knowledge Analysis framework (HKA), a multi-agent architecture that reasons over both kinds of knowledge and integrates the texts, figures, and tables into coherent multimodal reports. The key design is the Structured Knowledge Analyzer, which utilizes both coding and vision-language models to produce figures, tables, and corresponding insights. To support systematic evaluation, we construct KDR-Bench, which covers 9 domains, includes 41 expert-level questions, and incorporates a large number of structured knowledge resources (e.g., 1,252 tables). We further annotate the main conclusions and key points for each question and propose three categories of evaluation metrics including general-purpose, knowledge-centric, and vision-enhanced ones. Experimental results demonstrate that HKA consistently outperforms most existing DR agents on general-purpose and knowledge-centric metrics, and even surpasses the Gemini DR agent on vision-enhanced metrics, highlighting its effectiveness in deep, structure-aware knowledge analysis. Finally, we hope this work can serve as a new foundation for structured knowledge analysis in DR agents and facilitate future multimodal DR studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。