让大模型像专家一样攻克超复杂问题,通过深度广度结合的科研流程。
Super Research: Answering Highly Complex Questions with Large Language Models through Super Deep and Super Wide Research
- 将复杂问题拆解为研究计划,结合超广检索与深度迭代查询。
- 需100+次检索、1000+网页,整合矛盾信息生成可验证报告。
- 适合评估大模型真实科研能力,是衡量其通用推理水平的标尺。
尽管大型语言模型在深度研究或广域搜索中已展现能力,但其解决高度复杂问题——需要长周期规划、大规模证据搜集及跨异构来源综合——的能力仍基本未被探索。我们提出「Super Research」任务,整合三项核心:(i) 结构化分解形成研究计划,(ii) 超广度检索获取多元视角,(iii) 超深度调查通过迭代查询化解不确定性。为评估该能力,我们构建了300个专家撰写的跨领域问题基准,每个问题需最多100+次检索步骤和1000+网页以调和冲突证据。Super Research生成带有细粒度引用和中间产物(如大纲、表格)的可验证报告,确保推理可追溯。此外,我们提出基于图锚定的审计协议,从覆盖度、逻辑一致性、报告实用性、客观性与引用健康度五个维度评估。尽管超复杂问题在常规应用中罕见,但Super Research作为关键天花板测试,能有效检验大模型的综合研究能力;在此任务中的成功,是其具备强健泛化研究能力的重要指标。排行榜详见:https://cnsdqd-dyb.github.io/Super-Research-Benchmark/
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have demonstrated proficiency in Deep Research or Wide Search, their capacity to solve highly complex questions-those requiring long-horizon planning, massive evidence gathering, and synthesis across heterogeneous sources-remains largely unexplored. We introduce Super Research, a task for complex autonomous research tasks that integrates (i) structured decomposition into a research plan, (ii) super wide retrieval for diverse perspectives, and (iii) super deep investigation to resolve uncertainties through iterative queries. To evaluate this capability, we curated a benchmark of 300 expert-written questions across diverse domains, each requiring up to 100+ retrieval steps and 1,000+ web pages to reconcile conflicting evidence. Super Research produces verifiable reports with fine-grained citations and intermediate artifacts (e.g., outlines and tables) to ensure traceable reasoning. Furthermore, we present a graph-anchored auditing protocol that evaluates Super Research along five dimensions: Coverage, Logical Consistency, Report Utility, Objectivity and Citation Health. While super-complex questions may be infrequent in standard applications, Super Research serves as a critical ceiling evaluation and stress test for LLM capabilities. A model's proficiency within Super Research acts as a powerful proxy for its general research competence; success here suggests the robustness necessary to navigate nearly any subordinate research task. Leaderboard is available at: https://cnsdqd-dyb.github.io/Super-Research-Benchmark/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。