Nomad能自动发现数据中隐藏的洞察,突破人类提问的局限。
Nomad: Autonomous Exploration and Discovery
- 构建探索地图,系统性遍历数据以平衡广度与深度
- 生成并验证假设,产出有数据支撑、可行动的报告
- 适合需要深度洞察但不知从何入手的研究者
我们提出Nomad,一个自主数据探索与洞察发现系统。面对文档、数据库等数据源,用户常无法预知所有可探究的问题或关联。传统基于查询或提示的研究受限于人类设定框架,难以覆盖全面洞察空间。Nomad采用探索优先架构,构建领域内的显式探索地图,并系统遍历以平衡广度与深度。它生成并筛选假设,由探查代理利用文档检索、网页搜索和数据库工具进行验证。候选洞察经独立验证后进入报告流水线,生成带引用的报告及更高层次的元报告。我们还提出一套全面评估框架,衡量可信度、报告质量与多样性。在联合国与世卫组织报告、以及关于大模型智能体的arXiv论文数据集上,实验显示Nomad产出的报告具备更强的数值依据,整体质量和可行动性优于基线,且多轮运行中展现更丰富的洞察多样性。Nomad迈向了不仅能回答问题、还能主动发现值得探索的研究方向的自主系统。
原文摘要 · Abstract (English)
We introduce Nomad, a system for autonomous data exploration and insight discovery. Given a corpus of documents, databases, or other data sources, users rarely know the full set of questions, hypotheses, or connections that could be explored. As a result, query-driven question answering and prompt-driven deep-research systems remain limited by human framing and often fail to cover the broader insight space. Nomad addresses this problem with an exploration-first architecture. It constructs an explicit Exploration Map over the domain and systematically traverses it to balance breadth and depth. It generates and selects hypotheses and investigates them with an explorer agent that can use document search, web search, and database tools. Candidate insights are then checked by an independent verifier before entering a reporting pipeline that produces cited reports and higher-level meta-reports. We also present a comprehensive evaluation framework for autonomous discovery systems that measures trustworthiness, report quality, and diversity. Using corpora of selected UN and WHO reports and arXiv papers on LLM agents, we show that Nomad produces reports with strong numeric grounding, higher overall quality and actionability than baselines, and more diverse insights over several runs. Nomad is a step toward autonomous systems that not only answer user questions or conduct directed research, but also discover which questions, research directions, and insights are worth surfacing in the first place.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。