让大模型像数据科学家一样自主探索数据库,发现关键洞察。
Hunt Instead of Wait: Evaluating Deep Data Research on Large Language Models
- 设计开放任务DDR,让大模型自主从数据库中挖掘信息。
- 前沿模型具备初步自主性,但长期探索仍存挑战。
- 适合研究智能体自主性与数据分析能力的学者。
大语言模型的智能不应仅限于准确回答问题,更需具备自主设定目标和决定探索方向的能力。我们称之为调查式智能,区别于仅完成指定任务的执行式智能。数据科学是天然的测试场景,因为真实分析始于原始数据而非明确查询,但现有评估基准极少关注此领域。为此,我们提出深度数据研究(DDR)这一开放性任务,使大模型能自主从数据库中提取关键见解,并构建了大规模、基于清单的评测基准DDR-Bench,支持可验证评估。结果表明,尽管前沿模型已展现出初步自主性,但长周期探索仍具挑战。分析显示,有效的调查式智能不仅依赖智能体框架或单纯规模扩展,更取决于智能体自身的内在策略。
原文摘要 · Abstract (English)
The agency expected of Agentic Large Language Models goes beyond answering correctly, requiring autonomy to set goals and decide what to explore. We term this investigatory intelligence, distinguishing it from executional intelligence, which merely completes assigned tasks. Data Science provides a natural testbed, as real-world analysis starts from raw data rather than explicit queries, yet few benchmarks focus on it. To address this, we introduce Deep Data Research (DDR), an open-ended task where LLMs autonomously extract key insights from databases, and DDR-Bench, a large-scale, checklist-based benchmark that enables verifiable evaluation. Results show that while frontier models display emerging agency, long-horizon exploration remains challenging. Our analysis highlights that effective investigatory intelligence depends not only on agent scaffolding or merely scaling, but also on intrinsic strategies of agentic models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。