构建首个面向AI研究的细粒度论文问答数据集,支持多任务多模态评估。
AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation
- 人工标注13,956篇论文+1,246个问题,支持实例级评估
- 多数大模型在该数据集上表现不佳,验证其挑战性
- 自动生成对话轨迹框架,助力小模型提升工具使用能力
学术论文数量激增,研究人员高效提取关键信息愈发困难。尽管基于大语言模型(LLMs)的智能体可自动化科学论文问答流程,但缺乏全面且真实的评估基准。同时,训练交互式智能体受限于高质量交互轨迹数据的匮乏。为此,本文提出AirQA——一个面向人工智能领域的、人工标注的综合性论文问答数据集,包含13,956篇论文和1,246个问题,涵盖多任务、多模态及实例级评估。此外,提出ExTrActor自动指令数据合成框架,通过三个基于LLM的智能体实现无须人工干预的示例生成与轨迹收集。对多个开源及专有模型的评估显示,多数模型在AirQA上表现欠佳,证明了数据集的质量。大量实验表明,ExTrActor持续提升小型模型的多轮工具使用能力,使其性能接近大型模型。
原文摘要 · Abstract (English)
The growing volume of academic papers has made it increasingly difficult for researchers to efficiently extract key information. While large language models (LLMs) based agents are capable of automating question answering (QA) workflows for scientific papers, there still lacks a comprehensive and realistic benchmark to evaluate their capabilities. Moreover, training an interactive agent for this specific task is hindered by the shortage of high-quality interaction trajectories. In this work, we propose AirQA, a human-annotated comprehensive paper QA dataset in the field of artificial intelligence (AI), with 13,956 papers and 1,246 questions, that encompasses multi-task, multi-modal and instance-level evaluation. Furthermore, we propose ExTrActor, an automated framework for instruction data synthesis. With three LLM-based agents, ExTrActor can perform example generation and trajectory collection without human intervention. Evaluations of multiple open-source and proprietary models show that most models underperform on AirQA, demonstrating the quality of our dataset. Extensive experiments confirm that ExTrActor consistently improves the multi-turn tool-use capability of small models, enabling them to achieve performance comparable to larger ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。