用人类与AI协作方式,找出未来最具影响力的科研问题。
HybridQuestion: Human-AI Collaboration for Identifying High-Impact Research Questions
- AI处理海量文献生成基础信息,再由多模型投票提出候选问题
- 人类逐步介入筛选,最终选出2025年十大突破与2026年十大关键问题
- AI在识别已知突破上表现接近人类,但在预测未来问题时仍需人判断
AI科学家范式正在重塑科研流程,涵盖从选题到写作的各个环节。然而,一个核心问题仍未明确:AI能否识别有意义的研究问题?尽管大语言模型在特定任务中表现出色,但其对过往重大突破和未来研究方向的战略性、长期性评估潜力仍待探索。为此,我们提出一种人机协同方案,结合AI的海量数据处理能力与人类专家的价值判断。方法分为三阶段:第一阶段,利用AI加速文献信息搜集,构建混合信息库;第二阶段,通过六种不同LLM组成的集合,基于合成数据提出初始候选问题,并以跨模型投票过滤;第三阶段,通过多级过滤逐步增强人类干预。为验证系统,我们在五个主要学科领域开展实验,目标是识别2025年十大科学突破与2026年十大科学问题。结果表明,尽管AI在识别既定突破方面与人类高度一致,但在预测前瞻性问题时存在显著分歧,说明人类判断在评估主观性、前瞻性的挑战中依然不可或缺。
原文摘要 · Abstract (English)
The "AI Scientist" paradigm is transforming scientific research by automating key stages of the research process, from idea generation to scholarly writing. This shift is expected to accelerate discovery and expand the scope of scientific inquiry. However, a key question remains unclear: can AI scientists identify meaningful research questions? While Large Language Models (LLMs) have been applied successfully to task-specific ideation, their potential to conduct strategic, long-term assessments of past breakthroughs and future questions remains largely unexplored. To address this gap, we explore a human-AI hybrid solution that integrates the scalable data processing capabilities of AI with the value judgment of human experts. Our methodology is structured in three phases. The first phase, AI-Accelerated Information Gathering, leverages AI's advantage in processing vast amounts of literature to generate a hybrid information base. The second phase, Candidate Question Proposing, utilizes this synthesized data to prompt an ensemble of six diverse LLMs to propose an initial candidate pool, filtered via a cross-model voting mechanism. The third phase, Hybrid Question Selection, refines this pool through a multi-stage filtering process that progressively increases human oversight. To validate this system, we conducted an experiment aiming to identify the Top 10 Scientific Breakthroughs of 2025 and the Top 10 Scientific Questions for 2026 across five major disciplines. Our analysis reveals that while AI agents demonstrate high alignment with human experts in recognizing established breakthroughs, they exhibit greater divergence in forecasting prospective questions, suggesting that human judgment remains crucial for evaluating subjective, forward-looking challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。