arXiv:2507.08870cs.LGcs.MA2025-07ACL被引 2

小模型+结构化推理,高效评估科研创意并提升录用率

GUIDE: Towards Scalable Advising for Research Ideas

  • 用压缩文献库+结构化推理框架,让小模型替代大模型
  • 在ICLR 2025测试集上高置信预测达90%以上录用率
  • 适合科研人员快速优化研究设想,提升投稿成功率

人工智能研究发展迅猛,推动了跨生物学、数学和人工智能等领域的自动化假设生成与实验设计。然而,缺乏可扩展的高质量反馈系统来优化提出的假设与实验方案。为此,我们探究了模型规模、上下文长度、置信度估计和结构化推理过程对智能指导系统的影响。研究发现,配备压缩文献库与结构化推理框架的小模型,在ICLR 2025自评排名前30%的投稿中,其录用率超过通用大模型Deepseek-R1。当仅保留高置信度预测时,该系统在ICLR 2025测试集上的录用率超过90%,展现出显著提升假设生成与实验设计质量与效率的潜力。代码已开源:https://github.com/HowardLiu0830/GUIDE-Research-Idea-Evaluation。

原文摘要 · Abstract (English)

The field of AI research is advancing at an unprecedented pace, enabling automated hypothesis generation and experimental design across diverse domains such as biology, mathematics, and artificial intelligence. Despite these advancements, there remains a significant gap in the availability of scalable advising systems capable of providing high-quality, well-reasoned feedback to refine proposed hypotheses and experimental designs. To address this challenge, we explore key factors that underlie the development of robust advising systems, including model size, context length, confidence estimation, and structured reasoning processes. Our findings reveal that a relatively small model, when equipped with a well-compressed literature database and a structured reasoning framework, can outperform powerful general-purpose language models such as Deepseek-R1 in terms of acceptance rates for self-ranked top-30% submissions to ICLR 2025. Moreover, when limited to high-confidence predictions, our system achieves an acceptance rate exceeding 90% on the ICLR 2025 test set, underscoring its potential to significantly enhance the quality and efficiency of hypothesis generation and experimental design. The code is released at https://github.com/HowardLiu0830/GUIDE-Research-Idea-Evaluation.

科研助手智能评审推理框架论文投稿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。