Dolphin用闭环机制自动迭代科研,从想法到验证全流程自动化。
Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback
- 基于实验反馈与文献排序生成新研究思路。
- 通过异常追踪优化代码结构,实现自动执行与分析。
- 在3D点云分类等任务上达到顶尖水平,适合自动化研究场景。
科学研究所处范式正因人工智能发展而深刻变革。近期工作表明,多种AI辅助方法可通过提升数据分析效率、加速计算和激发新想法显著提高研究效率。为更接近自动科研的终极目标,本文提出Dolphin——一个闭环驱动的大型语言模型框架,以提升科研自动化水平。Dolphin首先根据前序实验反馈及按主题与任务属性排序的相关论文生成新想法;随后利用经异常追踪引导的本地代码结构优化并调试代码模板,实现想法的自动执行;最后对每轮结果进行自动分析,并将结果反馈至下一轮想法生成。在多个主题的基准数据集及MLE-bench子集上的实验表明,Dolphin可在循环中持续提升输入主题的表现。我们发现,Dolphin能自动生成在某些任务(如3D点分类)上媲美当前最优方法的新方案。
原文摘要 · Abstract (English)
The scientific research paradigm is undergoing a profound transformation owing to the development of Artificial Intelligence (AI). Recent works demonstrate that various AI-assisted research methods can largely improve research efficiency by improving data analysis, accelerating computation, and fostering novel idea generation. To further move towards the ultimate goal (i.e., automatic scientific research), in this paper, we introduce Dolphin, a closed-loop LLM-driven framework to enhance the automation level of scientific research. Dolphin first generates novel ideas based on feedback from previous experiments and relevant papers ranked by the topic and task attributes. Then, the generated ideas can be implemented using a code template refined and debugged with the designed exception-traceback-guided local code structure. Finally, Dolphin automatically analyzes the results of each idea and feeds the results back to the next round of idea generation. Experiments are conducted on the benchmark datasets of different topics and a subset of MLE-bench. Results show that Dolphin can continuously improve the performance of the input topic in a loop. We highlight that Dolphin can automatically propose methods that are comparable to the state-of-the-art in some tasks such as 3D point classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。