用大模型自动分析大模型研究,效率提升93%以上
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
- 用大模型自动提取论文中的实验数据并结构化
- 发现上下文示例对编码和多模态任务帮助大,数学推理提升有限
- 可动态更新,适合持续追踪前沿大模型表现
大模型研究激增,人工梳理成果耗时费力。本文提出一种半自动化文献分析方法,利用大模型自动识别arXiv相关论文,提取实验结果及属性,构建结构化数据集LLMEvalDB。通过该方法,对前沿大模型的文献分析相比手动方式节省超过93%工作量。验证显示,LLMEvalDB能复现近期关于思维链(CoT)推理的手动分析关键结论,并发现新洞察:上下文示例在编程与多模态任务中显著提升性能,但在数学推理任务中相比零样本CoT提升有限。该数据集可随新论文持续更新,支持对目标模型的长期追踪分析。结合实证分析,为理解大模型行为提供新视角,并推动持续的文献研究。
原文摘要 · Abstract (English)
The surge of LLM studies makes synthesizing their findings challenging. Analysis of experimental results from literature can uncover important trends across studies, but the time-consuming nature of manual data extraction limits its use. Our study presents a semi-automated approach for literature analysis that accelerates data extraction using LLMs. It automatically identifies relevant arXiv papers, extracts experimental results and related attributes, and organizes them into a structured dataset, LLMEvalDB. We then conduct an automated literature analysis of frontier LLMs, reducing the effort of paper surveying and data extraction by more than 93% compared to manual approaches. We validate LLMEvalDB by showing that it reproduces key findings from a recent manual analysis of Chain-of-Thought (CoT) reasoning and also uncovers new insights that go beyond it, showing, for example, that in-context examples benefit coding & multimodal tasks but offer limited gains in math reasoning tasks compared to zero-shot CoT. Our automatically updatable dataset enables continuous tracking of target models by extracting evaluation studies as new data becomes available. Through LLMEvalDB and empirical analysis, we provide insights into LLMs while facilitating ongoing literature analyses of their behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。