评测大模型生成论文引言的能力,发现LLaMA-4 Maverick表现最佳。
Towards AI-Assisted Research Writing: Benchmarking LLMs for AI/ML Introduction Generation
- 构建新任务SciIG,基于标题、摘要和参考文献生成引言
- LLaMA-4 Maverick在语义相似性和忠实度上领先,三轮提示效果最好
- 适合想用AI辅助科研写作的研究者参考
随着研究人员越来越多地使用大模型作为写作助手,生成高质量的学术论文引言仍具挑战性且至关重要。本文提出科学引言生成(SciIG)任务,评估大模型根据论文标题、摘要和相关工作生成连贯引言的能力。基于NAACL 2025和ICLR 2025论文构建新数据集,评估了五种前沿模型,包括开源模型DeepSeek-v3、Gemma-3-12B、LLaMA 4-Maverick、MistralAI Small 3.1以及闭源GPT-4o系统。从词汇重叠、语义相似性、内容覆盖、忠实度、一致性、引用正确性和叙事质量等多个维度进行综合评估。采用自动化指标与大模型评判相结合的框架。结果表明,LLaMA-4 Maverick在多数指标上表现最优,尤其在语义相似性和忠实度方面突出。此外,三轮提示显著优于少轮提示。研究为开发高效科研写作助手提供实践指导,并设定了对大模型辅助学术写作的合理预期。为促进可复现性和未来研究,所有代码与数据集均已公开。
原文摘要 · Abstract (English)
As researchers increasingly adopt LLMs as writing assistants, generating high-quality research paper introductions remains both challenging and essential. We introduce Scientific Introduction Generation (SciIG), a task that evaluates LLMs' ability to produce coherent introductions from titles, abstracts, and related works. Curating new datasets from NAACL 2025 and ICLR 2025 papers, we assess five state-of-the-art models, including both open-source (DeepSeek-v3, Gemma-3-12B, LLaMA 4-Maverick, MistralAI Small 3.1) and closed-source GPT-4o systems, across multiple dimensions: lexical overlap, semantic similarity, content coverage, faithfulness, consistency, citation correctness, and narrative quality. Our comprehensive framework combines automated metrics with LLM-as-a-judge evaluations. Results demonstrate LLaMA-4 Maverick's superior performance on most metrics, particularly in semantic similarity and faithfulness. Moreover, three-shot prompting consistently outperforms fewer-shot approaches. These findings provide practical insights into developing effective research writing assistants and set realistic expectations for LLM-assisted academic writing. To foster re- producibility and future research, we publicly release all code and datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。