arXiv:2505.11336cs.CL2025-05ACL被引 4

XtraGPT让AI懂学术写作上下文,能精准改论文,提升研究表达质量。

XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration

  • 用准则引导意图对齐+上下文感知建模,实现深度学术润色
  • 在7000篇顶会论文上训练,14万条段落级修改数据验证效果
  • 开源1.5B到14B参数模型,性能媲美闭源产品,适合科研人员精修稿件

尽管大语言模型在学术工作流中日益普及,其在高质量科学写作支持方面仍显不足。现有系统多面向通用科学文本生成,难以满足研究沟通中深层一致性需求,如跨章节概念连贯性。此外,学术写作本质是迭代修订过程,但直接提示范式难以有效支持。为此,我们提出一种以人为中心的AI协作框架,聚焦准则引导的意图对齐与上下文感知建模。为验证框架有效性,我们构建了一个包含7000篇顶级会议论文的数据集,标注了14万条指令-响应对,反映真实、段落级的学术修改场景。我们基于此开发了XtraGPT,首个专为上下文感知学术论文修订而微调的开源大模型系列(1.5B至14B参数)。大量实验表明,XtraGPT显著优于同规模基线模型,且性能接近商业闭源方案。自动化偏好评估与人工评测均证实其在提升科学稿件质量上的有效性。代码与模型已公开于https://github.com/Xtra-Computing/XtraGPT 和 https://huggingface.co/collections/Xtra-Computing/xtragpt。

原文摘要 · Abstract (English)

Despite the growing adoption of large language models (LLMs) in academic workflows, their capabilities remain limited in supporting high-quality scientific writing. Most existing systems are designed for general-purpose scientific text generation and fail to meet the sophisticated demands of research communication beyond surface-level polishing, for example, maintaining conceptual coherence across sections. Furthermore, academic writing is inherently iterative and revision-driven, a process that is not well supported by direct prompting-based paradigms. To address these scenarios, we propose a human-AI collaboration framework for academic paper revision, centered on criteria-guided intent alignment and context-aware modeling. To validate the framework, we curate a dataset of 7,000 research papers from top-tier venues, annotated with 140,000 instruction--response pairs that reflect realistic, section-level scientific revisions. We instantiate the framework in XtraGPT, the first suite of open-source LLMs (1.5B to 14B parameters) specifically fine-tuned for context-aware academic paper revision. Extensive experiments show that XtraGPT significantly outperforms same-scale baselines and rivals the quality of proprietary counterparts. Both automated preference assessments and human evaluations confirm the effectiveness of XtraGPT in improving scientific drafts. Our code and models are available at https://github.com/Xtra-Computing/XtraGPT and https://huggingface.co/collections/Xtra-Computing/xtragpt.

论文修改AI协作大模型学术写作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。