首个端到端学术写作数据集,记录真实写稿全过程。
ScholaWrite: A Dataset of End-to-End Scholarly Writing Process
- 用浏览器插件无感采集Overleaf上的输入记录
- 涵盖5篇计算机科学预印本近6.2万次编辑
- 揭示大模型辅助与人类写作认知的差距
写作是高度认知负荷的活动,依赖工作记忆并频繁切换目标。为构建真正契合作者认知的写作助手,必须捕捉并解析从构思到成文的完整思维过程。我们提出ScholaWrite,首个端到端学术写作数据集,追踪从初稿到终稿的数月写作旅程。主要贡献包括:(1) 一款无感记录Overleaf键入行为的Chrome扩展;(2) 包含五篇计算机科学预印本的全新语料库,基于 exttt{LaTeX}的近6.2万次文本修改,附有精细标注的认知意图;(3) 对学术写作微观动态的分析,揭示当前大语言模型在提供有意义协助方面与人类写作过程的差距。ScholaWrite强调端到端写作数据对发展支持而非替代科研人员认知工作的未来写作助手的价值。
原文摘要 · Abstract (English)
Writing is a cognitively demanding activity that requires constant decision-making, heavy reliance on working memory, and frequent shifts between tasks of different goals. To build writing assistants that truly align with writers' cognition, we must capture and decode the complete thought process behind how writers transform ideas into final texts. We present ScholaWrite, the first dataset of end-to-end scholarly writing, tracing the multi-month journey from initial drafts to final manuscripts. We contribute three key advances: (1) a Chrome extension that unobtrusively records keystrokes on Overleaf, enabling the collection of realistic, in-situ writing data; (2) a novel corpus of full scholarly manuscripts, enriched with fine-grained annotations of cognitive writing intentions. The dataset includes \LaTeX-based edits from five computer science preprints, capturing nearly 62K text changes over four months; and (3) analyses and insights into the micro-dynamics of scholarly writing, highlighting gaps between human writing processes and the current capabilities of large language models (LLMs) in providing meaningful assistance. ScholaWrite underscores the value of capturing end-to-end writing data to develop future writing assistants that support, not replace, the cognitive work of scientists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。