首次系统梳理文本修订背后的编辑意图,揭示改写背后的动机与规律。
Making Revisions Understandable: A Survey of Edit Intentions, Methods, and Applications

- 从编辑意图视角整合修订研究,构建统一分析框架
- 归纳主流数据集与识别方法,涵盖从改写到摘要的全链条应用
- 适合关注写作辅助与文档演化分析的研究者阅读
文本修订是文档创作的核心过程,反映了作者如何迭代地改进、重组和优化内容。随着维基百科、arXiv等平台大规模修订历史数据的可获取性提升,自然语言处理研究已从单纯建模修改内容转向理解修改背后的动因,即编辑意图。据作者所知,这是首个从编辑意图角度系统综述文本修订研究的论文,提供了对数据集、分类体系、识别方法与应用的统一视图。本文回顾了修订全流程的前期工作,包括修订语料库构建、编辑意图分类体系设计及意图识别方法。进一步对代表性数据集与方法进行分类,总结了写作辅助、文档修订摘要等下游应用,并指出了关键开放研究方向。
原文摘要 · Abstract (English)
Text revision is a core process in document creation, capturing how authors iteratively refine, reorganize, and improve written content. With the increasing availability of large-scale revision histories from platforms such as Wikipedia and arXiv, NLP research has begun to move beyond modeling what changes are made to understanding why they are made, i.e., the underlying edit intentions. To our knowledge, this is the first survey that synthesizes text revision research through the lens of edit intentions, providing a unified view of datasets, taxonomies, identification methods, and applications. We review prior work across the full revision workflow, including revision corpus construction, edit intention taxonomy design, and edit intention identification. We further categorize representative datasets and methods, summarize downstream applications such as writing assistance and document edit summarization, and highlight key open research directions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。