arXiv:2505.24319cs.CL2025-05

提出分层因果框架,提升长文本修改的准确性和连贯性。

HiCaM: A Hierarchical-Causal Modification Framework for Long-Form Text Modification

  • 构建分层摘要树与因果图,系统化识别需修改内容
  • 在多领域数据集上实现最高79.5%胜率,显著优于现有模型
  • 适合需要精准长文本编辑的研究者和工业应用

大语言模型在多个领域取得显著成果,但在处理长文本修改任务时仍存在两大问题:(1) 不恰当地修改或摘要无关内容,产生不期望的改动;(2) 忽略对维持文档连贯性至关重要的隐含相关段落的必要修改。为此,我们提出HiCaM——一种分层因果修改框架,通过分层摘要树和因果图进行操作。此外,为评估HiCaM,我们从多个基准中构建了一个跨领域数据集,为评估其有效性提供资源。在该数据集上的综合评估表明,相比强基线模型,我们的方法实现显著改进,最高达79.50%胜率。结果凸显了该方法的全面性,在多种模型和领域中均表现出一致的性能提升。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable success in various domains. However, when handling long-form text modification tasks, they still face two major problems: (1) producing undesired modifications by inappropriately altering or summarizing irrelevant content, and (2) missing necessary modifications to implicitly related passages that are crucial for maintaining document coherence. To address these issues, we propose HiCaM, a Hierarchical-Causal Modification framework that operates through a hierarchical summary tree and a causal graph. Furthermore, to evaluate HiCaM, we derive a multi-domain dataset from various benchmarks, providing a resource for assessing its effectiveness. Comprehensive evaluations on the dataset demonstrate significant improvements over strong LLMs, with our method achieving up to a 79.50\% win rate. These results highlight the comprehensiveness of our approach, showing consistent performance improvements across multiple models and domains.

文本修改因果推理长文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。