arXiv:2412.08196cs.CLcs.CV2024-12被引 1

针对公文摘要难题,提出可适配领域的预训练方法。

DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization

  • 用OCR文本预训练+问答对微调,提升公文摘要能力。
  • 在真实业务场景下,摘要准确率显著优于基线模型。
  • 适合政府、企业等需处理大量公文的机构使用。

抽取式摘要在压缩和重述大段文本方面已取得显著进展。然而,行政文档因其领域特定术语、OCR生成错误以及标注数据稀缺,在摘要任务中面临独特挑战。现有模型难以适应此类文档的复杂结构与专业内容。为此,我们提出DocSum,一种专为行政文档设计的领域自适应摘要框架。通过在OCR转录文本上进行预训练,并结合创新的问答对微调策略,提升了摘要的准确性与相关性。该方法有效应对了行政内容的固有复杂性,确保输出符合实际业务需求。为评估其性能,我们定义了一个新的下游任务——文档抽象摘要,反映企业和组织的实际应用场景。大量实验表明,DocSum能生成高质量摘要,具备提升决策效率与运营流程的潜力。

原文摘要 · Abstract (English)

Abstractive summarization has made significant strides in condensing and rephrasing large volumes of text into coherent summaries. However, summarizing administrative documents presents unique challenges due to domain-specific terminology, OCR-generated errors, and the scarcity of annotated datasets for model fine-tuning. Existing models often struggle to adapt to the intricate structure and specialized content of such documents. To address these limitations, we introduce DocSum, a domain-adaptive abstractive summarization framework tailored for administrative documents. Leveraging pre-training on OCR-transcribed text and fine-tuning with an innovative integration of question-answer pairs, DocSum enhances summary accuracy and relevance. This approach tackles the complexities inherent in administrative content, ensuring outputs that align with real-world business needs. To evaluate its capabilities, we define a novel downstream task setting-Document Abstractive Summarization-which reflects the practical requirements of business and organizational settings. Comprehensive experiments demonstrate DocSum's effectiveness in producing high-quality summaries, showcasing its potential to improve decision-making and operational workflows across the public and private sectors.

文档摘要领域适配预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。