构建细粒度领域迁移基准,揭示摘要模型在不同文本类型下的性能变化。
DomainSum: A Hierarchical Benchmark for Fine-Grained Domain Shift in Abstractive Text Summarization
- 按文体、风格、主题分层设计领域迁移评估体系
- 实证发现领域迁移具有层级结构特征
- 适合研究模型跨域泛化能力的学者参考
当前抽象式摘要研究多集中于单一领域应用,常忽视文档间领域差异对模型性能与泛化能力的影响。为解决此问题,我们提出DomainSum,一个用于捕捉抽象式摘要中细粒度领域迁移的分层基准。我们将领域迁移划分为三个层次:体裁、风格和主题,并通过全面的基准分析证明其具有层级结构。此外,我们在同域与跨域设置下评估了常用预训练语言模型(PLMs)与大语言模型(LLMs)的领域泛化能力。
原文摘要 · Abstract (English)
Most research on abstractive summarization focuses on single-domain applications, often neglecting how domain shifts between documents affect performance and the generalization ability of summarization models. To address this issue, we introduce DomainSum, a hierarchical benchmark designed to capture fine-grained domain shifts in abstractive summarization. We categorize these shifts into three levels: genre, style, and topic, and demonstrate through comprehensive benchmark analysis that they follow a hierarchical structure. Furthermore, we evaluate the domain generalization capabilities of commonly used pre-trained language models (PLMs) and large language models (LLMs) in in-domain and cross-domain settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。