针对资源匮乏语言的查询聚焦摘要,提升摘要与用户意图的一致性
QFS-Composer: Query-focused summarization pipeline for less resourced languages
- 通过查询分解+问答生成+摘要的流水线,增强摘要对用户查询的响应性
- 在斯洛文尼亚语上实测,相比基线模型,摘要相关性和一致性显著提升
- 为低资源语言构建了可复用的自监督评估框架,适合多语言研究者参考
大语言模型在文本摘要任务中表现优异,但在训练数据受限的语言中性能显著下降。本文针对低资源语言的查询聚焦摘要(QFS)挑战,提出新框架QFS-Composer,整合查询分解、问题生成(QG)、问答(QA)和抽象摘要,以提升摘要与用户意图的语义一致性。我们在斯洛文尼亚语上测试该方法,基于斯洛文尼亚语大模型构建了高质量的问答与问题生成模型,并改进无参考评估方法。实验表明,基于QA引导的摘要流水线在一致性和相关性上优于基线大模型。本工作为低资源语言的查询聚焦摘要提供了可扩展的方法论。
原文摘要 · Abstract (English)
Large language models (LLMs) demonstrate strong performance in text summarization, yet their effectiveness drops significantly across languages with restricted training resources. This work addresses the challenge of query-focused summarization (QFS) in less-resourced languages, where labeled datasets and evaluation tools are limited. We present a novel QFS framework, QFS-Composer, that integrates query decomposition, question generation (QG), question answering (QA), and abstractive summarization to improve the factual alignment of a summary with user intent. We test our approach on the Slovenian language. To enable high-quality supervision and evaluation, we develop the Slovenian QA and QG models based on a Slovene LLM and adapt evaluation approaches for reference-free summary evaluation. Empirical evaluation shows that the QA-guided summarization pipeline yields improved consistency and relevance over baseline LLMs. Our work establishes an extensible methodology for advancing QFS in less-resourced languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。