arXiv:2509.26302cs.CLcs.AI2025-09EMNLP被引 2

用问答评估生成摘要,零样本实现任务导向对话摘要

QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization

  • 通过大模型生成多份摘要和任务问答对,以问答准确率筛选优质摘要
  • 在多个数据集上零样本表现媲美全监督最先进方法
  • 适合医疗等需精准任务信息的对话摘要场景

对话摘要旨在将对话核心内容浓缩为简洁文本,对降低对话密集型应用中的复杂性与噪声至关重要。现有方法通常依赖人工标注摘要进行训练,成本高且生成结果缺乏任务聚焦,限制了其在医疗等下游任务中的有效性。本文提出 extit{QUARTZ} 框架,采用零样本方式利用大模型池生成多个摘要及任务相关问答对。通过让大模型回答任务问题,根据答案质量选择最优答案,并据此识别最具信息量的摘要。最后对最佳大模型进行微调。在多个数据集上的验证表明, extit{QUARTZ} 在多种零样本设置下表现优异,可媲美完全监督的最先进方法。

原文摘要 · Abstract (English)

Dialogue summarization aims to distill the core meaning of a conversation into a concise text. This is crucial for reducing the complexity and noise inherent in dialogue-heavy applications. While recent approaches typically train language models to mimic human-written summaries, such supervision is costly and often results in outputs that lack task-specific focus limiting their effectiveness in downstream applications, such as medical tasks. In this paper, we propose \app, a framework for task-oriented utility-based dialogue summarization. \app starts by generating multiple summaries and task-oriented question-answer pairs from a dialogue in a zero-shot manner using a pool of large language models (LLMs). The quality of the generated summaries is evaluated by having LLMs answer task-related questions before \textit{(i)} selecting the best candidate answers and \textit{(ii)} identifying the most informative summary based on these answers. Finally, we fine-tune the best LLM on the selected summaries. When validated on multiple datasets, \app demonstrates its effectiveness by achieving competitive results in various zero-shot settings, rivaling fully-supervised State-of-the-Art (SotA) methods.

对话摘要零样本大模型任务导向

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。