让对话总结看清观点与理由,并量化支持强度
ARQUSUMM: Argument-aware Quantitative Summarization of Online Conversations
- 用大模型识别句子中的观点与理由关系
- 通过聚类量化不同观点的支持程度
- 适合需要客观分析网络争论的用户
在线讨论在公共平台(如 Reddit)日益普遍。面对越来越多的争议话题,仅总结表面信息已不足,还需揭示观点背后的论据与合理性。早期文本摘要研究侧重提取文档中的关键信息,忽略了在线对话的论辩特性;近期对话摘要虽考虑句间论辩关系,但未能深入解析句内结构。本文提出「论点感知的数量化摘要」新任务,旨在揭示对话中论点-理由结构并量化论证强度。为此,我们提出 ARQUSUMM 框架:利用基于论证理论的大模型少样本学习,识别句子内的命题及其论点-理由关系;再通过考虑论点结构的聚类算法聚合观点并量化支持度。实验表明,ARQUSUMM 在对话与数量化摘要任务上均优于现有模型,生成的摘要在论点结构呈现、文本质量和量化准确性方面表现更优。
原文摘要 · Abstract (English)
Online conversations have become more prevalent on public discussion platforms (e.g. Reddit). With growing controversial topics, it is desirable to summarize not only diverse arguments, but also their rationale and justification. Early studies on text summarization focus on capturing general salient information in source documents, overlooking the argumentative nature of online conversations. Recent research on conversation summarization although considers the argumentative relationship among sentences, fail to explicate deeper argument structure within sentences for summarization. In this paper, we propose a novel task of argument-aware quantitative summarization to reveal the claim-reason structure of arguments in conversations, with quantities measuring argument strength. We further propose ARQUSUMM, a novel framework to address the task. To reveal the underlying argument structure within sentences, ARQUSUMM leverages LLM few-shot learning grounded in the argumentation theory to identify propositions within sentences and their claim-reason relationships. For quantitative summarization, ARQUSUMM employs argument structure-aware clustering algorithms to aggregate arguments and quantify their support. Experiments show that ARQUSUMM outperforms existing conversation and quantitative summarization models and generate summaries representing argument structures that are more helpful to users, of high textual quality and quantification accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。