arXiv:2410.13667cs.CL2024-10EMNLP被引 10

首个中文辩论语料库,支持立场检测与对话摘要任务。

ORCHID: A Chinese Debate Corpus for Target-Independent Stance Detection and Argumentative Dialogue Summarization

  • 构建真实中文辩论数据集,覆盖476个主题
  • 含14,133条标注对话,2,436份立场摘要
  • 适合研究中文论辩理解与智能摘要的学者

近年来,对话代理受到广泛关注,大语言模型的发展进一步推动了这一趋势。立场检测与对话摘要作为论辩对话中的核心任务,却因公开数据集不足而受限,尤其在非英语语言中更为明显。为弥补中文语料的空白,我们提出了ORCHID(Oral Chinese Debate),这是首个用于评估目标无关立场检测与辩论摘要的中文数据集。该数据集包含1,218场真实中文辩论,涉及476个独特话题,包含2,436份立场相关的摘要和14,133条完全标注的发言。除了提供多用途测试平台外,我们还对该数据集进行了实证研究,并提出一个整合任务。结果表明该数据集具有挑战性,且将立场检测融入摘要生成具有潜在价值。

原文摘要 · Abstract (English)

Dialogue agents have been receiving increasing attention for years, and this trend has been further boosted by the recent progress of large language models (LLMs). Stance detection and dialogue summarization are two core tasks of dialogue agents in application scenarios that involve argumentative dialogues. However, research on these tasks is limited by the insufficiency of public datasets, especially for non-English languages. To address this language resource gap in Chinese, we present ORCHID (Oral Chinese Debate), the first Chinese dataset for benchmarking target-independent stance detection and debate summarization. Our dataset consists of 1,218 real-world debates that were conducted in Chinese on 476 unique topics, containing 2,436 stance-specific summaries and 14,133 fully annotated utterances. Besides providing a versatile testbed for future research, we also conduct an empirical study on the dataset and propose an integrated task. The results show the challenging nature of the dataset and suggest a potential of incorporating stance detection in summarization for argumentative dialogue.

中文语料立场检测对话摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。