让对话主题自动识别更灵活,用户可自定义主题粒度。
Controllable Conversational Theme Detection Track at DSTC 12
- 基于用户偏好控制主题聚类的粗细程度。
- 在DSTC12比赛中验证了方法的有效性。
- 适合需要定制化对话分析的客服与销售场景。
对话分析正因语音与自然语言处理技术的进步而快速发展。大型语言模型在该领域的应用,使可自动化的任务达到了前所未有的复杂度与规模。本文提出对话主题检测作为对话分析中的关键任务,旨在自动识别并分类对话中的主题,显著减少大规模对话分析所需的人工工作量,尤其适用于客服或销售领域。与依赖固定意图集的传统对话意图检测不同,主题是面向用户的对话核心内容摘要,具有更高的表达灵活性和个性化空间。本文将可控对话主题检测作为对话系统技术挑战赛(DSTC12)的一个公开竞赛赛道,其任务为联合聚类与主题标注对话语句,并通过提供的用户偏好数据实现对主题簇粒度的可控调节。本文概述了问题定义、数据集与评估指标(自动与人工评价)。最后分析了参赛团队的提交结果并提供洞察。相关数据与代码已开源于GitHub仓库。
原文摘要 · Abstract (English)
Conversational analytics has been on the forefront of transformation driven by the advances in Speech and Natural Language Processing techniques. Rapid adoption of Large Language Models (LLMs) in the analytics field has taken the problems that can be automated to a new level of complexity and scale. In this paper, we introduce Theme Detection as a critical task in conversational analytics, aimed at automatically identifying and categorizing topics within conversations. This process can significantly reduce the manual effort involved in analyzing expansive dialogs, particularly in domains like customer support or sales. Unlike traditional dialog intent detection, which often relies on a fixed set of intents for downstream system logic, themes are intended as a direct, user-facing summary of the conversation's core inquiry. This distinction allows for greater flexibility in theme surface forms and user-specific customizations. We pose Controllable Conversational Theme Detection problem as a public competition track at Dialog System Technology Challenge (DSTC) 12 -- it is framed as joint clustering and theme labeling of dialog utterances, with the distinctive aspect being controllability of the resulting theme clusters' granularity achieved via the provided user preference data. We give an overview of the problem, the associated dataset and the evaluation metrics, both automatic and human. Finally, we discuss the participant teams' submissions and provide insights from those. The track materials (data and code) are openly available in the GitHub repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。