用生成式模型预测对话脱轨,比传统方法更准。
Forecasting Conversation Derailments Through Generation
- 通过微调大模型生成多种未来对话路径
- 在多个英文数据集上准确率领先现有方法
- 适合内容审核与冲突预警场景使用
预测对话脱轨对在线内容审核、冲突调解和商业谈判等实际场景具有重要意义。尽管语言模型在识别对话中的攻击性言论方面表现良好,但在预测未来对话脱轨方面仍存在困难。与以往仅基于历史对话预测结果的方法不同,本文提出一种新方法:利用微调后的LLM,基于已有对话历史采样多个未来对话轨迹,并根据这些轨迹的共识判断对话结局。同时,我们还尝试引入反映回合级对话动态的社会语言学属性作为生成引导。实验表明,该方法在英文对话脱轨预测基准上超越现有最先进水平,并在消融实验中展现出显著的性能提升。
原文摘要 · Abstract (English)
Forecasting conversation derailment can be useful in real-world settings such as online content moderation, conflict resolution, and business negotiations. However, despite language models' success at identifying offensive speech present in conversations, they struggle to forecast future conversation derailments. In contrast to prior work that predicts conversation outcomes solely based on the past conversation history, our approach samples multiple future conversation trajectories conditioned on existing conversation history using a fine-tuned LLM. It predicts the conversation outcome based on the consensus of these trajectories. We also experimented with leveraging socio-linguistic attributes, which reflect turn-level conversation dynamics, as guidance when generating future conversations. Our method of future conversation trajectories surpasses state-of-the-art results on English conversation derailment prediction benchmarks and demonstrates significant accuracy gains in ablation studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。