arXiv:2507.19470cs.CLcs.HC2025-07被引 2

首个统一评估框架,让对话预测模型能力可比。

Conversations Gone Awry, But Then? Evaluating Conversational Forecasting Models

  • 构建统一评测基准,支持不同模型直接对比
  • 引入动态修正能力新指标,衡量预测更新性
  • 适配最新语言模型进展,推动对话预测研究

我们常依赖直觉预判对话走向。赋予自动化系统类似预见能力,可助力人与人互动。近年来,相关研究聚焦于‘对话失控’(Conversations Gone Awry, CGA)任务:预测正在进行的对话是否会偏离正轨。本文重新审视该任务,提出首个统一评估框架,建立基准以实现不同架构间的直接、可靠比较。该框架使我们能基于语言建模最新进展,全面梳理当前CGA模型的研究进展。同时,框架引入一种新指标,捕捉模型随对话推进而修正预测的能力。

原文摘要 · Abstract (English)

We often rely on our intuition to anticipate the direction of a conversation. Endowing automated systems with similar foresight can enable them to assist human-human interactions. Recent work on developing models with this predictive capacity has focused on the Conversations Gone Awry (CGA) task: forecasting whether an ongoing conversation will derail. In this work, we revisit this task and introduce the first uniform evaluation framework, creating a benchmark that enables direct and reliable comparisons between different architectures. This allows us to present an up-to-date overview of the current progress in CGA models, in light of recent advancements in language modeling. Our framework also introduces a novel metric that captures a model's ability to revise its forecast as the conversation progresses.

对话预测评估框架语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。