提出可延迟预警的对话脱轨预测机制,降低误报率。
Wait! There's a Way Out: A Decision Mechanism for Forecasting Conversational Derailment

- 用前瞻模拟判断紧张时刻是否有恢复可能,再决定是否预警。
- 在不降低准确率前提下,显著减少误报次数。
- 适合需要低误报的实时对话监控场景。
对话脱轨预测旨在随对话进行判断其是否会演变为人身攻击。现有模型在线决策时仅依据当前话语的脱轨概率,假设未来轨迹不可变,忽略后续修复可能,导致误报过高。本文提出将预警决策与脱轨概率估计解耦的方法,受人类基准启发——人类通过推迟决策,在预判紧张将缓解时避免过早预警。我们设计了一种前向模拟的延迟能力,评估紧张时刻是否存在合理恢复路径。将其融入先进预测模型后,显著降低误报率,同时保持预测精度。本工作强调将决策机制作为预测系统的核心组件,具有广泛适用性。
原文摘要 · Abstract (English)
Forecasting conversational derailment is the task of predicting, as the conversation unfolds, whether it will eventually derail into personal attacks. Since forecasting models operate in an online fashion, they must decide whether to "trigger" an alert after each utterance--for example, to notify participants or a moderator that the conversation is at risk of derailing. Existing approaches make this decision solely based on the estimated likelihood of derailment given the preceding utterances, implicitly assuming that the conversation's future trajectory is fixed. As a result, they ignore the possibility of future recovery and incur an unnecessarily high rate of false positives. In this work we propose a method for decoupling the decision to trigger from derailment likelihood estimation. Our approach is inspired by the first human baseline on this task, which shows that humans achieve dramatically lower false positive rates by selectively deferring their decision to trigger when they anticipate that tension is likely to subside. We operationalize this insight with a deferral mechanism that uses forward-looking simulations to assess whether a tense moment admits plausible paths to recovery. Incorporating this mechanism into a state-of-the-art forecasting model substantially reduces false positives without sacrificing forecasting accuracy. More broadly, this work highlights the value of treating decision-making as a first-class component of forecasting systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。