用多轮推理和6W框架提升目击证词矛盾检测准确率
Incongruence Identification in Eyewitness Testimony
- 基于6W框架与多跳推理,将矛盾检测转为封闭式问答
- 在MIND数据集上比传统方法提升5.63%的F1分数
- 适合司法取证、AI辅助审讯等需要验证证词一致性的场景
目击证词中的不一致检测对判断证言可靠性至关重要,但传统方法难以捕捉其内在复杂矛盾。本文提出一项新任务:在两名当事人提供的多组问答对中识别语境相关的不一致之处,并标注具体矛盾片段。为此构建了MIND(MultI-EyewitNess Deception)数据集,包含2927对上下文相关答案,涵盖显性和隐性矛盾。提出INTEND(Instruction-Tuned Incongruity Detection)框架,结合6W分析法与多跳推理,将问题转化为封闭式问答,聚焦人物、事件、时间、地点、原因等方面的矛盾。实验表明,使用该框架进行提示调优,相较微调和普通提示调优方法,F1分数提升5.63个百分点。在MLMs和LLMs上的实证结果均显示显著性能优势,验证了方法的有效性。
原文摘要 · Abstract (English)
Incongruence detection in eyewitness narratives is critical for understanding the reliability of testimonies, yet traditional approaches often fail to address the nuanced inconsistencies inherent in such accounts. In this paper, we introduce a novel task of incongruence detection in eyewitness testimonies. Given a pair of testimonies containing of multiple pairs of question and answer by two subjects, we identify contextually related incongruence between the two subjects. We also mark the span of incongruences in the utterances. To achieve this, we developed MIND(MultI-EyewitNess Deception) - a comprehensive dataset consisting of 2927 pairs of contextually related answers designed to capture both explicit and implicit contradictions. INstruction - TunEd iNcongruity Detection framework based on 6W and multi-hop reasoning approach, aka. INTEND. Drawing from investigative techniques, INTEND address the task as a close-style problem, contradicting on the who, what, when, where and why aspect of the content. Our findings shows that prompt tuning, especially when utilizing our framework, enhances the detection of incongruences by a margin of +5.63 percent. We compare our approach with multiple fine-tuning and prompt tuning techniques on MLMs and LLMs. Emperical results demonstrate convincing performance improvement in F1-score over fine-tuned and regular prompt-tuning techniques, highlighting the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。