让AI自修正解释,避免伪造风格误导判断。
REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control
- 用判决锚定推理风格,分离事实与表达方式。
- 仅需465个样本即达顶尖效果,真实数据提升7.54%。
- 适合需要可信解释的实时谣言检测场景。
社交媒体上虚假新闻泛滥,亟需能提供准确结论与忠实解释的自动化核查系统。现有基于大语言模型的方法忽视生成解释中的欺骗性表达风格,导致解释不忠实,可能误导人类判断;同时过度依赖外部知识源,引发幻觉并增加延迟,影响可靠性与实时性。为此,我们提出一种自修正范式——基于隐式解释的推理引导事实核查(REFLEX),通过主干模型与其微调版本之间的自相矛盾真实性信号构建控制向量,自然解耦事实与风格。在真实世界数据集上的实验表明,REFLEX在LLaMA系列模型下仅需465个自精炼样本即可达到当前最佳性能;且因其可迁移性,在野外数据上提升高达7.54%。结果还显示,该方法有效缓解了忠实性幻觉,使模型得出比以往更准确的判断。
原文摘要 · Abstract (English)
The prevalence of fake news on social media demands automated fact-checking systems to provide accurate verdicts with faithful explanations. However, existing large language model (LLM)-based approaches ignore deceptive misinformation styles in LLM-generated explanations, resulting in unfaithful rationales that can mislead human judgments. They rely heavily on external knowledge sources, introducing hallucinations and even high latency that undermine reliability and responsiveness, which is crucial for real-time use. To address these challenges, we propose REason-guided Fact-checking with Latent EXplanations (REFLEX), a self-refining paradigm that explicitly controls reasoning style anchored on verdict. REFLEX utilizes self-disagreement veracity signals between the backbone model and its fine-tuned variant to construct steering vectors, naturally disentangling fact from style. Experiments on the real-world dataset show REFLEX achieves state-of-the-art performance under LLaMA-series models with only 465 self-refined samples. Moreover, owing to its transferability, REFLEX yields up to a 7.54% gain on in-the-wild data. Our results further demonstrate that our method effectively mitigates faithful hallucination, thereby guiding the model toward more accurate verdicts than previous works in explainable fact-checking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。