arXiv:2505.09807cs.CLcs.AI2025-05被引 1

发现大模型真话方向在不同对话格式间泛化能力差异,提出末尾加关键词改善效果。

Exploring the generalization of LLM truth directions on conversational formats

  • 通过线性探测分析模型激活空间中的真话方向。
  • 短对话末尾说谎时泛化良好,长对话早期说谎时效果差。
  • 末尾添加固定关键词可显著提升跨格式泛化能力。

近期研究指出,大语言模型在激活空间中存在一个通用真话方向,真实与虚假陈述可线性分离。已有实验证明,仅需单层隐藏状态的线性探针即可在多种主题上泛化,甚至可用于检测模型对话中的谎言。本文探究该真话方向在不同对话格式间的泛化性能。结果发现:在以谎言结尾的短对话中泛化良好;但在谎言出现在输入前部的长对话中表现不佳。为此,我们提出在每段对话末尾添加固定关键词的方案,显著提升了跨格式泛化能力。研究揭示了构建可靠大模型谎言检测器在新场景下仍面临挑战。

原文摘要 · Abstract (English)

Several recent works argue that LLMs have a universal truth direction where true and false statements are linearly separable in the activation space of the model. It has been demonstrated that linear probes trained on a single hidden state of the model already generalize across a range of topics and might even be used for lie detection in LLM conversations. In this work we explore how this truth direction generalizes between various conversational formats. We find good generalization between short conversations that end on a lie, but poor generalization to longer formats where the lie appears earlier in the input prompt. We propose a solution that significantly improves this type of generalization by adding a fixed key phrase at the end of each conversation. Our results highlight the challenges towards reliable LLM lie detectors that generalize to new settings.

大模型真相方向对话生成谎言检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。