研究发现模型对熟悉信息更稳定,陌生信息易引发信念动摇。
Epistemic Familiarity is Associated With Belief Stability in Large Language Models
- 通过语义扰动测试模型信念稳定性,对比熟悉虚构与陌生合成信息。
- 陌生合成信息导致信念撤回率超50%,远高于熟悉虚构信息。
- 适合关注大模型可靠性、鲁棒性评估的研究者参考。
大型语言模型(LLMs)广泛用作信息来源,但微小的语义假设变化可能使其信念失稳。本文提出P-StaT(真理扰动稳定性)框架,在表示和行为两个层面评估模型在匹配语义扰动下的信念稳定性。我们在21个LLMs和三个领域中,比较了熟悉虚构陈述与合成生成的陌生合成陈述的扰动影响。结果表明,陌生合成扰动通常比熟悉虚构扰动引发更大认知不稳定性,行为层面的信念撤回率常超过0.5(50%)。探索性聚类分析揭示被撤回陈述的共性主题,包括模糊性、技术术语和晦涩概念。这些结果表明,认知熟悉度与语义重构下的稳定性系统相关,提示以稳定性为基础的分析可补充基于准确性的基准,用于评估LLM的鲁棒性和可靠性。代码与数据见https://github.com/samanthadies/P-StaT。
原文摘要 · Abstract (English)
Large language models (LLMs) are widely used as information sources, yet small changes in semantic assumptions can destabilize their beliefs. We introduce P-StaT (Perturbation Stability of Truth), a framework for evaluating belief stability under matched semantic perturbations in both representational and behavioral settings. Across 21 LLMs and three domains, we compare perturbations involving familiar Fictional statements against synthetically generated and unfamiliar Synthetic statements. Unfamiliar Synthetic perturbations generally induce greater epistemic instability than familiar Fictional perturbations, with behavioral belief retraction rates frequently exceeding 0.5 (50%). Finally, exploratory clustering analyses reveal recurring themes among retracted statements, including ambiguity, technical terminology, and obscure concepts. These results show that epistemic familiarity is systematically associated with stability under semantic reframing, suggesting that stability-based analyses can complement accuracy-based benchmarks when evaluating LLM robustness and reliability. Code and data are available at https://github.com/samanthadies/P-StaT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。