arXiv:2510.16712cs.CLcs.AI2025-10被引 3

发现大模型在搜索增强下多轮对话中立场易变,影响可信决策。

The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models

  • 构建首个多轮立场不稳定性基准数据集,涵盖12个争议领域
  • 三款主流模型均出现严重立场漂移,最高不稳定性得分为0.511
  • 揭示模型依赖提问方式、知识来源单一导致立场反复,适合安全敏感场景研究

将大型语言模型与搜索/检索引擎集成已成常态,但这类系统存在严重可靠性缺陷。本文首次系统性研究了语言模型的“变色龙行为”:在多轮对话中面对矛盾问题时,其立场极易发生转变(尤其在搜索增强型模型中)。通过自建的变色龙基准数据集,包含17,770个精心设计的问题-答案对,覆盖1,180次多轮对话和12个争议领域,暴露了当前先进系统的根本缺陷。我们提出两个理论基础的度量指标:立场不稳定性得分(0-1)和源内容复用率(0-1),用于量化知识多样性。对Llama-4-Maverick、GPT-4o-mini和Gemini-2.5-Flash的评估显示,所有模型均表现严重变色龙行为(得分0.391–0.511),其中GPT-4o-mini最差。关键的是,跨温度变化极小(<0.004),说明该现象非采样误差所致。分析表明,源复用率与置信度(r=0.627)及立场变化(r=0.429)显著相关(p<0.05),证实知识多样性不足使模型病态地依从问题表述。这些发现凸显了在医疗、法律和金融等需保持立场一致的关键系统部署前,必须进行全面一致性评估。

原文摘要 · Abstract (English)

Integration of Large Language Models with search/retrieval engines has become ubiquitous, yet these systems harbor a critical vulnerability that undermines their reliability. We present the first systematic investigation of "chameleon behavior" in LLMs: their alarming tendency to shift stances when presented with contradictory questions in multi-turn conversations (especially in search-enabled LLMs). Through our novel Chameleon Benchmark Dataset, comprising 17,770 carefully crafted question-answer pairs across 1,180 multi-turn conversations spanning 12 controversial domains, we expose fundamental flaws in state-of-the-art systems. We introduce two theoretically grounded metrics: the Chameleon Score (0-1) that quantifies stance instability, and Source Re-use Rate (0-1) that measures knowledge diversity. Our rigorous evaluation of Llama-4-Maverick, GPT-4o-mini, and Gemini-2.5-Flash reveals consistent failures: all models exhibit severe chameleon behavior (scores 0.391-0.511), with GPT-4o-mini showing the worst performance. Crucially, small across-temperature variance (less than 0.004) suggests the effect is not a sampling artifact. Our analysis uncovers the mechanism: strong correlations between source re-use rate and confidence (r=0.627) and stance changes (r=0.429) are statistically significant (p less than 0.05), indicating that limited knowledge diversity makes models pathologically deferential to query framing. These findings highlight the need for comprehensive consistency evaluation before deploying LLMs in healthcare, legal, and financial systems where maintaining coherent positions across interactions is critical for reliable decision support.

大模型可靠性立场不一致搜索增强一致性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。