测试大模型在长期对话中观点变化的稳定性与合理性
BeliefShift: Benchmarking Temporal Belief Consistency and Opinion Drift in LLM Agents
- 设计纵向对话基准,追踪多轮交互中的信念演化
- 7个模型表现参差,过度个性化易导致观点漂移
- 提出4项新指标,评估信念更新的准确性和证据敏感性
大型语言模型正越来越多地作为长期对话代理使用,但现有评测均将用户信息视为静态事实,这一模型存在根本缺陷。人们会改变想法,长期互动中观点漂移、过度对齐和确认偏见等问题日益显著。本文提出BeliefShift,一个专门用于评估多会话交互中信念动态的纵向基准,包含三个方向:时间信念一致性、矛盾检测与证据驱动修正。数据集包含2,400条人工标注的跨会话对话轨迹,涵盖健康、政治、个人价值观及产品偏好等领域。我们评估了GPT-4o、Claude 3.5 Sonnet、Gemini 1.5 Pro、LLaMA-3和Mistral-Large等7个模型,在零样本与检索增强生成(RAG)设置下的表现。结果揭示明确权衡:高度个性化的模型抗漂移能力差,而事实导向模型则难以识别合理信念更新。此外,我们引入四项新评估指标:信念修订准确率(BRA)、漂移连贯性得分(DCS)、矛盾解决率(CRR)和证据敏感性指数(ESI)。
原文摘要 · Abstract (English)
LLMs are increasingly used as long-running conversational agents, yet every major benchmark evaluating their memory treats user information as static facts to be stored and retrieved. That's the wrong model. People change their minds, and over extended interactions, phenomena like opinion drift, over-alignment, and confirmation bias start to matter a lot. BeliefShift introduces a longitudinal benchmark designed specifically to evaluate belief dynamics in multi-session LLM interactions. It covers three tracks: Temporal Belief Consistency, Contradiction Detection, and Evidence-Driven Revision. The dataset includes 2,400 human-annotated multi-session interaction trajectories spanning health, politics, personal values, and product preferences. We evaluate seven models including GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, LLaMA-3, and Mistral-Large under zero-shot and retrieval-augmented generation (RAG) settings. Results reveal a clear trade-off: models that personalize aggressively resist drift poorly, while factually grounded models miss legitimate belief updates. We further introduce four novel evaluation metrics: Belief Revision Accuracy (BRA), Drift Coherence Score (DCS), Contradiction Resolution Rate (CRR), and Evidence Sensitivity Index (ESI).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。