arXiv:2609.07662cs.CLcs.AI2026-09

研究大模型如何应对用户质疑,揭示其权威管理策略的复杂性。

How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement

论文配图:How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement
图 1 · 摘自论文原文
  • 构建六类挑战类型与四层响应框架,系统分析模型应对分歧策略。
  • 85%回应认可用户意见,但65%仍坚持原观点,存在明显矛盾。
  • 健康与法律建议中更愿转移权威,说明任务性质影响模型态度。

大型语言模型在高风险场景中日益作为信息与建议来源,但其面对用户质疑时如何管理自身知识权威尚不明确。本文基于会话分析,提出六类挑战类型与四层响应分析框架:是否维持原主张、权威归属、社会性协商方式及证据支持形式。构建包含2,310个受控挑战场景与32,340条响应的新数据集,涵盖14个模型,并采用大模型作为评判者的方法进行分析。结果发现,模型在85%的回应中认可用户,但65%仍坚持原主张;33%的回应包含道歉,其中59%伴随原观点维持;在健康与法律建议任务中,权威转移比例分别达57%与49%,远高于事实类(6%)和解释类(3%)任务;仅0.8%至40%的回应放弃原主张,完全替换原答案的情况极少,总体仅占1.5%。

原文摘要 · Abstract (English)

Large language models are increasingly used as sources of advice and information, including in high-stakes settings, yet little is known about how they respond to user disagreement. We study how a model manages its epistemic authority, referring here to its claim to knowledge, competence, or the right to advise, once a user challenges its answer. Building on Conversation Analysis, we introduce a taxonomy of six challenge types and a four-layer framework for analysing each response: whether the original claim is maintained or changed, where authority is located, how the disagreement is socially managed, and what kind of evidential support is offered. We construct a new dataset of 2,310 controlled challenge scenarios and 32,340 corresponding responses from 14 models, and analyse them using our framework with an LLM-as-judge pipeline, providing a vocabulary which future evaluation and benchmark design can build on. We find that models show conflicting behaviour: they validate users in 85% of responses but maintain their original claim in 65%. They explicitly apologise in 33% of responses, yet 59% of those apologies accompany maintenance of the original claim. They transfer authority most often in advice tasks, doing so in 28% of responses and reaching 57% in health advice and 49% in legal advice, compared with 6% in fact and 3% in explanation tasks. Abandonment of the original claim ranges from 0.8% for GPT-5.2 to 40% for DeepSeek 7B, while complete replacement of the original claim is rare overall at 1.5%.

大模型人机交互权威管理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。