LLM错误常因表达流畅被人类误判,形成人机共构的认知偏差。
Plausibility as Failure: How LLMs and Humans Co-Construct Epistemic Error
- 通过多轮评估发现,模型错误从预测性失误转为语义扭曲的解释性谬误。
- 人类易将语言流畅与内容可信混淆,80%以上判断依赖表面形式。
- 适合关注AI可信度、认知偏差与人机协作的研究者阅读。
大型语言模型(LLMs)日益作为日常推理中的认知伙伴使用,但其错误仍主要通过预测指标分析,而非对其对人类判断的解释性影响进行研究。本研究考察了不同形式的认知失败如何在人机互动中浮现、被掩盖并被容忍,其中失败被视为由模型生成的合理性与人类解释性判断共同作用的相对性断裂。我们采用跨学科任务和逐步细化的评估框架,在三轮多模型评估中观察评价者在语言、认知和可信度维度上对模型输出的解读。结果表明,LLM错误从预测性问题演变为解释性问题:语言流畅性、结构连贯性和看似合理的引用掩盖了深层意义扭曲。评价者频繁混淆正确性、相关性、偏见、依据性与一致性等标准,显示人类判断将分析标准坍缩为受形式和流畅度影响的直觉启发式。随着任务复杂度上升,系统性验证负担和认知漂移现象显现,评价者更依赖表面线索,导致错误但结构良好的回答被误认为可信。这说明错误不仅是模型行为的属性,更是生成合理性与人类认知捷径共同构建的结果。因此,理解人工智能认知失败需将评估重构为一种关系性的解释过程,系统故障与人类误判的边界变得模糊。研究为LLM评估、数字素养及可信人机沟通设计提供启示。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as epistemic partners in everyday reasoning, yet their errors remain predominantly analyzed through predictive metrics rather than through their interpretive effects on human judgment. This study examines how different forms of epistemic failure emerge, are masked, and are tolerated in human AI interaction, where failure is understood as a relational breakdown shaped by model-generated plausibility and human interpretive judgment. We conducted a three round, multi LLM evaluation using interdisciplinary tasks and progressively differentiated assessment frameworks to observe how evaluators interpret model responses across linguistic, epistemic, and credibility dimensions. Our findings show that LLM errors shift from predictive to hermeneutic forms, where linguistic fluency, structural coherence, and superficially plausible citations conceal deeper distortions of meaning. Evaluators frequently conflated criteria such as correctness, relevance, bias, groundedness, and consistency, indicating that human judgment collapses analytical distinctions into intuitive heuristics shaped by form and fluency. Across rounds, we observed a systematic verification burden and cognitive drift. As tasks became denser, evaluators increasingly relied on surface cues, allowing erroneous yet well formed answers to pass as credible. These results suggest that error is not solely a property of model behavior but a co-constructed outcome of generative plausibility and human interpretive shortcuts. Understanding AI epistemic failure therefore requires reframing evaluation as a relational interpretive process, where the boundary between system failure and human miscalibration becomes porous. The study provides implications for LLM assessment, digital literacy, and the design of trustworthy human AI communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。