人类更易受内容来源标签影响,而大模型评估逻辑谬误更稳定。
Label Over Logic? How Source Cues Bias Human Fallacy Judgments More Than LLMs

- 用逻辑谬误作为测试场景,比较人类与大模型对不同来源内容的判断差异。
- 505名参与者中,人类对标注为人类或人机协作的内容评分更高,明显受标签影响。
- 大模型评估结果受来源标签影响小,适合用于客观推理评价任务。
随着人工智能生成和辅助内容大量涌入网络空间,附带的内容来源标签会扭曲人类的判断,对内容审核、评估和决策产生下游影响。大模型是否同样易受此类影响,或具备更独立于来源的评估能力,仍是开放问题,直接关系到人机协作的可行性。本研究以逻辑谬误为控制场景,剥离领域知识干扰,仅考察来源标签对推理质量判断的影响。我们开展一项在线实验(N=505),参与者被分配至五种来源条件(人类、AI、人机协作、机人协作、无披露),并评估包含逻辑谬误的评论,同时对比GPT-5.2、Gemini 2.5 Flash、Claude Sonnet 4.5三款大模型在相同条件下的评估表现。结果显示,人类评估者显著受标签影响,对标注为人类或人机协作的内容赋予更高信任度和评分。而大模型的评估结果在不同来源条件下保持相对稳定,尽管模型间表现有差异。人类与大模型在各条件下的信心水平均很高,且与是否存在谬误无关。研究发现,来源标签偏差主要存在于人类判断中,凸显了大模型在日益智能化的人机协同环境中提供客观评估的潜力。
原文摘要 · Abstract (English)
As AI-generated and AI-assisted content floods online spaces, source labels attached to such content can distort human reasoning judgments, with downstream consequences for moderation, evaluation, and decision-making. Whether LLMs share this vulnerability, or offer more source-agnostic evaluation, remains an open question with direct implications for human-AI collaboration. We examine this issue using logical fallacies as a controlled setting to isolate source-label effects on reasoning quality, independent of domain knowledge. We conduct an online study (N=505) where participants are assigned to a source condition (human, AI, human with AI assistance, AI with human assistance, or no disclosure) and evaluate comments containing logical fallacies, comparing their judgments with those of LLMs (GPT-5.2, Gemini 2.5 Flash, Claude Sonnet 4.5), who were evaluated across the same source conditions. Human evaluators were significantly more susceptible to fallacies labeled as written by human or human with AI assistance and assigned higher trust and evaluation ratings in these conditions. LLM evaluations remained comparatively stable across source labels, though performance varied across models. Confidence levels were similarly high across conditions for both humans and LLMs, regardless of fallacy presence. Our findings indicate that source-label bias in reasoning evaluation is primarily a human vulnerability and highlight the potential of human-LLM collaboration in increasingly AI-mediated environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。