LLM和人类一样,更信任标为人工生成的内容。
Label Effects: Shared Heuristic Reliance in Trust Assessment by Humans and LLM-as-a-Judge
- 通过反事实实验发现,标签影响判断,人工标签更易获信。
- 人类与LLM均对人工标签投入更多注意力,且决策不确定性更低。
- 适合关注AI评估公正性与模型对齐风险的研究者。
大型语言模型(LLMs)被广泛用作自动化评估工具(LLM-as-a-Judge)。本研究挑战其可靠性,发现人类与LLM在信任判断中均受披露来源标签的影响:相同内容若标注为人工生成,则信任度高于标注为AI生成。眼动追踪数据显示,人类高度依赖源标签作为判断启发线索。分析LLM内部状态发现,在不同标签条件下,模型对标签区域的注意力密度显著高于内容区域,且该现象在人工标签下更强,与人类注视模式一致;同时,AI标签下的预测置信度(logits)不确定性更高。结果表明,源标签是人类与LLM共有的显著启发式线索。这引发对标签敏感型LLM-as-a-Judge评估有效性的担忧,并警示:若模型对齐人类偏好,可能将人类启发式依赖引入模型,亟需推动去偏评估与对齐策略。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used as automated evaluators (LLM-as-a-Judge). This work challenges its reliability by showing that trust judgments by LLMs are biased by disclosed source labels. Using a counterfactual design, we find that both humans and LLM judges assign higher trust to information labeled as human-authored than to the same content labeled as AI-generated. Eye-tracking data reveal that humans rely heavily on source labels as heuristic cues for judgments. We analyze LLM internal states during judgment. Across label conditions, models allocate denser attention to the label region than the content region, and this label dominance is stronger under Human labels than AI labels, consistent with the human gaze patterns. Besides, decision uncertainty measured by logits is higher under AI labels than Human labels. These results indicate that the source label is a salient heuristic cue for both humans and LLMs. It raises validity concerns for label-sensitive LLM-as-a-Judge evaluation, and we cautiously raise that aligning models with human preferences may propagate human heuristic reliance into models, motivating debiased evaluation and alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。