不同标签会显著影响模型对上下文的信任程度,标签选择需谨慎。
Discourse-Role Labels as Presentation-Time Variables for Context Use in Language Models
- 用同一错误答案搭配不同标签测试模型采纳率
- 标签差异导致采纳率波动达56-84个百分点
- 指令类标签易诱使模型采纳错误信息
上下文增强的语言模型常在输入内容前添加如Reference:、Evidence:、Instruction:等标签,但这些标签对模型行为的影响尚未被充分研究。本文设计了一项针对500万条MMLU-Pro题目的配对固定内容探测:相同误导性陈述在不同话语角色标签下呈现,测量模型是否采纳错误选项。在GPT-5.5、DeepSeek V4 Pro、Llama-3-8B-Instruct和Qwen2.5-7B-Instruct中,误导采纳率随标签变化达56-84个百分点。指令类或来源类标签(如Instruction:、Reference:)显著提升采纳率,而Example:则持续抑制采纳。配对检验、置信区间、最终指令消融及Qwen的步骤级概率探针均支持标签条件下的候选偏好机制。边界探测显示:算术任务降低采纳率,外部上下文为段落形式时标签差异较小,短答案评估排除了选项字母复制干扰,嵌套标签冲突表明说明性框架可限定采纳范围。200例人工审计验证短答案对比结果在保守判断下仍稳定。结论虽有限制,但具实用性:上下文使用与读者端RAG基准应报告并控制包装标签,因展示方式会改变对上下文依赖性的测量结果。
原文摘要 · Abstract (English)
Context-augmented language model systems often wrap supplied content with labels such as Reference:, Evidence:, Instruction:, Note:, or Example:, but the effect of these labels on reader-model behavior remains underexplored. We introduce a paired fixed-content probe over 500 MMLU-Pro items: each item receives the same misleading answer-bearing assertion under different discourse-role labels, and adoption is measured by whether the model outputs the injected wrong option. Across GPT-5.5, DeepSeek V4 Pro, Llama-3-8B-Instruct, and Qwen2.5-7B-Instruct, Misleading Adoption Rate shifts by 56-84 percentage points. Binding or source-like labels such as Instruction: and Reference: produce high adoption, whereas Example: consistently suppresses it. Paired tests, bootstrap intervals, final-instruction ablations, and Qwen final-step log-probability probes support a label-conditioned candidate preference. Boundary probes show where the effect weakens or persists: arithmetic tasks reduce adoption, passage-shaped external context preserves smaller label gaps, short-answer evaluation rules out option-letter copying, and nested-label conflicts suggest that illustrative framing can delimit adoption scope. A 200-case single-author manual audit confirms that the short-answer contrasts are stable under conservative adjudication. The resulting claim is bounded but practical: context-utilization and reader-side RAG benchmarks should report and control wrapper labels, because presentation choices can change measured reliance on supplied context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。