arXiv:2505.16134cs.CLcs.LG2025-05被引 5

研究多语言大模型的位置偏差,发现模型和语言共同影响信息权重。

Beyond Early-Token Bias: Model-Specific and Language-Specific Position Effects in Multilingual LLMs

  • 分析五种语言五类模型,发现位置偏好由模型主导但受语言影响
  • 中段信息最易被忽略,但模型仍表现自信,输出熵未上升
  • 指令强调关键信息位置反降准确率,挑战常规提示工程做法

大型语言模型普遍存在位置偏差,即对上下文信息的位置依赖性不均,但该偏差在不同语言和模型间的差异尚不明确。本研究针对五种类型迥异的语言(英语、俄语、德语、印地语、越南语)与五种模型架构展开分析,考察位置偏差如何与提示策略交互并影响输出熵。主要发现:(1) 位置偏差以模型为主导,但存在语言特异性;值得注意的是,Qwen2.5-7B-Instruct、DeepSeek 7B Chat 和 Mistral 7B 均显著偏好后置位置,挑战了普遍存在的早期令牌偏好假设。(2) 在存在无关干扰项时,显式指令模型‘关键信息标记为1’反而导致所有语言下准确率下降,质疑现有提示工程实践。(3) 当相关信息位于上下文中间时,准确率下降最显著,但输出熵未相应升高,表明模型在错误使用信息时仍保持高度自信。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit position bias systematically underweighting information based on its location in the context but how this bias varies across languages and models remains unclear. We conduct a multilingual study across five typologically diverse languages (English, Russian, German, Hindi, Vietnamese) and five model architectures, analyzing how position bias interacts with prompting strategies and affects output entropy. Our key findings are: (1) Position bias is primarily model-driven but shows language-specific nuances. Notably, Qwen2.5-7B-Instruct, DeepSeek 7B Chat and Mistral 7B consistently favor late positions challenging the common assumption of universal early-token preference. (2) Explicitly instructing the model, in the presence of irrelevant distractors, that "the most relevant context to the query is marked as 1" unexpectedly reduces accuracy across all languages, questioning standard prompt-engineering practices. (3) Accuracy consistently drops most when relevant information appears in the middle of the context, yet this is not reflected in a corresponding increase in output entropy, suggesting the model remains confident even when it fails to use mid-context cues.

大模型位置偏差多语言提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。