大模型在标注任务中易受内部先验影响,提示修正效果有限。
On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance

- 提出定义特异性熟悉度(DSF)衡量模型内部概念与任务定义匹配度
- 零样本错误中近三分之二无法通过提示纠正,整体修复率仅34.8%
- 模型对错误任务定义仍保持高置信度,适合关注任务对齐的研究者
大型语言模型(LLMs)越来越多地用于零样本标注和模型作为裁判的任务,但其可靠性取决于模型内部先验与用户指令的交互。我们研究了三个维度的交互:(1)模型对数据和任务定义的熟悉度如何影响性能;(2)提示中额外信息纠正零样本错误的能力(“决策固着”);(3)模型对任务定义错位的敏感性。在涵盖社交媒体、游戏、新闻和论坛的多样化数据集上,使用密集型和专家混合模型进行实验,发现近三分之二的零样本错误无法纠正,整体修复率(初始错误被提示纠正的比例)仅为34.8%。高置信度错误尤其难以纠正。当给出错位的任务定义时,模型仍会遵循,且置信度水平与正确定义条件下的结果无异。关键的是,我们引入了定义特异性熟悉度(DSF),衡量模型内部概念与任务定义的对齐程度。在控制数据集级混杂因素后,DSF与模型性能呈正相关(偏相关系数 r = +0.41),而三种不同的记忆化度量(ROUGE-L、BERTScore、嵌入余弦相似度)均未表现出正相关。这些发现揭示了提示修正在标注任务中的局限性,强调任务对齐的重要性超过文本级记忆。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used for zero-shot annotation and LLM-as-a-judge tasks, yet their reliability hinges on how model-internalized priors interact with user-provided instructions. We investigate three dimensions of this interaction: (1) how an LLM's familiarity with data and task definitions affects performance, (2) the extent to which additional information in prompts can correct zero-shot errors ("decision stickiness"), and (3) model susceptibility to misaligned task definitions. Through experiments on toxicity detection across diverse datasets (spanning social media, gaming, news, and forums) using both dense and mixture-of-experts models, we find that nearly two-thirds of zero-shot errors are resistant to correction, with an overall rescue rate (fraction of initial errors corrected by prompting) of only 34.8%. High-confidence errors prove especially resistant to correction. When given misaligned definitions, LLMs follow them while maintaining confidence levels unchanged from the aligned condition. Crucially, we introduce Definition-Specific Familiarity (DSF), which measures alignment between a model's internal concept and the task definition. After controlling for dataset-level confounds, DSF shows a positive association with model performance (partial r = +0.41), while three distinct memorization metrics (ROUGE-L, BERTScore, and embedding cosine similarity) all fail to show a positive association. These findings show the limitations of prompt-based correction in annotation tasks, highlighting the importance of definition alignment over text-level memorization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。