不同问题类型下大模型对提示变化的敏感度差异显著。
Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

- 对比客观题与主观题在提示扰动下的回答一致性
- 主观题回答受提示形式影响更大,一致性更低
- 结果表明评估需区分任务类型,避免误判模型信念
大语言模型的测评常将提示响应视为其价值观或信念的体现,这一假设在涉及政治立场、社会态度等主观内容时尤为脆弱。本文考察了客观题(固定答案)与主观题(观点或价值表达)在提示鲁棒性上的差异。在四个指令微调模型家族上,分别评估了三组客观数据集(MMLU、ARC、CulturalBench)和三组主观数据集(Political Compass Test、ValueBench、World Values Survey)。对每道题目/陈述施加多种提示变化(如措辞、框架、格式),并测量模型在不同变体下的回答一致性。通过二项广义估计方程分析发现,模型、数据集、提示类别及其交互作用均具有显著影响。数据集类型效应显著,且与提示类别的交互作用较大。结果表明,提示鲁棒性取决于问题类型、提示变化方式及模型本身。
原文摘要 · Abstract (English)
Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs. This assumption is particularly fragile when responses are read as evidence of political values, social attitudes, or beliefs. We ask whether prompt robustness differs between objective questions with fixed answers and subjective questions that ask for opinions or values. We evaluate four instruction-tuned model families on three objective datasets (MMLU, ARC, and CulturalBench) and three subjective datasets (Political Compass Test, ValueBench, and World Values Survey). For each question/statement, we apply multiple types of prompt changes, such as variations in wording, framing, and format, and measure whether the model gives the same answer across variants. Using a binomial generalized estimating equation, we find significant effects of model, dataset, prompt category, and their interactions. The dataset type effect is also significant, and the interaction between dataset type and prompt category is large. These results show that prompt robustness depends on the question type, the prompt change, and the model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。