arXiv:2510.26193cs.CL2025-10EMNLP被引 1

提出RCScore框架,量化指令风格对大模型输出一致性的影响。

RCScore: Quantifying Response Consistency in Large Language Models

  • 通过变换指令风格系统评估模型响应差异。
  • 指令风格变化可导致准确率波动最高达16.7个百分点。
  • 发现模型规模越大、确定性解码越稳定,一致性与准确性强相关。

当前大语言模型评估常依赖单一指令模板,忽略了模型对指令风格的敏感性,而这在实际部署中至关重要。我们提出RCScore,一个量化指令表述如何影响模型响应的多维框架。通过将基准任务系统性地转化为多种指令风格,RCScore揭示了传统指标未能捕捉到的表现差异。在四个推理基准上对十种LLM的实验表明,指令风格变化可使准确率波动高达16.7个百分点。我们引入跨响应相似性(CRS),利用RCScore指标衡量模型在不同风格下的自一致性,并发现其与任务准确率高度相关,表明一致性可作为模型可靠性的重要代理指标。额外发现显示,确定性解码产生更稳定的风格输出,且模型规模与跨风格一致性呈正相关。RCScore为评估指令鲁棒性提供了原则性方法。

原文摘要 · Abstract (English)

Current LLM evaluations often rely on a single instruction template, overlooking models' sensitivity to instruction style-a critical aspect for real-world deployments. We present RCScore, a multi-dimensional framework quantifying how instruction formulation affects model responses. By systematically transforming benchmark problems into multiple instruction styles, RCScore reveals performance variations undetected by conventional metrics. Our experiments across ten LLMs on four reasoning benchmarks demonstrate that instruction style can shift accuracy by up to 16.7% points. We introduce Cross-Response Similarity (CRS), a method applying RCScore metrics to measure stylistic self-consistency, and establish its strong correlation with task accuracy, suggesting consistency as a valuable proxy for model reliability. Additional findings show that deterministic decoding produces more stylistically stable outputs, and model scale correlates positively with cross-style consistency. RCScore offers a principled approach to assess instruction robustness.

大模型评估一致性指令风格

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。