大模型在序数分类中受位置偏倚影响,需综合评估性能与稳定性。
Position Bias in Ordinal Classification: A Systematic Evaluation

- 通过三种探针系统测试十种前沿大模型的位置偏倚问题。
- 准确率与稳定性常不一致,仅低基数标签能同时提升两者。
- 基于比较的列表式推理效果最佳,但跨模型泛化能力差。
大型语言模型在序数分类任务中的应用日益广泛,但提示顺序的语义等价变化可能改变其预测结果。我们系统性地实验评估了标签顺序、示范顺序和示范位置三种位置偏倚来源的影响。首先,在十种前沿大模型上对同一序数分类任务进行测试,发现所有模型均对三类位置因素敏感,表明该问题普遍存在。其次,我们在五个数据集上变化八类提示、任务和模型层面的因子,结果显示准确率与稳定性常不一致,仅较低基数标签水平能稳定提升两者表现。第三,我们对比了点级、成对、列表级推理方式,以及多种聚合与去偏方法及联合配置,发现现有修正手段无法可靠缓解偏倚,而基于比较的列表式方法虽表现最优,但在不同模型与偏倚来源间迁移效果不一。研究结果表明,位置鲁棒性取决于完整系统配置,而非仅由模型决定。因此,序数分类系统应综合考虑预测性能与稳定性进行选择。
原文摘要 · Abstract (English)
Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predictions. We conduct systematic experiments to characterize positional bias from label order, demonstration order, and demonstration placement. First, we apply the three probes to ten frontier LLMs on a common ordinal-classification task; every model is sensitive to all three positional sources, showing that the problem is pervasive. Second, we vary eight prompt-, task-, and model-level factors across five datasets; accuracy and stability are often misaligned, and only lower scale cardinality consistently improves both. Third, we compare pointwise, pairwise, and listwise inference, alternative aggregation and debiasing methods, and joint configurations; the tested corrections do not provide a reliable remedy, while a comparison-based listwise formulation offers the best balance but transfers unevenly across models and bias sources. These findings show that positional robustness depends on the full system configuration rather than the model alone. Ordinal-classification systems should therefore be selected jointly for predictive performance and stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。