用真实对话行为评估大模型公平性,比标准测试更可靠
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
- 将测试题转为对话种子,观察模型在多角色交流中的行为变化
- 800万条对话显示模型有稳定的不公平行为模式
- 适合关注模型真实社会交互公平性的研究者
大模型的公平性应通过真实对话中的行为模式评估,而非标准化问答测试。我们发现,标准化测试范式存在结构性不可靠:表面提示设计虽与公平性无关,却主导了分数方差,改变公平性结论的方向和程度,并导致模型排名严重不一致。为此,我们提出MAC-Fairness框架,将可控变量嵌入真实对话评估中,考察模型在身份变化时的差别对待行为。将标准化测试题作为对话引子,而非评估工具,分析跨模型、多身份情境下800万条对话记录中的自我立场坚持度(从自身视角)和对他人观点的接受度(从对方视角)。结果显示,基于真实行为的评估揭示出稳定、模型特异的差别对待模式,可跨不同公平性基准泛化,这是标准化测试无法提供的证据。
原文摘要 · Abstract (English)
LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks. We show that the standardized-test paradigm can be structurally unreliable: surface-level prompt construction choices, although entirely orthogonal to the fairness question being tested, account for the majority of score variance, shift fairness conclusions in both the direction and the magnitude, and result in severe discordance in model rankings. We develop MAC-Fairness, a framework that embeds controlled variation factors into in-situ behavioral evaluation, examining how models' disparate-treatment behaviors shift when identity is varied as part of natural multi-agent conversation. Repurposing standardized-test questions as conversation seeds rather than as the evaluation instrument, we evaluate within-model differences in position persistence (how they hold positions, from the self-perspective) and peer receptiveness (how receptive they are to peers, from the other-perspective) across 8 million conversation transcripts spanning multiple models and identity presence configurations. In-situ behavioral evaluation reveals stable, model-specific, disparate-treatment behavioral signatures that could generalize across different fairness benchmarks, a form of evidence the standardized-test paradigm does not offer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。