首次揭示大模型推荐系统中隐藏表示与输出的公平性脱节问题
Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders

- 构建双层公平性评测框架,同时检测可见输出与隐藏表征的偏差
- 发现78%用户存在输出稳定但内部表示剧烈变化的现象,传统审计无法捕捉
- 揭示内部公平与输出公平存在根本矛盾,适合算法公平性研究者参考
针对基于大模型的推荐系统,现有公平性审计多聚焦于可观测输出,隐含假设稳定推荐反映稳定内部处理。本文提出FairGap,首个联合评估推荐公平性的基准,涵盖可观测输出偏移(OBS)与隐藏表征偏移(IBS),通过控制反事实身份探测(性别、年龄、种族)进行测量。二者关系以表征-输出对齐度(ROA)量化,结合象限诊断识别用户级隐藏-输出不一致。在六种开源大模型家族、三个领域上的应用显示:普遍存在隐藏-输出解耦现象,ROA极少超过0.22;部分用户虽输出稳定但内部表示显著变化,此类模式输出审计无法识别。此外,激活操控可使IBS降低至1/8,却同步恶化OBS,暴露内部与输出公平间的根本张力,现有框架难以诊断。
原文摘要 · Abstract (English)
Fairness audits for LLM-based recommenders have largely focused on observable outputs, implicitly assuming that stable recommendations reflect stable internal processing. We challenge this assumption with FairGap, the first benchmark to jointly evaluate recommendation fairness at two levels: observable output shift (OBS) and hidden representation shift (IBS), measured through controlled counterfactual identity probes across gender, age, and race. Their relationship is summarized via Representation-Output Alignment (ROA), with quadrant diagnostics for identifying user-level hidden-output mismatch. Applied to six open-weight LLM families across three domains, FairGap reveals pervasive hidden-output decoupling: ROA rarely exceeds 0.22, and a non-negligible user population shows stable outputs despite substantial internal shifts, a mode that output-only audits cannot detect by design. Further, activation steering that reduces IBS by up to 8x simultaneously worsens OBS, demonstrating a fundamental tension between internal and output-level fairness that existing frameworks are unequipped to diagnose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。