揭示PRF对部分查询有害的原因,提出可审计的智能筛选方法。
Explaining When PRF Fails: Participatory Auditing for Selective Query Expansion

- 通过用户参与式审计发现仅20.9%查询受益于PRF
- 25.6%查询因查询漂移体验恶化,避害比增益更关键
- 用LLM重排序器自动复制用户判断,可追溯文档证据
伪相关反馈(PRF)虽平均提升检索效果,但导致大量查询出现查询漂移,损害用户体验,而这一问题被整体离线指标掩盖。现有选择性PRF(sPRF)方法依赖同源排名统计的查询性能预测(QPP),无法解决透明度缺失问题。本文提出‘审计-自动化’两阶段框架:第一阶段在43个TREC Deep Learning 2019查询上开展108名用户的参与式审计,发现仅20.9%查询受益,25.6%遭受体验下降,且避免伤害的价值几乎相当于利用成功扩展的两倍;第二阶段将基于LLM的重排序器改造为系统偏好预测器,基于可检视的文档证据自动复现用户标签。该框架解释了哪些查询受PRF伤害、为何做出筛选决策,并实现大规模可审计性,使原本不透明的检索组件变为以用户为中心、可验证的系统。
原文摘要 · Abstract (English)
Pseudo-Relevance Feedback (PRF) improves retrieval effectiveness on average, but harms a substantial fraction of queries through query drift, an asymmetry hidden by aggregate offline metrics. Existing Selective PRF (sPRF) approaches typically rely on Query Performance Prediction (QPP) methods derived from the same ranking statistics, and therefore inherit, rather than resolve, this opacity. We argue that this is a core explainability problem in IR, and propose a two-stage audit-then-automate framework. In Stage 1, a participatory audit with 108 users across 43 TREC Deep Learning 2019 queries shows that only 20.9% of queries benefit from PRF, while 25.6% suffer a degraded user experience, and that avoiding harm is nearly twice as valuable as exploiting successful expansion. In Stage 2, we repurpose LLM-based rerankers as system preference predictors that replicate these user-derived labels automatically, grounded in inspectable document evidence. Together, the two stages explain which queries PRF harms, why an sPRF decision is made, and how the decision can be inspected at scale, turning an opaque retrieval component into an auditable, user-grounded one.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。