arXiv:2608.17781cs.CL2026-08

不同读者对证据好坏的判断差异大,且稳定偏好不能直接用于预测干预效果。

Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility

  • 区分证据活性、排序偏好和方向性,发现读者排序偏好在多场景下稳定。
  • 在二分类事实核查中,读者方向性偏好强(ρ=0.75),但开放问答中弱。
  • 即使排序相似,跨读者干预效果也无法可靠转移,偏好不等于可迁移决策。

机器学习系统越来越多地依据下游模型身份做决策,但这只有在模型差异具有可复用结构而非仅依赖输入时才有效。我们在检索增强生成(RAG)中通过受控干预检验这一问题:固定查询、证据、任务、评分与干预方式,九个读者在33%的共同影响单元中对效果方向存在分歧;读者×查询交互解释了29.8%的证据效用方差,远超8.4%的随机基线;自选证据使F1提升+0.031(t=3.39)。进一步追问:哪些异质性是读者跨查询的稳定属性?我们分离出三个可测量项——证据活性、序数偏好与条件符号方向——发现序数几何在四个独立设置中稳定(分半一致性ρ=0.60–0.83):留一法干预、PRISM偏好、RAMDocs与RAGuard。符号几何则具任务依赖性:在开放式问答中较弱(0.14, 0.35),尤其对误导或无关证据;但在二分类事实核查中较强(0.75),且无显著序数差距,但仍低于其稀疏性匹配的上限。稀疏性、解码噪声与指标伪影无法解释主要的序数-符号差距。稳定序数相似性无法预测跨读者干预转移效果(奥兰多距离ρ=-0.27;后悔可靠性-0.28)。读者特定效用存在,但偏好≠干预:稳定排序相似性不支持帮助/伤害决策的转移。

原文摘要 · Abstract (English)

ML systems increasingly condition decisions on downstream model identity, but this is useful only if model-specific differences form reusable structure rather than input-local interactions. We test this in retrieval-augmented generation (RAG), where evidence utility can be measured under controlled interventions. Holding query, evidence, task, scoring, and intervention fixed, nine readers disagree on effect sign in 33\% of jointly affected cells; reader$\times$query interaction explains 29.8\% of utility variance versus an 8.4\% permutation null; and self-selected evidence improves F1 by $+0.031$ ($t=3.39$). We then ask the sharper question: \emph{which components of this heterogeneity are stable reader properties across queries?} Separating three measurable objects---evidence \emph{activity}, \emph{ordinal preference}, and \emph{conditional signed direction}---we find ordinal reader geometry stable across four independent settings (split-half $ρ=0.60$--$0.83$): leave-one-out interventions, PRISM preferences, RAMDocs, and RAGuard. Signed geometry is task-bounded: weak in open-ended QA (0.14, 0.35), especially for misleading and irrelevant evidence, but strong in binary fact-checking (0.75) with no significant ordinal gap, though still below its sparsity-matched ceiling. Sparsity, decoding noise, and metric artifacts do not explain the main ordinal--signed gap. Finally, stable ordinal similarity fails to predict cross-reader intervention transfer (oracle-distance $ρ=-0.27$; regret reliability $-0.28$). Reader-specific utility exists, but preference is not intervention: stable ranking similarity does not license transfer of help/harm decisions.

RAG偏好分析可迁移性证据效用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。