arXiv:2608.10441cs.LGcs.CL2026-08

检测信号有效不等于学会用它,低信噪比下根本学不会精准决策。

Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents

论文配图:Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents
图 1 · 摘自论文原文
  • 区分平均效应检测与个体决策学习,揭示学习路由的理论极限
  • 三个数据集均低于信噪比阈值,学习路由全败于随机,仿噪声可复现全部假性收益
  • 提出结构化假设嵌入(SHE),适合设计时确定使用策略而非实时决策

许多系统需支付单例成本获取模型生成的辅助观测(如结构化推理、慢速人工标注、高成本测量),并决定何时使用该信号。本文指出一个易被忽视的关键区别:检测信号平均有效,并不等于学会在每个实例中正确决策;而奖励信噪比(reward-SNR)存在一个可检测下限。即使信号忠实且样本内最优选择者(选前b个)表现出显著收益,实际部署策略仍无法学习何时获取该信号:在每印象、聚类、模式和增益树等粒度上,学习路由始终不如随机,且匹配时刻的噪声伪对照组可复现≥100%的伪增益——所谓“可学习结构”实为噪声的顺序统计特征。我们通过区分平均效应检测与个体策略学习,建立奖励信噪比下限ρ*(N) ≈ 2.8/√N。以结构化假设嵌入(SHE)为例,其将用户历史转化为排序、置信度评分、证据支撑的意图假设,并融合进推荐系统。在MIND、REES46、Amazon-Beauty三个公开数据集上,SHE表现忠实且可校准,但其价值依赖主干模型与运行场景(有序GRU上+0.0114,95% CI [+0.0030, +0.0209],而全局冗余差距接近零),所有数据集均低于信噪比下限,导致学习获取策略失效。真正的可行单元是设计时设定的场景门控,而非实时决策策略。代码与一键复现已发布。

原文摘要 · Abstract (English)

Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expensive measurement -- and then must decide when the acquired signal is worth using. Our thesis is a distinction that is easy to miss: detecting that such a signal helps on average is not the same as learning to act on it per instance, and a reward-SNR floor governs when the second is even possible. Even when the signal is faithful and an in-sample oracle picking the top-b examples by realized reward shows a sizable apparent gain, no deployable policy can learn when to acquire it: across per-impression, cluster, regime, and uplift-tree granularities, learned routing never beats random, and a matched-moment noise placebo reproduces >=100% of the oracle's apparent gain -- the apparent "learnable structure" is order statistics of noise. We explain this with one distinction, detecting a mean effect vs. learning a per-instance acquisition policy, and a reward-SNR detectability floor: routing is estimable offline only if the reward SNR rho clears rho*(N) ~= 2.8/sqrt(N), with a positive control confirming a true low-SNR limit rather than a broken pipeline. As a concrete instantiation we introduce Structured Hypothesis Embeddings (SHE): a frozen LLM turns a user history into ranked, confidence-scored, evidence-grounded intent hypotheses, fused into a recommender. On three public datasets (MIND, REES46, Amazon-Beauty), SHE is faithful and calibratable, yet its value is backbone- and regime-conditional (significant over an ordered GRU, +0.0114, 95% CI [+0.0030, +0.0209], but a global redundancy gap indistinguishable from zero), and learned acquisition collapses at every granularity because all three datasets sit below the floor. The realizable unit is a design-time regime gate, not a per-instance policy. We release code and a one-command reproduction.

强化学习信号检测决策机制信噪比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。