arXiv:2608.04552cs.CLcs.LG2026-08

提出关系响应场理论,量化大模型回复修复的内在难度。

Relational Response Fields: A General Theory of Black-Box LLM Response Consistency and Recovery

论文配图:Relational Response Fields: A General Theory of Black-Box LLM Response Consistency and Recovery
图 1 · 摘自论文原文
  • 用关系场建模问答变换下的响应一致性规律。
  • 定义可修复性指标γ_k,证明修复误差下界不可超越。
  • 验证理论预测:一致但错误的回复仍可修复,适合可信推理研究者。

黑箱语言模型的可靠性常通过采样、提示、投票或迭代修正单个答案来提升。我们提出一个前置问题:一组黑箱回复是否具备可恢复性?将查询经类型变换后的响应表示为关系响应场(RRF)。边传输编码了响应在改写、缩放、分解、重构等任务对称性下的变化规则;锚点编码独立可信证据,如执行结果或验证器输出。对于关系算子D、锚点算子A,以及最多k个受损响应节点,我们定义γ_k(D,A)为黑箱回复恢复的内在难度。当且仅当所有k节点扰动均可识别时,γ_k为正;它给出与1/γ_k成比例的确定性稳定界;匹配的两点极小极大下界表明,任何估计器都无法改善该依赖关系。因此一致性不等于正确性:仅依赖关系的方法对零空间方向(如共通幻觉)无感。我们推导出稀疏场修复算法,同时区分信息论可识别性与凸优化所需的更强零空间条件。受控定理测试与黑箱数学/代码实验验证了四个理论固定结论:一致性-真实性分离、锚点相变、冗余饱和,以及跨模型、跨任务对修复难度的预测能力。结果支持γ_k(D,A)是响应恢复实例的可测量属性,而非某一修复启发式附带的评分。

原文摘要 · Abstract (English)

Black-box language-model reliability is commonly pursued by sampling, prompting, voting, verifying, or iteratively revising individual answers. We ask a prior question: \emph{what determines whether a collection of black-box responses is recoverable at all?} We represent responses to typed transformations of a query as a \emph{relational response field} (RRF). Edge transports encode how valid responses must change under paraphrase, scaling, decomposition, refactoring, or other task symmetries; anchors encode independently trusted evidence such as execution or a verifier. For relation operator $D$, anchor operator $A$, and at most $k$ corrupted response nodes, we identify $γ_k(D,A)$ as the intrinsic difficulty of black-box response recovery. It is positive exactly when every $k$-node corruption is identifiable; it gives a deterministic stability bound proportional to $1/γ_k$; and a matching two-point minimax lower bound shows that no estimator can improve this dependence. Thus consistency is not truth: relation-only methods are blind to null directions, including shared hallucinations. We derive sparse field-repair algorithms while separating information-theoretic identifiability from the stronger null-space conditions required by convex optimization. Controlled theorem tests and black-box mathematics/code experiments evaluate four theory-fixed consequences: consistency--truth separation, anchor phase transitions, redundancy saturation, and cross-model, cross-task prediction of repair difficulty. The results support $γ_k(D,A)$ as a measurable property of a response-recovery instance, rather than a score attached to one repair heuristic.

大模型可靠性一致性分析可修复性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。