arXiv:2512.19717cs.LGcs.AI2025-12

用热力学聚焦法高效找稀有好解,适用于语言生成与强化学习。

Thermodynamic Focusing for Inference-Time Search: Practical Methods for Target-Conditioned Sampling and Prompted Inference

  • 将搜索转为目标条件重加权,复用现有采样器和相似度函数。
  • 在语言生成和稀疏奖励导航中显著降低样本需求,提升效率。
  • 适合需要高效推理的生成任务,如长文本生成与复杂规划。

在语言生成、规划和强化学习中,从巨大候选空间中寻找稀有但有用的解是一个常见挑战。本文提出实用框架——反因果聚焦算法(ICFA),将搜索视为目标条件重加权过程。ICFA利用已有提议采样器和任务相关的相似度函数,构建聚焦采样分布,并自适应控制聚焦强度以避免退化。我们提供了清晰的操作指南,基于有效样本量的稳定性诊断,简洁的理论分析说明何时可减少样本量,并完成两项可复现实验:约束语言生成与稀疏奖励导航。此外,我们展示了结构化提示如何实现近似语言级的ICFA,并提出结合提示推理与算法重加权的混合架构。

原文摘要 · Abstract (English)

Finding rare but useful solutions in very large candidate spaces is a recurring practical challenge across language generation, planning, and reinforcement learning. We present a practical framework, \emph{Inverted Causality Focusing Algorithm} (ICFA), that treats search as a target-conditioned reweighting process. ICFA reuses an available proposal sampler and a task-specific similarity function to form a focused sampling distribution, while adaptively controlling focusing strength to avoid degeneracy. We provide a clear recipe, a stability diagnostic based on effective sample size, a compact theoretical sketch explaining when ICFA can reduce sample needs, and two reproducible experiments: constrained language generation and sparse-reward navigation. We further show how structured prompts instantiate an approximate, language-level form of ICFA and describe a hybrid architecture combining prompted inference with algorithmic reweighting.

生成模型强化学习推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。