用多粒子流图实现高效在线反馈搜索,兼顾全局探索与偏好对齐。
Sequentially-Controlled Interactive Multi-Particle Flow-Maps for Online Feedback-Driven Search

- 通过多粒子流图动态调整样本分布,支持持续反馈下的智能搜索。
- 在多个任务上显著优于基线,避免模式坍塌并提升探索效率。
- 适合需要在线学习和多样化输出的生成式搜索场景。
尽管生成模型已实现无需训练的奖励对齐,但现有方法通常仅在分布的狭窄区域中表现良好。当偏好未知且需通过序列反馈揭示时,需广泛探索以发现高价值区域。为此,我们提出顺序控制的交互式多粒子流图(IMPFM),一种样本高效的在线反馈驱动搜索框架。IMPFM逐步将一组交互粒子导向目标分布,保持广泛覆盖以适应异构偏好对齐。其引入基于流映射的有原则且高效的后验样本共享机制,在每次重采样步骤中利用整个粒子集合的集体后验样本校正个体粒子漂移,最大化样本效用,实现全局探索并主动缓解标准控制框架中的奖励过优化问题。结合多粒子交互的探索-利用重加权机制,该顺序修正的多粒子动力学显式保留结构多样性,克服了标准SMC采样器固有的权重退化问题。关键的是,我们证明该采样框架生成了一种多粒子交互感知的Feynman-Kac校正器,逐步引导多粒子系统趋向KL倾斜的目标分布,促进全局探索并防止模式坍塌。在多种搜索与对齐任务上的广泛实验及严格消融分析证实IMPFM优于现有基线。
原文摘要 · Abstract (English)
While generative models have enabled training-free reward alignment, current methods typically excel in local exploration within narrow regions of the underlying distribution. These approaches struggle when preferences are unknown a priori and only revealed through sequential feedback-a scenario demanding broad exploration to uncover high-utility regions. To address this, we propose Sequentially-Controlled Interactive Multi-Particle Flow-Maps (IMPFM), a framework for sample-efficient online feedback-driven search. IMPFM progressively transports a group of interactive particles toward the target distribution, maintaining the broad coverage essential for heterogeneous preference alignment. IMPFM introduces a principled and efficient posterior sample sharing mechanism across particles powered by flow maps. By correcting individual particle drift with the collective posterior samples of the entire ensemble at each resampling step, the framework maximizes sample utility to enable global exploration while actively mitigating reward over-optimization, typical of standard control frameworks. Paired with a principled exploration-exploitation reweighting mechanism involving multi-particle interaction, this sequentially corrected multi-particle dynamics explicitly preserves structural diversity and overcomes the weight degeneracy inherent to standard SMC samplers. Crucially, we prove that the resulting sampling framework yields a multi-particle interaction-aware Feynman-Kac corrector that progressively steers the multi-particle system toward a KL-tilted target distribution, facilitating global exploration and preventing mode collapse. Extensive empirical evaluations and rigorous ablations across diverse search and alignment tasks confirm the efficacy of IMPFM over existing baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。