用文字反馈优化文本,无需修改模型权重
Feedback Descent: Open-Ended Text Optimization via Pairwise Comparison
- 通过结构化文字反馈生成方向性优化信号
- 在分子筛选中超越99.9%已有化合物
- 适用于提示、代码、分子等多场景
我们提出反馈下降(Feedback Descent)框架,通过结构化文本反馈而非单一标量奖励来优化文本内容(如提示、代码、分子)。该方法保留详细批评信息,避免压缩为二元偏好,扩大了偏好学习中的信息瓶颈,实现文本空间内的定向优化而非权重空间。利用上下文学习将结构化反馈转化为类梯度方向信息,可在不修改模型权重的情况下完成纯推理时迭代优化,且任务无关。在三个不同领域评估显示,其性能优于最先进的提示优化(GEPA)、强化学习方法(GRPO、REINVENT)以及专用图神经分子优化器。在DOCKSTRING分子发现基准测试中,成功识别出在六个蛋白靶点上超过26万种化合物中99.9百分位的新型药物分子。
原文摘要 · Abstract (English)
We introduce \textit{Feedback Descent}, a framework that optimizes text artifacts -- prompts, code, and molecules -- through structured textual feedback, rather than relying solely on scalar rewards. By preserving detailed critiques instead of compressing them to binary preferences, Feedback Descent widens the information bottleneck in preference learning, enabling directed optimization in text space rather than weight space. We show that in-context learning can transform structured feedback into gradient-like directional information, enabling targeted edits. Unlike prior approaches that collapse judgments into single bits, our evaluators pair each comparison with textual feedback, which functions as high-bandwidth supervision. The iteration loop is done purely at inference time, without modifying any model weights, and is task-agnostic. We evaluate Feedback Descent on three diverse domains and find that it outperforms state-of-the-art prompt optimization (GEPA), reinforcement learning methods (GRPO, REINVENT), and even specialized graph-based molecular optimizers. In the DOCKSTRING molecule discovery benchmark, Feedback Descent identifies novel drug-like molecules surpassing the $99.9$th percentile of a database with more than $260{,}000$ compounds across six protein targets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。