arXiv:2604.23988cs.LGcs.AI2026-04中稿 · ICLR被引 1

用事后结果训练模型生成带理由的金融建议,无需人工标注。

Hindsight Preference Optimization for Financial Time Series Advisory

论文配图:Hindsight Preference Optimization for Financial Time Series Advisory
图 1 · 摘自论文原文
  • 用实际结果回溯判断建议优劣,生成无监督偏好对
  • 40亿参数模型超越2350亿参数教师模型表现
  • 适合需要可解释金融决策的量化研究者

时间序列模型预测数值;决策者需要的是带有推理、可操作建议和风险管理的方向性指导。训练语言模型生成此类预测建议面临根本挑战:质量依赖于预测时未知的结果。我们结合强化学习中的两个思想——利用执行时不可用的信息回溯生成训练信号,以及偏好对齐——提出事后偏好优化(Hindsight Preference Optimization):通过观察到的实际结果,让大语言模型在无法用标量指标捕捉的维度上对候选建议进行排序,从而生成无需人工标注的偏好对用于直接偏好优化(DPO)。我们在基于视觉-语言模型的标普500股票时间序列预测建议任务上应用该方法,实验表明,一个40亿参数的模型在准确率和建议质量上均优于其2350亿参数的教师模型。

原文摘要 · Abstract (English)

Time series models predict numbers; decision-makers need advisory -- directional signals with reasoning, actionable suggestions, and risk management. Training language models for such predictive advisory faces a fundamental challenge: quality depends on outcomes unknown at prediction time. We bridge two ideas from reinforcement learning -- using information unavailable during execution to retrospectively generate training signal, and preference alignment -- and propose Hindsight Preference Optimization: observed outcomes let an LLM judge rank candidate advisories on dimensions that scalar metrics cannot capture, producing preference pairs for DPO without human annotation. We apply this to Vision-Language-Model-based predictive advisories on S&P 500 equity time series, demonstrated by a 4B model outperforming its 235B teacher on both accuracy and advisory quality.

金融预测偏好优化大模型时序分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。