让大模型理解时间序列决策,通过动态修正证据提升判断准确率。
Trajectory-Aware Retrieval Agents for Temporal Decision- Making
- 用轨迹感知的闭环框架迭代优化检索证据,保留时间脉络
- 在医疗问答中显著超越零样本模型和传统检索方法
- 适合需要长期时序推理的医疗、金融场景
我们研究使用大语言模型代理从长篇时序文本中进行决策的问题。标准检索增强生成(RAG)流程将时间结构打碎为孤立片段,丢失对正确下游决策至关重要的时间信息。我们提出TLM(轨迹语言模型),一种闭环智能体框架,通过SHAP引导反馈迭代优化证据集。关键技术是检索片段嵌入上的潜在增长曲线模型(LGCM),可解释地检测轨迹趋势、拐点与信息缺口。在评分校准假设下(实践中近似成立),迭代优化过程单调不减正确标签的概率。实验评估在三个时序决策任务上:医疗问答、财报惊喜预测和隔夜股价跳空预测。TLM在医疗任务上显著优于零样本模型和标准RAG,在两个金融任务上均实现一致且具有经济意义的提升。
原文摘要 · Abstract (English)
We study the problem of decision-making from long-form, temporally structured text using large language model (LLM) agents. Standard retrievalaugmented generation (RAG) pipelines fragment chronological context into isolated snippets, discarding the temporal structure that is often critical for correct downstream decisions. We introduce TLM (Trajectory Language Model), a closed-loop agentic framework that iteratively refines the evidence set using SHAP-guided feedback. The key technical contribution is the latent growth curve model (LGCM) over retrieved chunk embeddings, which provides an interpretable mechanism for detecting trajectory trends, turning points, and information gaps. We show that, under a scorer-calibration assumption (which holds approximately in practice), the iterative refinement procedure is monotonically non-decreasing in the probability assigned to the correct label. Empirically, TLM is evaluated on three temporally grounded decision tasks: medical question answering, earnings call surprise prediction, and overnight stock gap prediction. TLM substantially outperforms both zero-shot LLM baselines and standard retrieval-augmented approaches on the medical task, and yields consistent, economically meaningful gains on the two financial tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。