构建金融情感推理的全流程数据集,助力模型精准理解与修正。
SenseAI: A Human-in-the-Loop Dataset for RLHF-Aligned Financial Sentiment Reasoning
- 通过人类反馈构建包含推理链的金融情绪数据集
- 发现模型存在隐性推理漂移和置信度偏差等系统性错误
- 适合金融AI对齐、评估与模型优化研究者使用
我们提出SenseAI,一个基于人机协同验证的金融情绪推理数据集,不仅包含模型输出,还完整记录其推理过程。不同于现有资源,SenseAI整合了推理链条、置信度评分、人工校正信号及真实市场结果,契合强化学习从人类反馈(RLHF)范式。数据集涵盖40只美股标的、13类金融数据,共1,439个标注样本,可直接用于现代大模型微调流程。分析揭示模型行为中存在系统性模式,包括一种新型失败模式——隐性推理漂移(即引入输入未包含的信息),以及持续的置信度误判与前瞻性预测倾向。这些发现表明,金融推理中的大模型错误并非随机,而具有可预测且可纠正的规律,支持利用结构化人机协同数据进行针对性优化。本文讨论其对金融AI系统的影响,并指出SenseAI在模型评估与对齐中的应用潜力。
原文摘要 · Abstract (English)
We introduce SenseAI, a human-in-the-loop (HITL) validated financial sentiment dataset designed to capture not only model outputs but the full reasoning process behind them. Unlike existing resources, SenseAI incorporates reasoning chains, confidence scores, human correction signals, and real-world market outcomes, providing a structure aligned with Reinforcement Learning from Human Feedback (RLHF) paradigms. The dataset consists of 1,439 labelled data points across 40 US-listed equities and 13 financial data categories, enabling direct integration into modern LLM fine-tuning pipelines. Through analysis, we identify several systematic patterns in model behavior, including a novel failure mode we term Latent Reasoning Drift, where models introduce information not grounded in the input, as well as consistent confidence miscalibration and forward projection tendencies. These findings suggest that LLM errors in financial reasoning are not random but occur within a predictable and correctable regime, supporting the use of structured HITL data for targeted model improvement. We discuss implications for financial AI systems and highlight opportunities for applying SenseAI in model evaluation and alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。