新基准评估金融推荐模型是否理性决策而非盲目模仿用户行为。
Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation
- 用多视角参考区分用户行为与理性投资判断。
- 模型在理性评分上表现好时,常偏离用户实际选择。
- 适合研究智能投顾、大模型决策可信度的学者和开发者。
现有推荐评测多以用户行为模仿为标准,但在金融咨询中,用户行为可能受市场波动干扰,与长期目标冲突。将用户选择视为唯一真实标准,会混淆行为模仿与决策质量。我们提出 Conv-FinRe,一个对话式、长期性的股票推荐评测基准,要求模型基于开户访谈、逐步市场上下文和咨询对话,在固定投资周期内生成股票排名。关键在于,该基准提供多视角参考,区分描述性行为与基于投资者风险偏好的规范性效用,可诊断模型是依据理性分析、模仿用户噪声,还是受市场趋势驱动。数据源自真实市场与人类决策轨迹,构建了受控咨询对话,评估了多种前沿大模型。结果揭示理性决策质量与行为一致性之间存在持续矛盾:高效用排名的模型往往不匹配用户选择;而行为对齐的模型则易过度拟合短期噪音。数据集已公开于 Hugging Face,代码库可在 GitHub 获取。
原文摘要 · Abstract (English)
Most recommendation benchmarks evaluate how well a model imitates user behavior. In financial advisory, however, observed actions can be noisy or short-sighted under market volatility and may conflict with a user's long-term goals. Treating what users chose as the sole ground truth, therefore, conflates behavioral imitation with decision quality. We introduce Conv-FinRe, a conversational and longitudinal benchmark for stock recommendation that evaluates LLMs beyond behavior matching. Given an onboarding interview, step-wise market context, and advisory dialogues, models must generate rankings over a fixed investment horizon. Crucially, Conv-FinRe provides multi-view references that distinguish descriptive behavior from normative utility grounded in investor-specific risk preferences, enabling diagnosis of whether an LLM follows rational analysis, mimics user noise, or is driven by market momentum. We build the benchmark from real market data and human decision trajectories, instantiate controlled advisory conversations, and evaluate a suite of state-of-the-art LLMs. Results reveal a persistent tension between rational decision quality and behavioral alignment: models that perform well on utility-based ranking often fail to match user choices, whereas behaviorally aligned models can overfit short-term noise. The dataset is publicly released on Hugging Face, and the codebase is available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。