构建长期投资决策评估框架,检验AI顾问的真正影响力
ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors

- 用多智能体模拟器追踪用户投资轨迹,基于真实行为数据校准
- 2021-2026年中文基金市场数据验证,部分大模型顾问表现更优
- 关注从对话到决策的长期路径,适合评估金融AI顾问有效性
对话式投资顾问不仅影响用户认知,还改变其随市场变化的决策行为。现有评估多聚焦回答质量或结果,难以审计从顾问语言到投资者行为的长期路径。我们提出ShiJianBench,一个基于历史市场反馈的离线评估框架。核心是多智能体投资者模拟器,具备显式演化状态变量、动机驱动决策、长期记忆和对话引导更新。该模拟器基于7,199名真实用户的行为模式进行校准,通过投资者侧、服务侧、内容侧三类指标,在严格合规约束下评估顾问策略。基于2021至2026年中国基金市场数据的实验发现,部分大模型顾问在个性化内容上显著更强,同时保持与顶尖水平相当的投资者轨迹表现。结果揭示了高质量回复与有效长期干预之间的系统性差异,推动面向轨迹的对话顾问评估。
原文摘要 · Abstract (English)
Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily assess response quality or observed outcomes, leaving the long-horizon pathway from advisor language to investor behavior difficult to audit. We introduce ShiJianBench, an offline framework for evaluating conversational investment advisors through matched investor trajectories under fixed historical market feedback. At its core is a multi-agent investor simulator with explicit evolving state variables, motive-driven deliberation, long-term memory, and dialogue-grounded updates. The simulator is calibrated against aggregate behavioral patterns from 7,199 real users, and advisor policies are evaluated using separate investor-side, service-side, and content-side metrics under a hard compliance gate. Experiments on Chinese fund-market traces from 2021 to 2026 identify a stable leading group of LLM advisors that combines substantially stronger personalized content with competitive investor-side trajectory outcomes. These results reveal a systematic distinction between producing a high-quality response and delivering an effective long-horizon intervention, motivating trajectory-aware evaluation of conversational advisors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。