构建首个融合宏观场景的多任务金融推理基准,覆盖美股小盘股2021-2026年数据。
MacroLens: A Multi-Task Benchmark for Contextual Financial Reasoning under Macroeconomic Scenarios

- 整合价格、财报、文本与宏观经济四类信号,设计时序对齐的联合评估框架。
- 涵盖7个任务共4,416只股票,包含215,882篇新闻和1,130个自然语言描述的经济事件。
- 适合研究金融大模型、情境化预测与多源信息融合的学者与从业者使用。
金融决策具有情境性:价格预测、公司估值与事件影响评估需综合价格历史、会计基本面、宏观经济状态及同期文本信息。现有时间序列评估方法因四大限制难以构建统一基准:文本需按发布日期过滤以避免前瞻泄露;季度财报存在1至90天延迟;文件文本与数值字段部分重复;宏观经济状态会跨日历分割泄露。目前无公开基准同时覆盖这四类信号。MacroLens覆盖2021–2026年间4,416只美国中小市值股票,包含七项任务:共享一个时间点面板数据,涵盖4680万条XBRL会计事实、53个宏观经济系列、295,860份SEC文件、215,882篇新闻,以及由1,130个自动识别的宏观经济事件构成的场景层(49种类型),全部以自然语言呈现。任务范围包括情境化预测、公开与私有估值、基于基本面生成陈述、情景条件下的收益分析及房地产估值。我们评估了19种方法,涵盖从简单启发式到时间序列基础模型、微调大语言模型的时间序列模型、零样本大语言模型,还对两种前沿大语言模型与梯度提升基线进行了五步特征-上下文消融实验。数据集已开源:https://huggingface.co/datasets/DeepAuto-AI/MacroLens。
原文摘要 · Abstract (English)
Financial decision-making is contextual: forecasting prices, valuing companies, and assessing event exposure weigh price history, accounting fundamentals, macroeconomic regime, and contemporaneous text. A benchmark over these four signals is hard to build because finance violates four assumptions of time-series evaluation: text must be gated by its publication date to prevent look-ahead, quarterly fundamentals are reported with a one- to ninety-day lag, filing text is partly redundant with the numerical statement fields it accompanies, and macroeconomic regimes leak across calendar splits. No public benchmark addresses all four signals jointly. MacroLens covers 4,416 U.S. small- and micro-cap equities over 2021-2026. Seven tasks share one point-in-time panel of prices, 46.8M XBRL accounting facts, 53 macroeconomic series, 295,860 SEC filings, and 215,882 news articles, plus a scenario layer of 1,130 macroeconomic events across 49 types automatically detected and rendered as natural language. Tasks span contextual forecasting, public and private valuation, statement generation from fundamentals and descriptions, scenario-conditioned returns, and real-estate valuation. We evaluate 19 methods across six families spanning naive heuristics through time-series foundation models, fine-tuned LLM-based time-series models, and zero-shot large language models (LLMs), plus a five-step feature-context ablation on two frontier LLMs and a gradient-boosted baseline. MacroLens is released at https://huggingface.co/datasets/DeepAuto-AI/MacroLens.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。