构建首个跨市场财报因子研究基准,支持跨国预测与实证分析。
CrossAlpha: An Annual-Report Benchmark for Cross-Market Factor Research (with LLM Agents)

- 将多国财报统一为十类英文业务描述,解决语言与制度差异问题。
- 通过主成分去相关处理,生成1900万条跨市场公司配对评分。
- 在美日跨市场预测中表现优于本地文本基线,适合金融NLP研究者使用。
跨市场因子研究旨在检验一个或多个市场的公司级信号是否能预测目标市场的收益,但现有公开基准无法支持从披露信息到实际回报的评估。构建此类基准面临三大挑战:不同语言和监管体系下的文件差异、由共同披露内容引起的相似性偏差,以及需满足可行交易时间对齐的信号评估要求。本文提出 extbf{CrossAlpha},一个面向跨市场因子研究的公开年度报告基准。该基准通过三个核心组件应对上述挑战: extit{披露提炼}(Disclosure Distillation),将异构财报标准化为十类英文业务描述; extit{残差模式图构建}(Residual Schema Graph Construction),基于模式层面披露信息,构建经主成分分析去相关处理的跨市场公司对得分; extit{时序对齐评估}(Timing-Aligned Evaluation),将图结构与11年日度OHLCV数据对齐,依据可行的跨市场执行协议构建未来收益标签。CrossAlpha覆盖美国、日本、台湾、韩国和香港约3,600家公司的10,700份公司-年度报告,发布约1,900万条有向公司对得分。实验表明,在美→日设定下,基于披露的跨市场同行表现显著优于本地文本、行业代码及回报相关同行(ICIR 0.39 vs. 0.07–0.18);多数目标市场中,跨市场信号亦优于本地文本基线。CrossAlpha为跨市场金融自然语言处理提供了开源、可复用、以回报为基准的评估平台。
原文摘要 · Abstract (English)
Cross-market factor research studies whether firm-level signals from one or more markets can predict returns in a target market, but existing public benchmarks do not support cross-market disclosure-to-return evaluation. Building such a benchmark is challenging because filings differ across languages and regulatory systems, disclosure-derived similarity can be biased by common reporting components, and cross-market signals must be evaluated under feasible trading-time alignment. We introduce \textbf{CrossAlpha}, a public annual-report benchmark for cross-market factor research. CrossAlpha addresses these challenges through three corresponding components: \emph{Disclosure Distillation}, which standardises heterogeneous filings into ten-category English business descriptions; \emph{Residual Schema Graph Construction}, which builds PCA-whitened cross-market firm-pair scores from schema-level disclosures; and \emph{Timing-Aligned Evaluation}, which pairs the graph with 11 years of daily OHLCV data to construct forward-return labels under feasible cross-market execution protocols. CrossAlpha covers about 3,600 firms and 10,700 firm-year reports from the United States, Japan, Taiwan, South Korea, and Hong Kong, and releases about 19M directed firm-pair scores. In experiments, disclosure-derived cross-market peers outperform domestic text, industry-code, and return-correlation peers in the US-to-Japan setting (ICIR 0.39 versus 0.07--0.18), and cross-market sources beat the domestic text baseline in most target markets. CrossAlpha offers an open-sourced, reusable, return-grounded benchmark for cross-market financial NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。