测试中文开源模型在金融文本理解中的表现,发现部分模型已接近顶尖水平。
Can Open-Weight Models Compete on Financial Text Comprehension?

- 基于2967组财报问答数据,评估20个模型的金融理解能力
- 开源模型Kimi K2.6准确率达第三,非推理模型也进入前五
- 信息检索仍是主要瓶颈,且国产模型存在不合理的拒答现象
近期,中国AI实验室推出的开源语言模型在基准测试中逐渐逼近商业领先模型。然而其在真实金融任务中的可靠性仍待验证。我们更新了Financial Touchstone基准,包含495份国际年报中的2,967个问题-上下文-答案三元组。评估范围扩展至20个模型,覆盖10家厂商,包括GLM 4.7、GLM 5、Kimi K2.6、DeepSeek V3.2等新发布开源模型,以及阿里巴巴的专有旗舰Qwen3-Max。Anthropic的Claude Opus 4.6准确率最高(88.4%),Google Gemini 2.5 Pro hallucination率最低(0.08%)。值得注意的是,开源模型Kimi K2.6准确率位居第三,非推理模型GLM 5和Mistral 3分别位列第四和第五,挑战了“强金融理解需推理架构或专有权重”的假设。信息检索失败占所有错误的48.9%。此外,我们发现中国模型中的地缘政治内容过滤机制会拒绝合法金融问题(0.08%尝试),且拒答行为受访问路径影响显著。完整数据集与评估框架已公开。
原文摘要 · Abstract (English)
Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability on real-world financial tasks remains largely untested. We updated the Financial Touchstone benchmark, which now has 2,967 question context-answer triplets across 495 international annual reports. We also apply a new set of models on the benchmark, expanding coverage from eleven to twenty models across ten providers, including recent open-weight models such as GLM 4.7, GLM 5, Kimi K2.6, and DeepSeek V3.2, as well as Alibaba's proprietary flagship Qwen3-Max. Anthropic's Claude Opus 4.6 achieves the highest accuracy (88.4%), while Google's Gemini 2.5 Pro maintains the lowest hallucination rate (0.08%). Notably, the open-weight Kimi K2.6 ranks third in accuracy, and the non-reasoning models GLM 5 and Mistral 3 rank fourth and fifth, challenging the assumption that reasoning architectures or proprietary weights are a prerequisite for strong financial comprehension. Information retrieval remains the primary bottleneck, accounting for 48.9% of all failures. We also document a new finding: geopolitical content filters in Chinese models refuse legitimate financial questions (0.08% of attempts), sometimes without clear reason, and the refusal behavior depends on the access route as much as on the model. The complete dataset and evaluation framework are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。