将通用未来预测扩展至金融、零售等高价值领域,评估大模型在关键场景中的实际预测能力。
FutureX-Pro: Extending Future Prediction to High-Value Vertical Domains
- 构建面向金融、零售等领域的专用未来预测框架
- 在多个垂直领域验证大模型预测精度存在显著差距
- 适合关注大模型落地应用的工业界与政策研究者
基于已建立通用未来预测基准的FutureX,本文提出FutureX-Pro,涵盖FutureX-Finance、FutureX-Retail、FutureX-PublicHealth、FutureX-NaturalDisaster和FutureX-Search。该框架将智能体式未来预测拓展至高价值垂直领域。尽管通用智能体在开放域搜索中表现良好,其在资本密集型与安全敏感型行业的可靠性仍待验证。本研究聚焦金融、零售、公共卫生与自然灾害四大关键领域,对入门级但基础性的预测任务(如市场指标、供应链需求、疫情趋势、灾害演化)进行基准测试。通过复用FutureX的无污染实时评估流程,检验当前最先进的(SOTA)智能体大模型是否具备工业部署所需的领域知识根基。结果表明,通用推理模型在高价值场景下的预测精度仍有明显不足。
原文摘要 · Abstract (English)
Building upon FutureX, which established a live benchmark for general-purpose future prediction, this report introduces FutureX-Pro, including FutureX-Finance, FutureX-Retail, FutureX-PublicHealth, FutureX-NaturalDisaster, and FutureX-Search. These together form a specialized framework extending agentic future prediction to high-value vertical domains. While generalist agents demonstrate proficiency in open-domain search, their reliability in capital-intensive and safety-critical sectors remains under-explored. FutureX-Pro targets four economically and socially pivotal verticals: Finance, Retail, Public Health, and Natural Disaster. We benchmark agentic Large Language Models (LLMs) on entry-level yet foundational prediction tasks -- ranging from forecasting market indicators and supply chain demands to tracking epidemic trends and natural disasters. By adapting the contamination-free, live-evaluation pipeline of FutureX, we assess whether current State-of-the-Art (SOTA) agentic LLMs possess the domain grounding necessary for industrial deployment. Our findings reveal the performance gap between generalist reasoning and the precision required for high-value vertical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。