用真实加密市场动态测试大模型智能体,发现多数模型会查数据却不会预测。
CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
- 每月更新50道由专业人士设计的实操题,覆盖查数据与做判断
- 10个大模型评估中,多数在预测任务上表现远差于检索能力
- 专为高时效、强对抗的加密世界打造,适合研究智能体真实分析力
本文提出CryptoBench,首个由专业人员构建、动态更新的基准测试,用于严格评估大语言模型(LLM)智能体在快节奏、高对抗性加密货币环境中的真实能力。与通用搜索与预测类基准不同,专业加密分析面临极端时间敏感性、信息环境高度欺骗性,以及需整合链上平台、实时去中心化金融(DeFi)仪表盘等多元专业数据源的挑战。为此,我们构建了一个实时动态基准,每月发布50道题目,由精通加密领域的专业人士设计,模拟真实分析师工作流程。任务按四象限系统分类:简单检索、复杂检索、简单预测、复杂预测,实现对智能体基础数据获取与高级分析预测能力的精准评估。对10个大模型的直接调用与智能体框架下的评估揭示性能层级,并暴露一种‘检索-预测失衡’现象:许多领先模型虽擅长数据检索,但在需要综合推理的预测任务中表现显著不足,反映出智能体可能看似事实准确,实则缺乏深层分析能力。
原文摘要 · Abstract (English)
This paper introduces CryptoBench, the first expert-curated, dynamic benchmark designed to rigorously evaluate the real-world capabilities of Large Language Model (LLM) agents in the uniquely demanding and fast-paced cryptocurrency domain. Unlike general-purpose agent benchmarks for search and prediction, professional crypto analysis presents specific challenges: \emph{extreme time-sensitivity}, \emph{a highly adversarial information environment}, and the critical need to synthesize data from \emph{diverse, specialized sources}, such as on-chain intelligence platforms and real-time Decentralized Finance (DeFi) dashboards. CryptoBench thus serves as a much more challenging and valuable scenario for LLM agent assessment. To address these challenges, we constructed a live, dynamic benchmark featuring 50 questions per month, expertly designed by crypto-native professionals to mirror actual analyst workflows. These tasks are rigorously categorized within a four-quadrant system: Simple Retrieval, Complex Retrieval, Simple Prediction, and Complex Prediction. This granular categorization enables a precise assessment of an LLM agent's foundational data-gathering capabilities alongside its advanced analytical and forecasting skills. Our evaluation of ten LLMs, both directly and within an agentic framework, reveals a performance hierarchy and uncovers a failure mode. We observe a \textit{retrieval-prediction imbalance}, where many leading models, despite being proficient at data retrieval, demonstrate a pronounced weakness in tasks requiring predictive analysis. This highlights a problematic tendency for agents to appear factually grounded while lacking the deeper analytical capabilities to synthesize information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。