用实时预测市场评估大模型,看它像不像人类在不确定性中做决策。
TruthTensor: Evaluating LLMs through Human Imitation on Prediction Market under Drift and Holistic Reasoning
- 把大模型放进建立在真实市场的动态环境中,模拟人类决策过程。
- 500多个真实市场测试显示,准确率相似的模型在校准和风险敏感度上差异显著。
- 适合关注模型真实世界表现、需可复现评估的AI研究者与应用开发者。
评估语言模型和AI代理仍面临根本性挑战,因静态基准无法捕捉现实世界的不确定性、分布漂移,以及孤立任务准确率与人类对齐决策之间的差距。本文提出TruthTensor,一种新颖且可复现的评估范式,将推理模型不仅视为预测引擎,更视为在社会性高熵环境中运行的人类模仿系统。基于前瞻性的无污染任务,该框架以实时预测市场为锚点,结合概率评分,提供模型行为的全景视图。TruthTensor在500多个真实市场(政治、经济、文化、技术)中验证,证明即使预测准确率相近,模型在校准、漂移和风险敏感度上仍存在显著差异,凸显需多维度评估(准确率、校准、叙事稳定性、成本与资源效率)。该框架整合现代评估最佳实践:明确人机角色分工、标注协议、统计检验流程,实现结果可解释与可复现。包含清晰假设设定、谨慎度量选择、透明计算/成本报告、人机协同验证及开源版本化评估合约,确保在真实决策场景中对大模型的可信评估。项目已公开发布于https://truthtensor.com。
原文摘要 · Abstract (English)
Evaluating language models and AI agents remains fundamentally challenging because static benchmarks fail to capture real-world uncertainty, distribution shift, and the gap between isolated task accuracy and human-aligned decision-making under evolving conditions. This paper introduces TruthTensor, a novel, reproducible evaluation paradigm that measures reasoning models not only as prediction engines but as human-imitation systems operating in socially-grounded, high-entropy environments. Building on forward-looking, contamination-free tasks, our framework anchors evaluation to live prediction markets and combines probabilistic scoring to provide a holistic view of model behavior. TruthTensor complements traditional correctness metrics with drift-centric diagnostics and explicit robustness checks for reproducibility. It specify human vs. automated evaluation roles, annotation protocols, and statistical testing procedures to ensure interpretability and replicability of results. In experiments across 500+ real markets (political, economic, cultural, technological), TruthTensor demonstrates that models with similar forecast accuracy can diverge markedly in calibration, drift, and risk-sensitivity, underscoring the need to evaluate models along multiple axes (accuracy, calibration, narrative stability, cost, and resource efficiency). TruthTensor therefore operationalizes modern evaluation best practices, clear hypothesis framing, careful metric selection, transparent compute/cost reporting, human-in-the-loop validation, and open, versioned evaluation contracts, to produce defensible assessments of LLMs in real-world decision contexts. We publicly released TruthTensor at https://truthtensor.com.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。