arXiv:2605.00907cs.CVcs.AI2026-05

首个开放的交通多模态评测基准,专为大模型在交通场景中的可靠性评估而设计。

TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation

论文配图:TRIP-Evaluate: An Open Multimodal Benchmark for Evaluating Large Models in Transportation
图 1 · 摘自论文原文
  • 构建837个任务,覆盖车、路、人、规划四类交通功能,支持多模态诊断
  • 包含596文本、198图像、43点云数据,涵盖工程计算与规则推理等难点
  • 提供可复现、可诊断的评测标准,适合模型选型与安全部署验证

大语言模型和多模态大模型正被广泛应用于交通领域的法规问答、交通管理辅助、工程评审和自动驾驶场景理解。然而,交通工作流具有规则密集、计算密集、安全关键且天然多模态的特点。现有通用评测基准难以验证模型是否能正确应用法规、完成可验证的工程计算或可靠解析交通场景;而少数公开的交通评测基准范围狭窄,且极少支持对文本、图像和点云数据的细粒度故障诊断。为此,我们提出TRIP-Evaluate,一个面向交通领域大模型的开源多模态评测基准。该基准采用角色-任务-知识分类体系组织837个任务,覆盖车辆、交通管理、出行者及规划与设计功能。每个任务标注能力类型、模态类别和难度等级,支持从整体准确率到具体失效模式的逐层诊断。当前版本包含596个文本任务、198个图像任务和43个点云任务。TRIP-Evaluate还统一了任务构建、质量控制、提示设计、解码策略和评分标准,提升跨模型可比性。多模型实测结果表明,文本任务表现持续提升,但在多步工程计算、规则约束推理、多模态场景理解及点云解析方面仍存在显著短板。总体而言,TRIP-Evaluate为模型选型、回归测试和交通应用的安全部署提供了可复现、可诊断、工程对齐的评估基线。

原文摘要 · Abstract (English)

Large language models (LLMs) and multimodal large models (MLLMs) are increasingly used for transportation tasks such as regulation question answering, traffic management support, engineering review, and autonomous-driving scene reasoning. Yet transportation workflows are rule-intensive, computation-intensive, safety-critical, and inherently multimodal. Existing general benchmarks provide limited evidence of whether a model can apply regulations correctly, perform verifiable engineering calculations, or interpret traffic scenes reliably, while the small number of public transportation benchmarks remain narrow in scope and rarely support fine-grained diagnosis across text, images, and point-cloud data. To address this gap, we present TRIP-Evaluate, an open multimodal benchmark for large models in transportation. The benchmark organizes 837 items using a role-task-knowledge taxonomy that covers vehicle, traffic-management, traveler, and planning-and-design functions. Each item is annotated with capability, modality, and difficulty labels, enabling diagnosis from overall accuracy down to specific failure modes. The current release includes 596 text items, 198 image items, and 43 point-cloud items. TRIP-Evaluate also standardizes item construction, quality control, prompting, decoding, and scoring to improve cross-model comparability. Results on a diverse panel of models show that text-based performance is improving, but substantial weaknesses remain in multi-step engineering calculation, rule-constrained reasoning, multimodal scene understanding, and point-cloud understanding. Overall, TRIP-Evaluate provides a reproducible, diagnosable, and engineering-aligned evaluation baseline for model selection, regression testing, and safer deployment in transportation applications.

多模态评测交通大模型点云理解可诊断评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。