首个可组合AI基准测试,评估模型协作解决复杂任务的能力。
CABENCH: Benchmarking Composable AI for Solving Complex Tasks through Composing Ready-to-Use Models
- 构建70个真实场景任务与700个可用模型的组合评估体系。
- 人类设计解法平均性能优于两个大模型方案,但仍有提升空间。
- 适合研究自动任务编排、多模型协作的AI开发者参考。
可组合AI通过将复杂任务分解为子任务,并利用现成的训练模型逐一解决,提供了一种可扩展且高效的方法。然而,该范式下的系统性评估仍基本空白。本文提出CABENCH,首个公开的基准,包含70个真实可组合AI任务及跨模态、跨领域的700个模型库,并设计端到端评估框架。为建立基线,我们提供人工设计的参考解决方案,并与两种基于大模型的方法进行对比。结果表明,可组合AI在应对现实复杂问题上具有潜力,但也凸显出需开发能自动生成高效执行流程的方法以充分释放其潜能。
原文摘要 · Abstract (English)
Composable AI offers a scalable and effective paradigm for tackling complex AI tasks by decomposing them into sub-tasks and solving each sub-task using ready-to-use well-trained models. However, systematically evaluating methods under this setting remains largely unexplored. In this paper, we introduce CABENCH, the first public benchmark comprising 70 realistic composable AI tasks, along with a curated pool of 700 models across multiple modalities and domains. We also propose an evaluation framework to enable end-to-end assessment of composable AI solutions. To establish initial baselines, we provide human-designed reference solutions and compare their performance with two LLM-based approaches. Our results illustrate the promise of composable AI in addressing complex real-world problems while highlighting the need for methods that can fully unlock its potential by automatically generating effective execution pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。