构建首个面向语义查询引擎的多模态基准测试
SemBench: A Benchmark for Semantic Query Processing Engines
- 用自然语言配置语义操作符,通过大模型执行跨模态查询
- 覆盖电影评论分析到车辆损伤检测等10类场景与多模态数据
- 评测学术与工业系统表现,揭示当前大模型查询引擎短板
我们提出一个针对新型系统——语义查询处理引擎的基准测试。这类系统依赖于前沿大语言模型(LLMs)的生成与推理能力,将SQL扩展为可通过自然语言指令配置的语义操作符,利用LLMs执行并支持对多模态数据的各种操作。本基准测试在三个关键维度上体现多样性:场景、模态与操作符。涵盖从电影评论分析到汽车损伤检测等多样化应用场景,涉及图像、音频和文本等多种数据模态。查询包含语义过滤、连接、映射、排序与分类等多种操作符。我们在三个学术系统(LOTUS、Palimpzest、ThalamusDB)和一个工业系统(Google BigQuery)上评估了该基准。尽管结果反映的是持续开发中的系统快照,但仍提供了关于其当前优劣势的关键洞察,为未来研究指明了方向。
原文摘要 · Abstract (English)
We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on generative and reasoning capabilities of state-of-the-art large language models (LLMs). They extend SQL with semantic operators, configured by natural language instructions, that are evaluated via LLMs and enable users to perform various operations on multimodal data. Our benchmark introduces diversity across three key dimensions: scenarios, modalities, and operators. Included are scenarios ranging from movie review analysis to car damage detection. Within these scenarios, we cover different data modalities, including images, audio, and text. Finally, the queries involve a diverse set of operators, including semantic filters, joins, mappings, ranking, and classification operators. We evaluated our benchmark on three academic systems (LOTUS, Palimpzest, and ThalamusDB) and one industrial system, Google BigQuery. Although these results reflect a snapshot of systems under continuous development, our study offers crucial insights into their current strengths and weaknesses, illuminating promising directions for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。