arXiv:2410.18001cs.AI2024-10

首个针对异常场景的模型评测数据集,覆盖多模态罕见任务。

Benchmarking Foundation Models on Exceptional Cases: Dataset Creation and Validation

  • 构建多模态异常案例数据集,涵盖漫画、书法等非常规输入。
  • 引入提示工程提升模型在异常任务上的表现,效果显著改善。
  • 适合研究模型泛化能力与鲁棒性的研究人员使用。

基础模型(FMs)在各类任务中取得显著进展,推动了推理能力评测基准的研究。然而,关于模型在异常场景下的表现仍缺乏研究,本文将其定义为分布外(OOD)推理任务。本文首次系统性地应对这一问题,构建了一个新颖的多模态评测数据集,涵盖图像小说、书法、新闻文章和歌词等类型,包含实例分类、角色识别、词元预测和文本生成等任务。同时提出链式思维(CoT)及CoT+少样本提示工程方法以提升性能。通过多种验证方法评估,结果显示模型表现得到明显改善。代码仓库已公开:https://github.com/MLAI-Yonsei/ExceptionalBenchmark。

原文摘要 · Abstract (English)

Foundation models (FMs) have achieved significant success across various tasks, leading to research on benchmarks for reasoning abilities. However, there is a lack of studies on FMs performance in exceptional scenarios, which we define as out-of-distribution (OOD) reasoning tasks. This paper is the first to address these cases, developing a novel dataset for evaluation of FMs across multiple modalities, including graphic novels, calligraphy, news articles, and lyrics. It includes tasks for instance classification, character recognition, token prediction, and text generation. The paper also proposes prompt engineering techniques like Chain-of-Thought (CoT) and CoT+Few-Shot to enhance performance. Validation of FMs using various methods revealed improvements. The code repository is accessible at: https://github.com/MLAI-Yonsei/ExceptionalBenchmark

基础模型异常检测多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。