用真实互动生成问题,评估多模态模型的协作能力。
"Is This It?": Towards Ecologically Valid Benchmarks for Situated Collaboration
- 用户在与智能系统实时交互中自然提出问题。
- 新基准问题更贴近真实场景,形式多样且内容复杂。
- 适合研究人机协作与真实环境下的模型评估者。
我们报告了构建生态有效基准的初步工作,用于评估大型多模态模型在情境化协作中的能力。与现有基准不同——后者通过模板、人工标注或大语言模型(LLMs)在预设数据集上生成问答对——我们提出并研究了一种由交互式系统驱动的方法:问题由用户在与端到端情境化AI系统的实际交互过程中自然生成。我们展示了这些生成问题在形式和内容上与传统具身问答(EQA)基准中的问题存在显著差异,并讨论了由此带来的新型现实挑战问题。
原文摘要 · Abstract (English)
We report initial work towards constructing ecologically valid benchmarks to assess the capabilities of large multimodal models for engaging in situated collaboration. In contrast to existing benchmarks, in which question-answer pairs are generated post hoc over preexisting or synthetic datasets via templates, human annotators, or large language models (LLMs), we propose and investigate an interactive system-driven approach, where the questions are generated by users in context, during their interactions with an end-to-end situated AI system. We illustrate how the questions that arise are different in form and content from questions typically found in existing embodied question answering (EQA) benchmarks and discuss new real-world challenge problems brought to the fore.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。