首个面向真实商业创意的多模态评测基准,提升AI生成创意的实用性和多样性。
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation

- 构建包含3万条数据的多模态基准,融合图文信息支持更真实的商业创意生成。
- 在6个维度上测试,MBA-k模型比纯文本基线提升77.1%的创意质量。
- 适用于需要视觉理解能力的商业智能、创业辅助等场景,适合研究多模态Agent的开发者。
由大语言模型驱动的智能体系统为商业创意生成带来了新机遇,但现有方法仍局限于纯文本范式,难以捕捉真实场景中固有的多模态特性。为此,我们提出MBA-Bench——首个用于训练与评估商业创意智能体的多模态基准,涵盖六个领域共3万样本,每个领域包含仅靠文字无法完整表达的视觉特征。具体而言,我们自动为图像生成标题,并利用GPT-4o通过检索查询生成、市场证据检索与证据增强合成,为每三个商业问题生成五个参考创意。沿用以往工作,采用多模态大模型作为裁判,在六个商业导向标准下评估智能体表现。针对评判标准是否公开的两种场景,我们设计MBA-b(盲评)和MBA-k(已知)两种版本。两者均采用两种新颖的奖励目标进行训练:创造力与可行性;MBA-k进一步优化六个公开标准,共八个目标。训练过程基于LoRA的监督微调,随后采用组相对策略优化。在MBA-Bench上的广泛实验表明,我们设置的两个基线分别支持仅文本或多模态输入,后者接近闭源模型性能。MBA-b与MBA-k分别比文本基线提升63.9%与77.1%,比多模态基线提升25.6%与35.8%。
原文摘要 · Abstract (English)
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently multimodal nature of real-world contexts. We thus introduce MBA-Bench, the first multimodal benchmark for training and evaluating business ideation agents, comprising 30K samples across six domains, each domain characterized by distinct visual cues not fully conveyed by text alone. Concretely, we automatically caption images and employ GPT-4o to generate five reference ideas for each of three business questions through retrieval query generation, market evidence retrieval, and evidence-augmented synthesis. Following prior work, we evaluate agents across six business-oriented criteria using MLLM-as-a-Judge. To consider settings where criteria are hidden or disclosed, we present MBA-b and MBA-k for blind and known, respectively. We train both with two novel reward objectives---creativity and feasibility---while MBA-k further optimizes the six disclosed criteria for eight in total. Both are trained via LoRA-based supervised fine-tuning followed by group relative policy optimization with these setting-specific rewards. For extensive experiments on MBA-Bench, we set up two baselines accommodating either captions only or multimodal inputs, with the latter nearing closed-source performance on several metrics. MBA-b and MBA-k outperform caption baselines by 63.9% and 77.1%, and multimodal baselines by 25.6% and 35.8%, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。