针对餐饮零售场景打造更精准的多模态模型,解决数据差、评测难问题。
Ostrakon-VL: Towards Domain-Expert MLLM for Food-Service and Retail Stores
- 基于Qwen3-VL-8B构建专用模型,引入自动数据清洗流程提升数据质量。
- 在首个公开基准ShopBench上达60.1分,超越同规模模型4.8分。
- 适合餐饮零售领域研究者,推动可复现的行业模型研发。
多模态大语言模型在通用感知与推理方面取得显著进展,但在餐饮零售服务(FSRS)场景中仍面临两大挑战:一是来自异构设备采集的现实数据噪声大、缺乏可审计的闭环数据治理,难以构建高质量、可控且可复现的训练语料;二是现有评估协议缺少覆盖单图、多图及视频输入的统一、细粒度、标准化基准,难以客观衡量模型鲁棒性。为此,我们首先开发了面向FSRS的Ostrakon-VL模型,基于Qwen3-VL-8B架构;其次提出首个公开的FSRS基准测试集ShopBench;第三设计了多阶段的高质量无偏自动化数据清洗方法QUAD。通过多阶段训练策略,Ostrakon-VL在ShopBench上实现平均60.1分,成为参数量相当的开源多模态模型中的新标杆,优于更大模型Qwen3-VL-235B-A22B(59.4)0.7分,也超过同规模的Qwen3-VL-8B(55.3)4.8分,展现显著的参数效率优势。结果表明,Ostrakon-VL具备更强的领域感知与决策能力。为促进可复现研究,我们将公开发布Ostrakon-VL与ShopBench。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have recently achieved substantial progress in general-purpose perception and reasoning. Nevertheless, their deployment in Food-Service and Retail Stores (FSRS) scenarios encounters two major obstacles: (i) real-world FSRS data, collected from heterogeneous acquisition devices, are highly noisy and lack auditable, closed-loop data curation, which impedes the construction of high-quality, controllable, and reproducible training corpora; and (ii) existing evaluation protocols do not offer a unified, fine-grained and standardized benchmark spanning single-image, multi-image, and video inputs, making it challenging to objectively gauge model robustness. To address these challenges, we first develop Ostrakon-VL, an FSRS-oriented MLLM based on Qwen3-VL-8B. Second, we introduce ShopBench, the first public benchmark for FSRS. Third, we propose QUAD (Quality-aware Unbiased Automated Data-curation), a multi-stage multimodal instruction data curation pipeline. Leveraging a multi-stage training strategy, Ostrakon-VL achieves an average score of 60.1 on ShopBench, establishing a new state of the art among open-source MLLMs with comparable parameter scales and diverse architectures. Notably, it surpasses the substantially larger Qwen3-VL-235B-A22B (59.4) by +0.7, and exceeds the same-scale Qwen3-VL-8B (55.3) by +4.8, demonstrating significantly improved parameter efficiency. These results indicate that Ostrakon-VL delivers more robust and reliable FSRS-centric perception and decision-making capabilities. To facilitate reproducible research, we will publicly release Ostrakon-VL and the ShopBench benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。