arXiv:2512.17436cs.CV2025-12被引 13

小米开源家庭场景视觉语言模型,兼顾智能家居理解与通用多模态推理。

Xiaomi MiMo-VL-Miloco Technical Report

  • 分两阶段训练:监督微调+基于组相对策略优化的强化学习
  • 在手势识别和家庭场景理解上达领先F1分数,视频/语言基准全面超越基线
  • 模型专精家庭任务同时提升文本推理,适合智能家电研发与部署

我们开源了MiMo-VL-Miloco-7B及其量化版本MiMo-VL-Miloco-7B-GGUF,一对面向家庭场景的视觉语言模型,在家庭场景理解和通用多模态推理方面表现优异。基于MiMo-VL-7B架构,该模型针对智能家居环境优化,在手势识别和常见家庭场景理解任务上取得领先F1得分,并在Video-MME、Video-MMMU、Charades-STA等视频基准以及MMMU-Pro、MMLU-Pro等语言理解基准上持续获得提升。实验表明,该模型在家庭场景理解及多个多模态推理基准上优于强闭源与开源自研基线。为平衡专精与通用性,我们设计了两阶段训练流程,结合监督微调与基于组相对策略优化(Group Relative Policy Optimization)的强化学习,利用高效多领域数据;进一步引入思维链监督与令牌预算感知推理,实现数据高效知识学习与高效推理。分析显示,针对性家庭场景训练不仅增强活动与手势理解,还小幅提升纯文本推理能力,仅对文档中心任务有轻微影响。模型检查点、量化GGUF权重及家庭场景评估工具包已公开于https://github.com/XiaoMi/xiaomi-mimo-vl-miloco,支持真实智能家庭应用的研究与部署。

原文摘要 · Abstract (English)

We open-source MiMo-VL-Miloco-7B and its quantized variant MiMo-VL-Miloco-7B-GGUF, a pair of home-centric vision-language models that achieve strong performance on both home-scenario understanding and general multimodal reasoning. Built on the MiMo-VL-7B backbone, MiMo-VL-Miloco-7B is specialized for smart-home environments, attaining leading F1 scores on gesture recognition and common home-scenario understanding, while also delivering consistent gains across video benchmarks such as Video-MME, Video-MMMU, and Charades-STA, as well as language understanding benchmarks including MMMU-Pro and MMLU-Pro. In our experiments, MiMo-VL-Miloco-7B outperforms strong closed-source and open-source baselines on home-scenario understanding and several multimodal reasoning benchmarks. To balance specialization and generality, we design a two-stage training pipeline that combines supervised fine-tuning with reinforcement learning based on Group Relative Policy Optimization, leveraging efficient multi-domain data. We further incorporate chain-of-thought supervision and token-budget-aware reasoning, enabling the model to learn knowledge in a data-efficient manner while also performing reasoning efficiently. Our analysis shows that targeted home-scenario training not only enhances activity and gesture understanding, but also improves text-only reasoning with only modest trade-offs on document-centric tasks. Model checkpoints, quantized GGUF weights, and our home-scenario evaluation toolkit are publicly available at https://github.com/XiaoMi/xiaomi-mimo-vl-miloco to support research and deployment in real-world smart-home applications.

视觉语言模型智能家庭多模态推理模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。