测试手机端实时翻译性能,为移动端AI部署提供实证参考
An Empirical Study of On-Device Translation for Real-Time Live-Stream Chat on Mobile Devices
- 构建1000对韩英直播弹幕数据集,评估手机端模型表现
- 五款手机实测显示,优化后模型性能接近GPT-5.1
- 揭示设备资源与模型选型的权衡,适合移动AI开发者参考
尽管高效,但关于在真实场景中部署手机端AI模型的实际问题研究仍较少,如CPU占用率和发热情况。本文通过大量实验,探讨了在真实服务中部署手机端模型必须解决的两个关键问题:(i) 模型选择及其资源消耗,(ii) 手机端模型在领域适应方面的潜力。为此,我们聚焦于实时直播弹幕翻译任务,人工构建了包含1,000组韩英平行语句的LiveChatBench基准数据集。在五款移动设备上的实验表明,尽管服务大规模异构用户需谨慎考虑严苛的部署环境和模型选择,但所提方法在目标任务上仍能达到与GPT-5.1相当的性能。我们期望这些发现能为手机端AI社区提供有意义的洞见。
原文摘要 · Abstract (English)
Despite its efficiency, there has been little research on the practical aspects required for real-world deployment of on-device AI models, such as the device's CPU utilization and thermal conditions. In this paper, through extensive experiments, we investigate two key issues that must be addressed to deploy on-device models in real-world services: (i) the selection of on-device models and the resource consumption of each model, and (ii) the capability and potential of on-device models for domain adaptation. To this end, we focus on a task of translating live-stream chat messages and manually construct LiveChatBench, a benchmark consisting of 1,000 Korean-English parallel sentence pairs. Experiments on five mobile devices demonstrate that, although serving a large and heterogeneous user base requires careful consideration of highly constrained deployment settings and model selection, the proposed approach nevertheless achieves performance comparable to commercial models such as GPT-5.1 on the well-targeted task. We expect that our findings will provide meaningful insights to the on-device AI community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。