arXiv:2409.19010cs.CLcs.AI2024-09

在手机端实现视觉问答、自动填表和多语言混用智能回复。

A comprehensive study of on-device NLP applications -- VQA, automated Form filling, Smart Replies for Linguistic Codeswitching

  • 基于屏幕理解,实现视觉问答与自动填表。
  • 支持多语言混用场景的智能回复,提升跨语言沟通效率。
  • 首个在设备端落地此类应用的研究,推动技术实用化。

大型语言模型的进展为设备端应用开启了新可能。本文提出三项新体验:基于屏幕理解的视觉问答、基于前一界面的自动表单填充,以及针对多语言混用场景的智能回复。代码切换指说话者在两种或以上语言间交替使用。据我们所知,这是首个将这些任务与解决方案应用于设备端应用的研究,旨在弥合前沿研究与实际应用之间的差距。

原文摘要 · Abstract (English)

Recent improvement in large language models, open doors for certain new experiences for on-device applications which were not possible before. In this work, we propose 3 such new experiences in 2 categories. First we discuss experiences which can be powered in screen understanding i.e. understanding whats on user screen namely - (1) visual question answering, and (2) automated form filling based on previous screen. The second category of experience which can be extended are smart replies to support for multilingual speakers with code-switching. Code-switching occurs when a speaker alternates between two or more languages. To the best of our knowledge, this is first such work to propose these tasks and solutions to each of them, to bridge the gap between latest research and real world impact of the research in on-device applications.

设备端NLP视觉问答多语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。