arXiv:2601.05851cs.CLcs.AI2026-01Conference of the …

让聊天助手根据图文上下文动态选择模型,更快更准地补全用户输入。

Router-Suggest: Dynamic Routing for Multimodal Auto-Completion in Visually-Grounded Dialogs

  • 根据对话上下文动态切换文本模型与视觉语言模型
  • 相比最优视觉语言模型提速2.3至10倍,保持高准确率
  • 适合资源受限场景,提升多轮对话中用户输入效率

实时多模态自动补全对数字助手、聊天机器人、设计工具和医疗咨询至关重要,其依赖共享视觉上下文。我们提出多模态自动补全(MAC)任务,利用部分输入文本和视觉线索预测聊天中的下一个字符。不同于传统文本补全(TAC),MAC基于多模态上下文捕捉用户意图。为此,我们改造MMDialog和ImageChat构建基准数据集。评估主流视觉语言模型(VLMs)与强文本基线,揭示精度与效率的权衡。提出Router-Suggest框架,根据对话上下文动态选择文本模型或VLM;并提供轻量级变体适配资源受限环境。该框架相较最优VLM实现2.3至10倍加速。用户研究显示,VLM显著优于文本模型,在用户满意度、减少输入负担和提升补全质量方面表现突出,尤其在多轮对话中。结果表明,多模态上下文对自动补全是必要且关键的,有助于打造更智能、更懂用户的助手。

原文摘要 · Abstract (English)

Real-time multimodal auto-completion is essential for digital assistants, chatbots, design tools, and healthcare consultations, where user inputs rely on shared visual context. We introduce Multimodal Auto-Completion (MAC), a task that predicts upcoming characters in live chats using partially typed text and visual cues. Unlike traditional text-only auto-completion (TAC), MAC grounds predictions in multimodal context to better capture user intent. To enable this task, we adapt MMDialog and ImageChat to create benchmark datasets. We evaluate leading vision-language models (VLMs) against strong textual baselines, highlighting trade-offs in accuracy and efficiency. We present Router-Suggest, a router framework that dynamically selects between textual models and VLMs based on dialog context, along with a lightweight variant for resource-constrained environments. Router-Suggest achieves a 2.3x to 10x speedup over the best-performing VLM. A user study shows that VLMs significantly excel over textual models on user satisfaction, notably saving user typing effort and improving the quality of completions in multi-turn conversations. These findings underscore the need for multimodal context in auto-completions, leading to smarter, user-aware assistants.

多模态自动补全视觉语言模型动态路由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。