arXiv:2510.08576cs.SEcs.AI2025-10中稿 · First Internationa…

对比开源大模型与GPT-4在理解用户意图并生成操作流程上的表现。

Comparative Analysis of Large Language Models for the Machine-Assisted Resolution of User Intentions

  • 用开源模型替代云端大模型,实现本地化意图解析。
  • 部分开源模型已接近GPT-4在流程生成任务中的表现。
  • 适合关注隐私、自主性和系统可控性的开发者与研究者。

大语言模型(LLMs)已成为自然语言理解与用户意图解析的变革性工具,支持翻译、摘要等任务,并逐步实现复杂工作流的编排。这一发展标志着从传统图形界面转向以语言为核心的交互范式:用户无需手动操作应用,只需用自然语言描述目标,由大模型动态协调多应用执行。然而,现有实现多依赖云上私有模型,带来隐私、自主与可扩展性限制。为使语言驱动交互成为可信可靠的未来操作系统基础,本地部署不可或缺。本研究评估多个开源开放模型在机器辅助用户意图解析中的能力,与OpenAI的GPT-4系统进行对比,分析其在生成各类用户意图对应工作流时的表现。研究提供了关于开源模型实用可行性、性能权衡及作为下一代操作系统中自治本地组件潜力的实证见解。结果推动了对人工智能基础设施去中心化与民主化的讨论,指向一个通过嵌入本地智能实现更无缝、自适应、隐私保护的用户-设备交互未来。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have emerged as transformative tools for natural language understanding and user intent resolution, enabling tasks such as translation, summarization, and, increasingly, the orchestration of complex workflows. This development signifies a paradigm shift from conventional, GUI-driven user interfaces toward intuitive, language-first interaction paradigms. Rather than manually navigating applications, users can articulate their objectives in natural language, enabling LLMs to orchestrate actions across multiple applications in a dynamic and contextual manner. However, extant implementations frequently rely on cloud-based proprietary models, which introduce limitations in terms of privacy, autonomy, and scalability. For language-first interaction to become a truly robust and trusted interface paradigm, local deployment is not merely a convenience; it is an imperative. This limitation underscores the importance of evaluating the feasibility of locally deployable, open-source, and open-access LLMs as foundational components for future intent-based operating systems. In this study, we examine the capabilities of several open-source and open-access models in facilitating user intention resolution through machine assistance. A comparative analysis is conducted against OpenAI's proprietary GPT-4-based systems to assess performance in generating workflows for various user intentions. The present study offers empirical insights into the practical viability, performance trade-offs, and potential of open LLMs as autonomous, locally operable components in next-generation operating systems. The results of this study inform the broader discussion on the decentralization and democratization of AI infrastructure and point toward a future where user-device interaction becomes more seamless, adaptive, and privacy-conscious through locally embedded intelligence.

大模型意图解析本地部署开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。