arXiv:2501.15727cs.HCcs.AI2025-01被引 9

用大模型让普通人轻松自定义智能视觉传感器。

Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning

  • 通过自然语言描述任务,大模型自动生成并调试传感器规则。
  • 支持用户并行测试规则,发现隐藏缺陷和未考虑场景。
  • 适合想自定义家庭监控等个性化视觉系统的普通用户。

多模态大语言模型(MLLM)具备广泛世界知识和推理能力,使普通用户能够创建可自主推理复杂情境的个性化AI传感器。用户可用自然语言描述任务(如“当我的幼儿开始调皮时报警”),由模型实时分析摄像头画面并响应。在前期研究中,我们发现用户虽高度认可自定义传感器的价值,但难以准确表达需求且仅靠提示难以调试。为此,我们开发了Gensors系统,利用MLLM的推理能力支持用户定义定制化传感器。该系统具备四项功能:1)通过自动生成与手动创建相结合的方式辅助用户明确需求;2)允许用户并行隔离和测试单个规则以实现高效调试;3)根据用户提供的图像建议补充规则;4)提出测试案例帮助用户对潜在未预见场景进行“压力测试”。用户研究表明,使用Gensors后,参与者在控制感、理解度和沟通便捷性方面均有显著提升。该系统不仅缓解了模型局限性,还帮助用户发现需求盲点、揭示意外失效模式。最后,我们讨论了MLLM特性(如幻觉和响应不一致)对传感器创建过程的影响。这些发现为未来面向普通用户的直观、可定制智能感知系统设计提供了重要参考。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs), with their expansive world knowledge and reasoning capabilities, present a unique opportunity for end-users to create personalized AI sensors capable of reasoning about complex situations. A user could describe a desired sensing task in natural language (e.g., "alert if my toddler is getting into mischief"), with the MLLM analyzing the camera feed and responding within seconds. In a formative study, we found that users saw substantial value in defining their own sensors, yet struggled to articulate their unique personal requirements and debug the sensors through prompting alone. To address these challenges, we developed Gensors, a system that empowers users to define customized sensors supported by the reasoning capabilities of MLLMs. Gensors 1) assists users in eliciting requirements through both automatically-generated and manually created sensor criteria, 2) facilitates debugging by allowing users to isolate and test individual criteria in parallel, 3) suggests additional criteria based on user-provided images, and 4) proposes test cases to help users "stress test" sensors on potentially unforeseen scenarios. In a user study, participants reported significantly greater sense of control, understanding, and ease of communication when defining sensors using Gensors. Beyond addressing model limitations, Gensors supported users in debugging, eliciting requirements, and expressing unique personal requirements to the sensor through criteria-based reasoning; it also helped uncover users' "blind spots" by exposing overlooked criteria and revealing unanticipated failure modes. Finally, we discuss how unique characteristics of MLLMs--such as hallucinations and inconsistent responses--can impact the sensor-creation process. These findings contribute to the design of future intelligent sensing systems that are intuitive and customizable by everyday users.

多模态模型智能传感用户交互个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。