用大模型让摄像头听懂人话,边端部署还更准
Camera Control at the Edge with Language Models for Scene Understanding
- 用提示词生成关键词,小模型通过合成数据学大模型
- 比先进方法高35%性能,比Gemini Pro高20%准确率
- 无需编程,自然对话控制摄像头,适合智能安防场景
本文提出优化提示统一系统(OPUS),利用大语言模型(LLM)控制全景云台变焦(PTZ)摄像头,实现对自然环境的上下文理解。为提升成本效益,OPUS通过高层摄像头控制接口生成关键词,并在合成数据上使用监督微调(SFT)将大闭源模型知识迁移至小模型,实现高效边端部署,性能接近GPT-4。系统通过将多摄像头数据转为文本描述供语言模型处理,无需专用感知标记,增强环境感知能力。基准测试显示,该方法显著优于传统语言模型技术及复杂提示方法,较先进方案提升35%,任务准确率较Gemini Pro高20%。系统支持通过自然语言接口直观操控摄像头,无需显式编程,提供对话式交互方式,是用户控制与利用PTZ技术的重要进展。
原文摘要 · Abstract (English)
In this paper, we present Optimized Prompt-based Unified System (OPUS), a framework that utilizes a Large Language Model (LLM) to control Pan-Tilt-Zoom (PTZ) cameras, providing contextual understanding of natural environments. To achieve this goal, the OPUS system improves cost-effectiveness by generating keywords from a high-level camera control API and transferring knowledge from larger closed-source language models to smaller ones through Supervised Fine-Tuning (SFT) on synthetic data. This enables efficient edge deployment while maintaining performance comparable to larger models like GPT-4. OPUS enhances environmental awareness by converting data from multiple cameras into textual descriptions for language models, eliminating the need for specialized sensory tokens. In benchmark testing, our approach significantly outperformed both traditional language model techniques and more complex prompting methods, achieving a 35% improvement over advanced techniques and a 20% higher task accuracy compared to closed-source models like Gemini Pro. The system demonstrates OPUS's capability to simplify PTZ camera operations through an intuitive natural language interface. This approach eliminates the need for explicit programming and provides a conversational method for interacting with camera systems, representing a significant advancement in how users can control and utilize PTZ camera technology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。