AI眼镜通过双代理架构实现实时语音处理与跨平台任务执行。
An Intelligent AI glasses System with Multi-Agent Architecture for Real-Time Voice Processing and Task Execution
- 双代理设计:一个负责语音识别,一个运行本地大模型处理指令。
- 支持多语言语音命令,可实时传输音视频与眼动数据。
- 适合开发者构建智能眼镜应用,尤其关注隐私与低延迟场景。
本文提出一种集成实时语音处理、人工智能代理和跨网络流媒体能力的AI眼镜系统。系统采用双代理架构,Agent 01 负责自动语音识别(ASR),Agent 02 通过本地大型语言模型(LLMs)、模型上下文协议(MCP)工具和检索增强生成(RAG)进行AI处理。系统支持实时RTSP音视频流传输、眼动追踪数据采集,以及通过RabbitMQ消息队列实现远程任务执行。实验表明,该系统成功实现了多语言语音指令处理与跨平台任务执行能力。
原文摘要 · Abstract (English)
This paper presents an AI glasses system that integrates real-time voice processing, artificial intelligence(AI) agents, and cross-network streaming capabilities. The system employs dual-agent architecture where Agent 01 handles Automatic Speech Recognition (ASR) and Agent 02 manages AI processing through local Large Language Models (LLMs), Model Context Protocol (MCP) tools, and Retrieval-Augmented Generation (RAG). The system supports real-time RTSP streaming for voice and video data transmission, eye tracking data collection, and remote task execution through RabbitMQ messaging. Implementation demonstrates successful voice command processing with multilingual support and cross-platform task execution capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。