arXiv:2601.13801cs.RO2026-01中稿 · publication at LBR…被引 1

让无人机能看会说,实时回应人类指令和情绪。

HoverAI: An Embodied Aerial Agent for Natural Human-Drone Interaction

  • 用视觉语音识别+自适应对话系统实现自然交互。
  • 命令识别准确率90%,年龄估计误差仅5.14年。
  • 适合智能导览、人机协作等需要情感响应的场景。

在有人类活动空间中运行的无人机因通信机制不足而难以明确表达意图。本文提出HoverAI,一种集飞行能力、无基础设施视觉投影与实时对话AI于一体的具身空中代理。该系统配备微机电激光投影仪、机载半刚性屏幕及RGB摄像头,通过视觉与语音感知用户,并以唇同步动画角色进行响应,其外观会根据用户人口统计特征动态调整。系统采用多模态流程:结合语音活动检测(VAD)、自动语音识别(Whisper)、基于大模型的意图分类、检索增强生成(RAG)对话、人脸分析个性化处理以及语音合成(XTTS v2)。评估显示,系统在指令识别(F1: 0.90)、性别识别(F1: 0.89)、年龄估计(MAE: 5.14年)和语音转写(WER: 0.181)方面表现优异。通过融合空中机器人、自适应对话AI与自持视觉输出,HoverAI开创了一类空间感知、社会响应的新型具身智能体,适用于引导、辅助与以人为中心的交互应用。

原文摘要 · Abstract (English)

Drones operating in human-occupied spaces suffer from insufficient communication mechanisms that create uncertainty about their intentions. We present HoverAI, an embodied aerial agent that integrates drone mobility, infrastructure-independent visual projection, and real-time conversational AI into a unified platform. Equipped with a MEMS laser projector, onboard semi-rigid screen, and RGB camera, HoverAI perceives users through vision and voice, responding via lip-synced avatars that adapt appearance to user demographics. The system employs a multimodal pipeline combining VAD, ASR (Whisper), LLM-based intent classification, RAG for dialogue, face analysis for personalization, and voice synthesis (XTTS v2). Evaluation demonstrates high accuracy in command recognition (F1: 0.90), demographic estimation (gender F1: 0.89, age MAE: 5.14 years), and speech transcription (WER: 0.181). By uniting aerial robotics with adaptive conversational AI and self-contained visual output, HoverAI introduces a new class of spatially-aware, socially responsive embodied agents for applications in guidance, assistance, and human-centered interaction.

无人机交互对话系统具身智能多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。