构建可实时交互的数字孪生人,提升人机协作中的信任与上下文理解。
A Human Digital Twin Architecture for Knowledge-based Interactions and Context-Aware Conversations
- 用大语言模型+元认知机制实现个性化、情境感知对话响应。
- 系统支持语音识别、情绪建模、唇形同步等多模态交互能力。
- 适用于训练、部署到复盘全周期的人机协同任务场景。
人工智能与机器学习的发展为任务、使命及持续协作中的智能体与人类协同(HAT)带来了新机遇。核心挑战在于如何使人类保持对自主资产的感知与控制,同时建立信任并支持共享上下文理解。为此,我们提出一种实时人类数字孪生(HDT)架构,集成大语言模型(LLMs)用于知识报告、问答与推荐,并通过可视化界面呈现。系统采用元认知方法,实现与人类队友预期一致的个性化、上下文感知响应。该HDT作为视觉与行为上高度逼真的团队成员,贯穿任务全生命周期——从训练、部署到事后复盘。架构包含语音识别、上下文处理、AI驱动对话、情绪建模、唇形同步与多模态反馈。本文描述了系统设计、性能指标及未来发展方向,旨在推动更自适应、更真实的HAT系统演进。
原文摘要 · Abstract (English)
Recent developments in Artificial Intelligence (AI) and Machine Learning (ML) are creating new opportunities for Human-Autonomy Teaming (HAT) in tasks, missions, and continuous coordinated activities. A major challenge is enabling humans to maintain awareness and control over autonomous assets, while also building trust and supporting shared contextual understanding. To address this, we present a real-time Human Digital Twin (HDT) architecture that integrates Large Language Models (LLMs) for knowledge reporting, answering, and recommendation, embodied in a visual interface. The system applies a metacognitive approach to enable personalized, context-aware responses aligned with the human teammate's expectations. The HDT acts as a visually and behaviorally realistic team member, integrated throughout the mission lifecycle, from training to deployment to after-action review. Our architecture includes speech recognition, context processing, AI-driven dialogue, emotion modeling, lip-syncing, and multimodal feedback. We describe the system design, performance metrics, and future development directions for more adaptive and realistic HAT systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。