arXiv:2410.20116cs.HCcs.AI2024-10被引 3

打造多模态低延迟社交智能体框架,让实时互动更高效

Estuary: A Framework For Building Multimodal Low-Latency Real-Time Socially Interactive Agents

  • 构建模块化多模态框架,支持文本、音频等输入
  • 实现离云运行,响应速度更快且可复现
  • 适合研究实时社交智能体的开发者和学者

生成式人工智能技术的发展推动了社交互动智能体(SIAs)的研究。尽管当前已有诸多先进AI组件用于实时SIA研究,但缺乏统一标准框架导致开发重复投入。为此,我们提出Estuary:一个支持文本、音频及未来视频的多模态框架,旨在构建低延迟、实时响应的社交智能体。该框架通过模块化与互操作架构,无缝集成现有与未来组件,支持完全离云运行,显著提升实验可配置性、可控性、可复现性及响应速度,有效减少研究间重复工作。

原文摘要 · Abstract (English)

The rise in capability and ubiquity of generative artificial intelligence (AI) technologies has enabled its application to the field of Socially Interactive Agents (SIAs). Despite rising interest in modern AI-powered components used for real-time SIA research, substantial friction remains due to the absence of a standardized and universal SIA framework. To target this absence, we developed Estuary: a multimodal (text, audio, and soon video) framework which facilitates the development of low-latency, real-time SIAs. Estuary seeks to reduce repeat work between studies and to provide a flexible platform that can be run entirely off-cloud to maximize configurability, controllability, reproducibility of studies, and speed of agent response times. We are able to do this by constructing a robust multimodal framework which incorporates current and future components seamlessly into a modular and interoperable architecture.

社交智能体多模态低延迟框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。