arXiv:2510.00481cs.NIcs.AI2025-10

首次系统评测主流AI视频聊天性能,揭示体验关键因素。

Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps

  • 构建四维基准测试框架,覆盖质量、延迟、机制与开销。
  • 发现AI聊天中网络延迟影响小于真人通话,模型能力决定体验。
  • 适合研究AI交互、实时通信及用户体验优化的开发者与学者。

2025年,大型语言模型(LLM)服务推出了新功能——AI视频聊天,使用户可通过实时通信(RTC)与AI代理进行类真人对话。尽管该功能意义重大,但尚无系统性研究对其性能进行刻画。为此,本文提出一个涵盖质量、延迟、内部机制和系统开销四个维度的综合基准,并利用自研测试平台对六款主流AI视频聊天机器人进行了评估。同时,搭建了在线用户研究平台。测量结果揭示若干有趣发现:在AI视频聊天中,网络延迟的影响不如真人通话显著;模型能力是影响用户体验的核心因素。该基准还为未来优化提出了多个值得深入的研究问题。数据集与测试平台已开源,可访问 https://callarena.net/ 获取。

原文摘要 · Abstract (English)

In 2025, Large Language Model (LLM) services have launched a new feature -- AI video chat -- allowing users to interact with AI agents via real-time video communication (RTC), just like chatting with real people. Despite its significance, no systematic study has characterized the performance of existing AI video chat systems. To address this gap, this paper proposes a comprehensive benchmark across four dimensions: quality, latency, internal mechanisms, and system overhead. Using custom testbeds, we further evaluate six mainstream AI video chatbots with this benchmark. We also build an online platform for user study. The measurement leads to interesting findings that could be beneficial to the future optimizations. For example, the network latency of AI video chat matters not as much as human video chat. The capabilities of AI agents matters most in the user experience. Our benchmarking results also open up several research questions for future optimizations of AI video chatbots. Availability: https://callarena.net/ for the online evaluation platform and our open-sourced dataset and testbed.

AI视频用户体验实时通信评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。