arXiv:2506.03761cs.CL2025-06被引 3

用7500+交互实例评测大模型当虚拟宠物的表现

Pet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network Services

  • 设计多维度任务评估大模型自演化与互动能力
  • 28个模型表现差异大,大小和能力影响显著
  • 适合研究情感化交互与虚拟宠物的开发者

随着对大语言模型在互动性与情感丰富体验中应用的兴趣增长,虚拟宠物陪伴成为一项新颖但尚未充分探索的应用。现有方法仅关注基础宠物角色扮演交互,缺乏对大模型综合陪伴能力的系统性评测。本文提出Pet-Bench,一个专为评估大模型在自我互动与人际互动维度上表现而设计的基准。不同于以往工作,Pet-Bench强调自我进化与成长行为,同时注重交互参与度,更真实地反映宠物陪伴特性。其包含智能日程安排、基于记忆的对话及心理对话等多样化任务,设计超过7,500个交互实例以模拟宠物行为。对28个大模型的评估显示性能差异显著,且与模型规模和内在能力相关,凸显该领域需专门优化。Pet-Bench为评测宠物相关大模型能力提供了基础资源,推动情感沉浸式人宠互动发展。

原文摘要 · Abstract (English)

As interest in using Large Language Models for interactive and emotionally rich experiences grows, virtual pet companionship emerges as a novel yet underexplored application. Existing approaches focus on basic pet role-playing interactions without systematically benchmarking LLMs for comprehensive companionship. In this paper, we introduce Pet-Bench, a dedicated benchmark that evaluates LLMs across both self-interaction and human-interaction dimensions. Unlike prior work, Pet-Bench emphasizes self-evolution and developmental behaviors alongside interactive engagement, offering a more realistic reflection of pet companionship. It features diverse tasks such as intelligent scheduling, memory-based dialogues, and psychological conversations, with over 7,500 interaction instances designed to simulate pet behaviors. Evaluation of 28 LLMs reveals significant performance variations linked to model size and inherent capabilities, underscoring the need for specialized optimization in this domain. Pet-Bench serves as a foundational resource for benchmarking pet-related LLM abilities and advancing emotionally immersive human-pet interactions.

虚拟宠物大模型评测情感交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。