模型越大越聪明?实验发现短时人机互动中大模型效果不升反降。
Diminishing Returns of Intelligence: The Non-Linear Relationship Between LLM Scale and User Perception in Short-Duration Open-Ended Social Human-Robot Interactions

- 用4B、8B、30B参数的Qwen3-VL模型驱动机器人对话
- 用户对30B模型在智能感和自然度上无显著偏好提升
- 经验多的用户更易察觉模型差异,适合研究人机交互体验
大型语言模型(LLMs)正被广泛用于驱动具身社交代理,但其规模是否能提升用户在短暂人机互动中的感知体验仍不明确。本研究通过受试者内实验,让19名参与者与搭载Qwen3-VL模型(4B、8B、30B参数)的机器人面孔进行短时开放性社交互动,评估其感知智能、自然度、愉悦感与幽默感。结果表明,30B模型在整体上未显著优于小规模模型,尤其在自然度和智能感方面,与4B模型相比无明显优势。同时,高频率使用AI的用户对30B模型的智能评分更高,显示经验可能影响感知差异。总体而言,在简短开放性人机交互中,模型规模带来的收益递减,对话流畅性、响应速度与社交适切性可能比参数量更重要。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly used to drive embodied social agents, yet it remains unclear whether larger models improve user perception during brief human-robot encounters. This paper examines the effect of LLM parameter size on short-duration, open-ended social interactions with a robot interface. In a within-subjects study, 19 participants interacted with robot faces driven by Qwen3-VL models at 4B, 8B, and 30B parameters. Participants evaluated the interactions in terms of perceived intelligence, naturalness, enjoyment, and humor. Results showed no significant overall preference for the 30B model over the smaller variants, including no significant advantage over the 4B model in perceived naturalness or intelligence. A significant relationship between AI interaction frequency and intelligence rankings for the 30B model suggests that more experienced users may be more sensitive to differences in model capability. Overall, the findings indicate diminishing returns from model scaling in brief open-ended social HRI, where conversational flow, responsiveness, and socially appropriate behavior potentially matter as much as raw parameter count.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。