arXiv:2512.23217cs.AI2025-12

用人体舒适度评估AI的认知能力,发现当前大模型有基本推理但缺精准因果理解。

TCEval: Using Thermal Comfort to Assess Cognitive and Perceptual Abilities of AI

  • 设计虚拟人格的AI代理,通过穿衣选择和舒适反馈测试跨模态推理能力。
  • 模型生成的热舒适指标分布与人类数据差异显著,分类准确率接近随机。
  • 适合关注AI真实感知与决策能力、智能建筑等应用的研究者参考。

大型语言模型在特定任务上的评估存在关键空白。热舒适性作为环境因素与个人感知复杂互动的结果,涉及感官整合与适应性决策,是检验AI真实认知能力的理想范式。为此,我们提出TCEval——首个基于热舒适场景评估AI三大核心认知能力(跨模态推理、因果关联、自适应决策)的评估框架,利用大语言模型(LLM)代理进行测试。方法包括为LLM代理赋予虚拟人格属性,引导其生成衣物保温值选择及热舒适反馈,并以ASHRAE全球数据库和中国热舒适数据库验证输出结果。在四个主流大模型上的实验表明,尽管代理反馈与人类的精确匹配度有限,方向一致性在1 PMV容差下显著提升;统计分析显示,模型生成的PMV分布与人类数据存在显著差异,且在离散热舒适分类任务中表现接近随机。这些结果证实TCEval作为生态有效认知图灵测试的可行性,表明当前大模型具备基础的跨模态推理能力,但缺乏对热舒适中变量非线性关系的精确因果理解。TCEval弥补传统基准不足,推动AI评估从抽象任务能力转向具身化、情境感知的感知与决策能力,为智能建筑等以人为本的应用提供重要启示。

原文摘要 · Abstract (English)

A critical gap exists in LLM task-specific benchmarks. Thermal comfort, a sophisticated interplay of environmental factors and personal perceptions involving sensory integration and adaptive decision-making, serves as an ideal paradigm for evaluating real-world cognitive capabilities of AI systems. To address this, we propose TCEval, the first evaluation framework that assesses three core cognitive capacities of AI, cross-modal reasoning, causal association, and adaptive decision-making, by leveraging thermal comfort scenarios and large language model (LLM) agents. The methodology involves initializing LLM agents with virtual personality attributes, guiding them to generate clothing insulation selections and thermal comfort feedback, and validating outputs against the ASHRAE Global Database and Chinese Thermal Comfort Database. Experiments on four LLMs show that while agent feedback has limited exact alignment with humans, directional consistency improves significantly with a 1 PMV tolerance. Statistical tests reveal that LLM-generated PMV distributions diverge markedly from human data, and agents perform near-randomly in discrete thermal comfort classification. These results confirm the feasibility of TCEval as an ecologically valid Cognitive Turing Test for AI, demonstrating that current LLMs possess foundational cross-modal reasoning ability but lack precise causal understanding of the nonlinear relationships between variables in thermal comfort. TCEval complements traditional benchmarks, shifting AI evaluation focus from abstract task proficiency to embodied, context-aware perception and decision-making, offering valuable insights for advancing AI in human-centric applications like smart buildings.

认知评估热舒适大模型评测跨模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。