提出新模型提升大模型对多维认知状态的理解能力
Modeling Multi-Dimensional Cognitive States in Large Language Models under Cognitive Crowding

- 将认知状态建模从单一维度扩展到情绪、思维风格等四维联合分析
- 发现大模型在多维任务中性能骤降,因欧氏空间表达能力不足
- 创新使用双曲空间与对齐调优,让8B模型超越GPT-4o
建模人类认知状态对先进人工智能至关重要。现有大语言模型主要处理情感分析或立场检测等孤立任务,无法捕捉心理学定义的多个认知维度(情绪、思维风格、立场、意图)之间的交互关系。为此,我们构建了首个在上述四个维度上具统一标注的基准 CognitiveBench。在 CognitiveBench 上的实验表明,尽管大模型在单维度任务表现良好,但在联合多维建模中性能急剧下降。通过格罗莫夫 δ-双曲性分析,我们发现 CognitiveBench 具有强烈的层次结构。我们将其性能瓶颈归因于‘认知拥挤’:层次化认知状态需要指数级表示空间,而大模型的欧氏空间仅多项式增长,导致表示重叠、性能退化。为解决此不匹配问题,我们提出 HyCoLLM,将认知状态建模置于双曲空间,并通过双曲引导对齐调优对齐大模型表示。结果表明,HyCoLLM 显著提升多维认知理解能力,使 8B 参数模型优于 GPT-4o 等强基线。
原文摘要 · Abstract (English)
Modeling human cognitive states is essential for advanced artificial intelligence. Existing Large Language Models (LLMs) mainly address isolated tasks such as emotion analysis or stance detection, and fail to capture interactions among cognitive dimensions defined in psychology, including emotion, thinking style, stance, and intention. To bridge this gap, we construct CognitiveBench, the first benchmark with unified annotations across the above four dimensions. Experiments on CognitiveBench show that although LLMs perform well on single dimension tasks, their performance drops sharply in joint multi-dimensional modeling. Using Gromov $δ$-hyperbolicity analysis, we find that CognitiveBench exhibits a strong hierarchical structure. We attribute the performance bottleneck to ``Cognitive Crowding'', where hierarchical cognitive states require exponential representational space, while the Euclidean space of LLMs grows only polynomially, causing representation overlap and degraded performance. To address this mismatch, we propose HyCoLLM, which models cognitive states in hyperbolic space and aligns LLM representations via Hyperbolic Guided Alignment Tuning. Results show that HyCoLLM substantially improves multi-dimensional cognitive understanding, allowing 8B parameter model to outperform strong baselines, including GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。