arXiv:2505.04637cs.CLcs.AI2025-05被引 2

让AI像人一样分块理解图文信息,提升理解准确率。

Adaptive Token Boundaries: Integrating Human Chunking Mechanisms into Multimodal LLMs

  • 根据人类认知规律动态调整图文分块边界。
  • 在视觉问答任务上比现有模型高7.8%准确率。
  • 适合研究人机认知对齐与多模态AI优化的学者。

多模态大语言模型在处理多种数据类型方面取得显著进展,但其计算方法与人类认知过程仍存在明显差异。本研究系统探究了人类跨模态分块机制与多模态大模型分词方法之间的相似性。通过对比人类表现与模型行为在视觉-语言任务中的表现,发现传统静态分词方式严重限制了模型模拟人类动态、上下文敏感信息处理的能力。为此,提出一种融合自适应边界、层次化表示与认知科学对齐机制的动态跨模态分词框架。定量评估显示,该方法在基准任务上显著优于当前最先进模型:视觉问答任务提升7.8%,复杂场景描述任务提升5.3%,同时误差模式和注意力分布更贴近人类。研究成果深化了对人脑认知与人工智能关系的理论理解,并为构建更具认知合理性的人工智能系统提供了实证支持。

原文摘要 · Abstract (English)

Recent advancements in multimodal large language models (MLLMs) have demonstrated remarkable capabilities in processing diverse data types, yet significant disparities persist between human cognitive processes and computational approaches to multimodal information integration. This research presents a systematic investigation into the parallels between human cross-modal chunking mechanisms and token representation methodologies in MLLMs. Through empirical studies comparing human performance patterns with model behaviors across visual-linguistic tasks, we demonstrate that conventional static tokenization schemes fundamentally constrain current models' capacity to simulate the dynamic, context-sensitive nature of human information processing. We propose a novel framework for dynamic cross-modal tokenization that incorporates adaptive boundaries, hierarchical representations, and alignment mechanisms grounded in cognitive science principles. Quantitative evaluations demonstrate that our approach yields statistically significant improvements over state-of-the-art models on benchmark tasks (+7.8% on Visual Question Answering, +5.3% on Complex Scene Description) while exhibiting more human-aligned error patterns and attention distributions. These findings contribute to the theoretical understanding of the relationship between human cognition and artificial intelligence, while providing empirical evidence for developing more cognitively plausible AI systems.

多模态认知对齐分词优化视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。