用张量统一视觉任务,突破矩阵思维局限
Multidimensional Task Learning: A Unified Tensor Framework for Computer Vision Tasks
- 基于张量的GE-MLP直接操作多维数据,避免信息丢失
- 证明分类、分割、检测都属同一张量任务框架的不同配置
- 适合研究多维度任务设计与跨模态/时序预测的学者
本文提出多维任务学习(MTL),一种基于广义爱因斯坦多层感知机(GE-MLPs)的统一数学框架,通过爱因斯坦积直接在张量上操作。现有视觉任务受限于矩阵思维:标准架构使用矩阵权重和向量偏置,需结构展平,限制了自然可表达任务的范围。GE-MLPs通过张量参数突破此限制,可显式控制哪些维度保留或收缩,无信息损失。我们通过严格数学推导证明,分类、分割、检测均为MTL的特例,仅在形式化的任务空间中维度配置不同。进一步证明该任务空间严格大于矩阵基表述能力,支持如时空或跨模态预测等需破坏性展平的传统方法无法处理的任务。本工作为理解、比较与设计视觉任务提供了张量代数基础。
原文摘要 · Abstract (English)
This paper introduces Multidimensional Task Learning (MTL), a unified mathematical framework based on Generalized Einstein MLPs (GE-MLPs) that operate directly on tensors via the Einstein product. We argue that current computer vision task formulations are inherently constrained by matrix-based thinking: standard architectures rely on matrix-valued weights and vectorvalued biases, requiring structural flattening that restricts the space of naturally expressible tasks. GE-MLPs lift this constraint by operating with tensor-valued parameters, enabling explicit control over which dimensions are preserved or contracted without information loss. Through rigorous mathematical derivations, we demonstrate that classification, segmentation, and detection are special cases of MTL, differing only in their dimensional configuration within a formally defined task space. We further prove that this task space is strictly larger than what matrix-based formulations can natively express, enabling principled task configurations such as spatiotemporal or cross modal predictions that require destructive flattening under conventional approaches. This work provides a mathematical foundation for understanding, comparing, and designing computer vision tasks through the lens of tensor algebra.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。