通过可学习的灰-怀恩网络,实现多任务视觉信息共享,减少冗余。
Lossy Common Information in a Learnable Gray-Wyner Network
- 设计三通道可学习编码器,分离多任务共用与特定信息。
- 在六个视觉基准上,相比独立编码,冗余降低且性能更优。
- 将经典信息论与现代任务驱动学习结合,适合多任务模型研究者。
许多计算机视觉任务共享大量重叠信息,但传统编码器常忽略这一点,导致表示冗余且低效。灰-怀恩网络作为信息论中的经典概念,为分离共用信息与任务特异性信息提供了理论框架。受此启发,我们提出一种可学习的三通道编码器,能在多个视觉任务间解耦共享信息与任务特定细节。通过引入“有损共用信息”的概念,刻画该方法的极限,并设计优化目标以平衡学习过程中的固有权衡。在涵盖六项视觉基准的双任务场景中,对比三种编码器架构,结果表明本方法显著降低冗余,且性能持续优于独立编码。这些成果凸显了在现代机器学习背景下重新审视灰-怀恩理论的实际价值,弥合了经典信息论与任务驱动表征学习之间的鸿沟。
原文摘要 · Abstract (English)
Many computer vision tasks share substantial overlapping information, yet conventional codecs tend to ignore this, leading to redundant and inefficient representations. The Gray-Wyner network, a classical concept from information theory, offers a principled framework for separating common and task-specific information. Inspired by this idea, we develop a learnable three-channel codec that disentangles shared information from task-specific details across multiple vision tasks. We characterize the limits of this approach through the notion of lossy common information, and propose an optimization objective that balances inherent tradeoffs in learning such representations. Through comparisons of three codec architectures on two-task scenarios spanning six vision benchmarks, we demonstrate that our approach substantially reduces redundancy and consistently outperforms independent coding. These results highlight the practical value of revisiting Gray-Wyner theory in modern machine learning contexts, bridging classic information theory with task-driven representation learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。