提出任务空间复杂度新度量,用以评估零样本泛化性能。
Formalizing Task-Space Complexity for Zero-Shot Generalization

- 引入有向任务差异度量‘符号发散’,量化源任务到目标任务的性能差距。
- 证明最小源任务数可保证任意目标任务的泛化误差不超过ε,实现性能证书。
- 在弹簧-质量-阻尼与倒立摆系统上验证,贪心策略优于随机或均匀采样。
策略需在多种条件下运行,但单一策略通常保守,而完全自适应方案又过于复杂。本文研究上下文动态系统中的零样本泛化,提出一种以性能为中心、有方向性的任务差异度量——符号发散,该度量上界了从源上下文到目标上下文的泛化误差。符号发散诱导出ε容差集,用于验证源策略类是否可泛化,并给出任务空间复杂度的明确定义:使每个目标上下文的泛化误差不超过ε所需最少源上下文数量。在性能局部光滑的温和假设下,容差集具有认证的内/外球结构及实例相关的体积边界。在有限查询设置中,源任务选择等价于集合覆盖问题,贪心算法继承标准的H(n)近似保证。在使用线性二次调节器(LQR)控制器的弹簧-质量-阻尼系统和基于深度强化学习的非线性倒立摆系统上,实验表明贪心策略在达到相同ε覆盖时所需策略数少于均匀或随机基线。本方法提供基于性能的任务相似性度量及构建可泛化控制策略的实际证书。
原文摘要 · Abstract (English)
Policies must operate across diverse conditions, yet a single policy is often conservative while fully adaptive schemes can be complex. We study zero-shot generalization in contextual dynamical systems and introduce a performance-centric, directional task dissimilarity--the signed divergence--that upper bounds the generalization gap from a source context to a target context. The signed divergence induces $\varepsilon$-tolerance sets that certify when a source policy class generalizes, and it yields a concrete notion of task-space complexity: the minimum number of source contexts needed so that every target context incurs at most $\varepsilon$ generalization gap. Under a mild local smoothness assumption on performance, the induced tolerance sets admit certified inner/outer balls and instance-dependent volume bounds on task-space complexity. In the finite-oracle setting, source selection reduces to set cover; a greedy strategy inherits the standard $H(n)$ approximation guarantee. Using a Mass-Spring-Damper system with linear-quadratic regulator (LQR) controllers and a nonlinear CartPole system with deep reinforcement learning controllers, we show that greedy selection achieves the same $\varepsilon$-coverage with fewer policies than uniform or random baselines. Our approach delivers a performance-based task similarity measure and practical certificates for building generalizable control with simple policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。