提出以'实体'为学习基本单元,让模型理解事件背后的同一主体。
Toward Machine Learning with the Unit as a Primitive: Learning from Unit-Linked Events
- 用实体作为学习的显式基本单元,明确事件归属的主体。
- 通过上下文令牌和共享响应形式,实现对同一实体的统一预测。
- 适用于需区分个体差异的场景,如个性化推荐、身份识别等。
机器学习通常基于样本建模,但多个观测或可能事件所指向的持久个体往往隐含未明。本文提出将‘单位’(unit)作为任务语义层面的显式基本单元。学习任务首先声明一组持久指代对象及同一性标准;实际值 $u$ 表示被选中的指代对象。监督学习是主要的形式化特例,其语义对象是一族单位条件响应律。同质性是这些响应律一致的特殊情况;仅基于样本的条件概率无法判断世界是否同质,或所观察到的规律只是异质族的边际分布。从数据中学习的是一个配对 $(T_ϕ, R_θ)$:一个生成上下文单位标记的分词器和一个共享的响应律形式。结构化类别将该形式视为标记中的简单关系,线性预测器为其具体实例。标记是学习方表示,通过它任务方的单位影响预测;若学习规范忽略单位信息,则为单位无关;同质性仍为世界侧响应族的属性。当身份未确定时,世界侧规律混合单位条件目标,而学习者将其共享形式与标记组合。可信解析器可固定单位并提供查找标记;否则,‘单位反演’利用事实证据生成同类型标记。无关联的单行观测无法区分异质单位世界与同质合并世界;可信同单位对可提供受限见证。形式结果集中于这一监督特例。
原文摘要 · Abstract (English)
Machine learning is usually formalized through samples, while the persistent individual to which multiple observed or possible events refer often remains implicit. We propose the \emph{unit} as an explicit primitive at the level of task semantics. A learning task first declares a population of persistent referents and a sameness criterion; the realized value $u$ denotes the selected referent. Supervised learning is the main formal specialization. Its semantic object is a family of unit-conditioned response laws. Homogeneity is the special case in which those laws coincide; a sample-only conditional is silent as to whether the world is homogeneous or the observed law is only the marginal of a heterogeneous family. What is learned from data is a pair $(T_ϕ,R_θ)$: a tokenizer that produces a contextual unit token and one shared response-law form that reads it. The structured class takes that form to be a simple relation in the token; a linear predictor is the running instance. The token is the learner-side representation through which the task-side unit affects prediction, while a learner specification that omits unit information is unit-insensitive; homogeneity remains a property of the world-side response family. When identity is unresolved, the world-side law mixes unit-conditioned targets, while the learner composes its shared form with a token. A trusted resolver may fix the unit and supply a lookup token; otherwise \emph{unit abduction} forms a token of the same type from factual evidence. Unlinked single-row observations can fail to distinguish a heterogeneous unit world from a homogeneous pooled world; trusted same-unit pairs separate a restricted witness. The formal results concern this supervised specialization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。