arXiv:2606.22674stat.MLcs.AI2026-06

用哲学思想生成随时间演化的数据集,保持历史特征的相似性。

Data Evolution by Wittgenstein's Rule Following

论文配图:Data Evolution by Wittgenstein's Rule Following
图 1 · 摘自论文原文
  • 基于结构描述符捕捉数据几何与分布特性,而非点对点匹配
  • 通过轨迹外推与历史均值生成目标,再混合重构候选数据集
  • 支持样本量和维度变化,适合动态演化数据生成场景

本文提出哲数学中的维特根斯坦规则遵循(WRF)数据演化框架,用于从历史数据序列中生成新数据集。该方法受维特根斯坦《哲学研究》中规则遵循与家族相似性思想启发,不依赖固定分布采样或简单增强。WRF以结构描述符表征每个数据集,涵盖几何、分布、聚类及标签属性等特征。通过外推描述符轨迹预测规则延续目标,通过平均历史描述符生成家族相似目标。候选数据集通过平衡或有界混合重组生成,依据上述目标评分,并可选在描述符空间进行可微优化。框架允许样本数量和特征维度随时间变化,不要求下一数据集是前一数据集的直接变换。合成与图像数据集的模拟实验表明,WRF能在无监督与有监督设置下生成有意义的数据演化延续。

原文摘要 · Abstract (English)

This paper introduces Wittgenstein's Rule Following (WRF) data evolution, a framework in philomatics for evolving or generating a new dataset from a sequence of previously observed datasets. The method is inspired by Ludwig Wittgenstein's rule-following considerations and his notion of family resemblance in Philosophical Investigations. Unlike standard synthetic data generation, where the goal is usually to sample from or augment a fixed distribution, WRF aims to continue the implicit rule expressed by a historical sequence of datasets while preserving resemblance to the previous datasets. WRF represents each dataset by structural descriptors rather than pointwise correspondences. These descriptors summarize geometric, distributional, clustering, and, in the supervised case, label-based properties of the data. The method predicts a rule-following target by extrapolating descriptor trajectories and a family-resemblance target by averaging historical descriptors. Candidate datasets are then generated from the observed history through balanced or bounded mixture recombination, scored according to these targets, and optionally refined through differentiable optimization in descriptor space. The proposed framework allows both sample size and feature dimension to vary over time and does not assume that the next dataset is a direct transformation of the last one. Simulations on synthetic and image datasets show that WRF can generate meaningful continuations of evolving datasets in both unsupervised and supervised settings.

数据演化哲学建模生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。