通过稀疏PPMI图平均提升童话家族关系推理的嵌入效果
Sparse PPMI Graph Averaging for Random Indexing Embeddings
- 用稀疏PPMI图加权平均优化随机索引嵌入,融合残差结构
- 在童话家族类比任务上准确率从19.41%提升至30.74%,增益11.32个百分点
- 该方法仅对特定童话语料有效,不具普遍适用性
我们研究了针对小规模童话语料中家族关系类比问题的稀疏后处理流程。已发表成果采用200维、每行8个非零值的均匀随机索引上下文累积,随后进行一次残差图平均:$\mathbf{E}=(1-α)\mathbf{E}_0+α\mathbf{P}\mathbf{E}_0$,其中$\mathbf{P}$为行归一化的PPMI图,$α=0.3$。最后执行终端行归一化及每维度中位数与四分位距缩放。在Google类比基准的家庭部分,506题中有272题可验证。五组种子实验显示,完整流程将准确率从19.41%提升至30.74%,提升11.32个百分点,置信区间[6.93, 15.89]。鲁棒缩放单独贡献3.24点[1.25, 5.38],图平均无缩放贡献6.18点[2.63, 9.92]。在独立40题通用网格上未见提升:完整流程下降6.00点[-13.50, -0.50],无缩放平均下降6.50点[-14.50, -0.50]。因此,仅支持对童话家族类比集的有效性,不构成通用嵌入方法。
原文摘要 · Abstract (English)
We study a specific sparse post-processing pipeline for Random Indexing (RI) on kinship analogies in a small fairytales corpus. The published artifacts use uniform RI context accumulation with 200 dimensions and eight nonzeros, followed by one residual graph average, $\mathbf{E}=(1-α)\mathbf{E}_0+α\mathbf{P}\mathbf{E}_0$, where $\mathbf{P}$ is a row-normalized PPMI graph and $α=0.3$. Terminal row normalization and per-dimension median/IQR scaling are then applied. On the Google analogy benchmark's family section, 272 of 506 questions are valid for every seed. Across five paired seeds, the complete pipeline raises accuracy from 19.41\% to 30.74\%, a gain of 11.32 percentage points with a nested-bootstrap 95\% confidence interval of [6.93, 15.89]. Robust scaling alone contributes 3.24 points [1.25, 5.38], while graph averaging without robust scaling contributes 6.18 points [2.63, 9.92]. A separate 40-question general grid does not support a general improvement: the full pipeline changes accuracy by -6.00 points [-13.50, -0.50], and averaging without robust scaling changes it by -6.50 points [-14.50, -0.50]. The supported positive claim is therefore limited to the covered fairytales kinship analogy set; the results do not establish a generally effective embedding method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。