多人共读时,读者群体内部存在稳定分组,但跨文档一致性不明确。
Factions Within, Uncertain Across: Within-Document Reader Sub-Groups in Social Highlighting
- 通过零假设检验发现,读者在单篇文档内形成显著子群体。
- 88%文档中读者配对一致率远超可预测水平(z=+6.3)。
- 跨文档稳定性不足,可能受情境影响,无法确定是否为持久阅读习惯。
当多人共同标注同一文档时,群体是单一共识,还是内部形成标记不同内容的读者子群?基于个体标注行为体现选择性而非信号强度的前期研究,本工作在共读平台上提出一种保留边界的零模型。实验1:在单文档内,读者形成强子群体——配对一致率显著高于共享显著性、标记密度和句子流行度所能解释的范围(最近邻一致率z=+6.3,88%文档显著)。在八区块区域保持的零模型下,共同关注粗粒度区域解释了约40%的超额一致性,其余多数仍为细粒度读者特异性一致(z=+3.6,77%显著)。因此,从描述上看,文档内群体具有派系特征。实验2:这种分组是否为读者稳定特质?跨文档半样本重现性近乎为零(两样本分别+0.078和0.000),且功效校准显示仅在共读多文档的配对中测试才具信息量。在唯一有信息量的高重叠子集(k≥4)中,点估计为正但小样本、不精确,且在两个独立样本间差异大,从未达到显著,且在区域保持零模型下进一步减弱。因此,跨文档稳定性尚无定论:数据既支持情境性分组,也支持弱至中等稳定的读者特质。文档内群体是分裂的;其派系是否随读者跨越文档,则实难断言。
原文摘要 · Abstract (English)
When many people highlight the same document, is the crowd a single consensus, or is it internally structured into reader sub-groups that mark different things -- and is that structure a stable property of a reader or of the document? Building on prior work showing an individual's within-document highlighting signal is a whisper while individuality lives in selection, we ask the group-level question on a co-readership platform using a margin-preserving curveball null. Experiment 1: within a document, readers form strong sub-groups -- pairs agree far beyond what shared salience, mark density, and sentence popularity predict (nearest-neighbour agreement z=+6.3, significant in 88% of documents). Under an eight-block region-preserving null, shared engagement with the same coarse regions of the document accounts for about 40% of this excess; the majority survives as finer reader-specific agreement (z=+3.6, 77% significant). So the within-document crowd is, in a descriptive sense, factional. Experiment 2: is that grouping a stable reader trait? Here we are honest about power. The cross-document split-half reproducibility of a pair's agreement is near zero pooled (+0.078 and 0.000 in two separately drawn samples), and a power calibration shows the test is informative only for pairs that co-read many documents. In the only informative high-overlap subset (k>=4), point estimates are positive but small-sample, imprecise across the separately drawn samples, never significant, and attenuate under the region-preserving null. We therefore leave cross-document stability unresolved: the data is consistent with anything from situational grouping to a weak-to-moderate stable reader trait. The crowd is factional within a document; whether its factions follow the reader across documents is, honestly, beyond our reach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。