arXiv:2606.02946cs.LGcs.CR2026-06KDD被引 1

针对直播风险检测中恶意行为手法伪装问题,提出潜空间解耦的反事实推理方法。

Outsmarting the Chameleon: Counterfactual Decoupling for Tactical OOD Shifts in Live Streaming Risk Assessment

论文配图:Outsmarting the Chameleon: Counterfactual Decoupling for Tactical OOD Shifts in Live Streaming Risk Assessment
图 1 · 摘自论文原文
  • 在潜空间分离意图与叙事特征,实现对抗性手法伪装下的稳定风险判断。
  • 在真实工业数据上,对齐率提升17.3%,误报率降低32%。
  • 适合需要实时抗伪装风险评估的平台安全团队使用。

直播已成为社交互动与数字商业的核心媒介,但面临日益复杂的恶意风险。核心挑战是战术性分布外(OOD)漂移:攻击者保持不变的恶意意图,却持续更换叙述方式以逃避检测。现有泛化方法因难以满足意图与手法紧密耦合、原始层反事实不明确等条件而失效。本文从潜在因果视角出发,提出潜预测反事实解耦(LPCD)框架,通过在潜空间建模意图与叙事变化,并强制潜空间反事实一致性,使风险预测锚定于因果稳定的恶意意图。推理时引入轻量级无参校准,进一步缓解手法引发的分布偏移。在大规模工业数据集与线上生产流量上的实验表明,LPCD持续优于当前最优基线,验证了其在真实直播场景下应对动态对抗风险的有效性。

原文摘要 · Abstract (English)

Live streaming has emerged as a primary medium for social interaction and digital commerce, yet it is increasingly plagued by sophisticated risks. A fundamental challenge in this domain is \emph{tactical out-of-distribution (OOD) shift}: while malicious actors maintain stable underlying objectives, they continuously redesign narrative packaging to evade detection. Such adversarial shifts expose critical limitations of existing OOD generalization paradigms, whose assumptions are difficult to satisfy in the presence of tightly coupled intent-tactic evolution and ill-defined raw-level counterfactuals. In this paper, we tackle this issue from a \emph{latent causal} perspective and propose \underline{L}atent-\underline{P}redictive \underline{C}ounterfactual \underline{D}ecoupling~(LPCD), a plug-in framework for robust live streaming risk assessment. LPCD enables counterfactual reasoning under adversarial tactical re-packaging by modeling intent and narrative variation at the latent level, and enforces \emph{latent counterfactual consistency} to anchor risk prediction on causally stable malicious intent. At inference time, LPCD applies a lightweight, parameter-free calibration to further mitigate tactic-induced distribution shifts. Extensive experiments on large-scale industrial datasets and online production traffic demonstrate that LPCD consistently outperforms state-of-the-art baselines, validating its effectiveness in moderating evolving adversarial risks in real-world live streaming. The project page is available at https://qiaoyran.github.io/LiveStreamingRiskAssessment/.

风险评估反事实推理直播安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。