arXiv:2602.20224cs.LGcs.AI2026-02

用凸优化方法分析衰老研究,自动提炼出可解释的热点主题。

Exploring Anti-Aging Literature via ConvexTopics and Large Language Models

  • 基于凸优化选择数据典型样本,避免局部最优,保证结果稳定。
  • 从1.2万篇衰老文献中提取出分子机制、饮食补充剂等可验证主题。
  • 适合生物医学研究者追踪前沿趋势,也适用于构建可复现的知识图谱。

生物医学文献的快速扩张带来了知识组织和新兴趋势识别的挑战,亟需可扩展且可解释的方法。传统聚类与主题建模方法如K-means或LDA对初始化敏感,易陷入局部最优,影响可重复性与评估。本文提出一种基于凸优化的聚类算法重构,通过从数据中选取典型样本,确保全局最优,生成稳定且细粒度的主题。该方法应用于约1.2万篇关于衰老与长寿的PubMed文章,所发现的主题经医学专家验证,涵盖分子机制、膳食补充剂、体育活动及肠道菌群等方向。相比K-means、LDA与BERTopic,本方法在性能上表现良好,尤其在可复现性与可解释性方面更具优势。本工作为开发可扩展、网页可访问的知识发现工具奠定了基础。

原文摘要 · Abstract (English)

The rapid expansion of biomedical publications creates challenges for organizing knowledge and detecting emerging trends, underscoring the need for scalable and interpretable methods. Common clustering and topic modeling approaches such as K-means or LDA remain sensitive to initialization and prone to local optima, limiting reproducibility and evaluation. We propose a reformulation of a convex optimization based clustering algorithm that produces stable, fine-grained topics by selecting exemplars from the data and guaranteeing a global optimum. Applied to about 12,000 PubMed articles on aging and longevity, our method uncovers topics validated by medical experts. It yields interpretable topics spanning from molecular mechanisms to dietary supplements, physical activity, and gut microbiota. The method performs favorably, and most importantly, its reproducibility and interpretability distinguish it from common clustering approaches, including K-means, LDA, and BERTopic. This work provides a basis for developing scalable, web-accessible tools for knowledge discovery.

主题建模衰老研究凸优化可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。