Skip to main navigation Skip to search Skip to main content

Scalable semi-supervised clustering via structural entropy with different constraints

Guangjie Zeng, Hao Peng*, Angsheng Li, Jia Wu, Chunyang Liu, Philip S. Yu

*Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

Abstract

Semi-supervised clustering leverages prior information in the form of constraints to achieve higher-quality clustering outcomes. However, most existing methods struggle with large-scale datasets owing to their high time and space complexity. Moreover, they encounter the challenge of seamlessly integrating various constraints, thereby limiting their applicability. In this paper, we present Scalable Semi-supervised clustering via Structural Entropy (SSSE), a novel method that tackles scalable datasets with different types of constraints from diverse sources to perform both semi-supervised partitioning and hierarchical clustering, which is fully explainable compared to deep learning-based methods. Specifically, we design objectives based on structural entropy, integrating constraints for semi-supervised partitioning and hierarchical clustering. To achieve scalability on data size, we develop efficient algorithms based on graph sampling to reduce the time and space complexity. To achieve generalization on constraint types, we formulate a uniform view for widely used pairwise and label constraints. Extensive experiments on real-world clustering datasets at different scales demonstrate the superiority of SSSE in clustering accuracy and scalability with different constraints. Additionally, Cell clustering experiments on single-cell RNA-seq datasets demonstrate the functionality of SSSE for biological data analysis.

Original languageEnglish
Pages (from-to)478-492
Number of pages15
JournalIEEE Transactions on Knowledge and Data Engineering
Volume37
Issue number1
DOIs
Publication statusPublished - Jan 2025

Keywords

  • biological data analysis
  • scalable clustering
  • semi-supervised clustering
  • structural entropy

Fingerprint

Dive into the research topics of 'Scalable semi-supervised clustering via structural entropy with different constraints'. Together they form a unique fingerprint.

Cite this