Rxivist logo

mbkmeans: fast clustering for single cell data using mini-batch k-means

By Stephanie C Hicks, Ruoxi Liu, Yuwei Ni, Elizabeth F Purdom, Davide Risso

Posted 27 May 2020
bioRxiv DOI: 10.1101/2020.05.27.119438

Single-cell RNA-Sequencing (scRNA-seq) is the most widely used high-throughput technology to measure genome-wide gene expression at the single-cell level. One of the most common analyses of scRNA-seq data detects distinct subpopulations of cells through the use of unsupervised clustering algorithms. However, recent advances in scRNA-seq technologies result in current datasets ranging from thousands to millions of cells. Popular clustering algorithms, such as k -means, typically require the data to be loaded entirely into memory and therefore can be slow or impossible to run with large datasets. To address this problem, we developed the mbkmeans R/Bioconductor package, an open-source implementation of the mini-batch k -means algorithm. Our package allows for on-disk data representations, such as the common HDF5 file format widely used for single-cell data, that do not require all the data to be loaded into memory at one time. We demonstrate the performance of the mbkmeans package using large datasets, including one with 1.3 million cells. We also highlight and compare the computing performance of mbkmeans against the standard implementation of k -means. Our software package is available in Bioconductor at https://bioconductor.org/packages/mbkmeans. ### Competing Interest Statement The authors have declared no competing interest.

Download data

  • Downloaded 436 times
  • Download rankings, all-time:
    • Site-wide: 53,546
    • In bioinformatics: 5,498
  • Year to date:
    • Site-wide: 10,958
  • Since beginning of last month:
    • Site-wide: 15,467

Altmetric data

Downloads over time

Distribution of downloads per paper, site-wide


Sign up for the Rxivist weekly newsletter! (Click here for more details.)