Mbkmeans: Fast clustering for single cell data using mini-batch k-means

Stephanie C. Hicks; Ruoxi Liu; Yuwei Ni; Elizabeth Purdom; Davide Risso

doi:10.1371/JOURNAL.PCBI.1008625

Mbkmeans: Fast clustering for single cell data using mini-batch k-means

Stephanie C. Hicks, Ruoxi Liu, Yuwei Ni, Elizabeth Purdom, Davide Risso

Bloomberg School of Public Health

Research output: Contribution to journal › Article › peer-review

Abstract

Single-cell RNA-Sequencing (scRNA-seq) is the most widely used high-throughput technology to measure genome-wide gene expression at the single-cell level. One of the most common analyses of scRNA-seq data detects distinct subpopulations of cells through the use of unsupervised clustering algorithms. However, recent advances in scRNA-seq technologies result in current datasets ranging from thousands to millions of cells. Popular clustering algorithms, such as k-means, typically require the data to be loaded entirely into memory and therefore can be slow or impossible to run with large datasets. To address this problem, we developed the mbkmeans R/Bioconductor package, an open-source implementation of the mini-batch k-means algorithm. Our package allows for on-disk data representations, such as the common HDF5 file format widely used for single-cell data, that do not require all the data to be loaded into memory at one time. We demonstrate the performance of the mbkmeans package using large datasets, including one with 1.3 million cells. We also highlight and compare the computing performance of mbkmeans against the standard implementation of k-means and other popular single-cell clustering methods. Our software package is available in Bioconductor at https://bioconductor.org/packages/ mbkmeans.

Original language	English (US)
Article number	e1008625
Journal	PLoS computational biology
Volume	17
Issue number	1
DOIs	https://doi.org/10.1371/JOURNAL.PCBI.1008625
State	Published - Jan 26 2021

ASJC Scopus subject areas

Ecology, Evolution, Behavior and Systematics
Modeling and Simulation
Ecology
Molecular Biology
Genetics
Cellular and Molecular Neuroscience
Computational Theory and Mathematics

Access to Document

10.1371/JOURNAL.PCBI.1008625

Cite this

@article{faf28cbcbd8447dd9e908d0a13072ee3,

title = "Mbkmeans: Fast clustering for single cell data using mini-batch k-means",

abstract = "Single-cell RNA-Sequencing (scRNA-seq) is the most widely used high-throughput technology to measure genome-wide gene expression at the single-cell level. One of the most common analyses of scRNA-seq data detects distinct subpopulations of cells through the use of unsupervised clustering algorithms. However, recent advances in scRNA-seq technologies result in current datasets ranging from thousands to millions of cells. Popular clustering algorithms, such as k-means, typically require the data to be loaded entirely into memory and therefore can be slow or impossible to run with large datasets. To address this problem, we developed the mbkmeans R/Bioconductor package, an open-source implementation of the mini-batch k-means algorithm. Our package allows for on-disk data representations, such as the common HDF5 file format widely used for single-cell data, that do not require all the data to be loaded into memory at one time. We demonstrate the performance of the mbkmeans package using large datasets, including one with 1.3 million cells. We also highlight and compare the computing performance of mbkmeans against the standard implementation of k-means and other popular single-cell clustering methods. Our software package is available in Bioconductor at https://bioconductor.org/packages/ mbkmeans.",

author = "Hicks, {Stephanie C.} and Ruoxi Liu and Yuwei Ni and Elizabeth Purdom and Davide Risso",

note = "Publisher Copyright: {\textcopyright} 2021 Hicks et al.",

year = "2021",

month = jan,

day = "26",

doi = "10.1371/JOURNAL.PCBI.1008625",

language = "English (US)",

volume = "17",

journal = "PLoS computational biology",

issn = "1553-734X",

publisher = "Public Library of Science",

number = "1",

}

TY - JOUR

T1 - Mbkmeans

T2 - Fast clustering for single cell data using mini-batch k-means

AU - Hicks, Stephanie C.

AU - Liu, Ruoxi

AU - Ni, Yuwei

AU - Purdom, Elizabeth

AU - Risso, Davide

PY - 2021/1/26

Y1 - 2021/1/26

N2 - Single-cell RNA-Sequencing (scRNA-seq) is the most widely used high-throughput technology to measure genome-wide gene expression at the single-cell level. One of the most common analyses of scRNA-seq data detects distinct subpopulations of cells through the use of unsupervised clustering algorithms. However, recent advances in scRNA-seq technologies result in current datasets ranging from thousands to millions of cells. Popular clustering algorithms, such as k-means, typically require the data to be loaded entirely into memory and therefore can be slow or impossible to run with large datasets. To address this problem, we developed the mbkmeans R/Bioconductor package, an open-source implementation of the mini-batch k-means algorithm. Our package allows for on-disk data representations, such as the common HDF5 file format widely used for single-cell data, that do not require all the data to be loaded into memory at one time. We demonstrate the performance of the mbkmeans package using large datasets, including one with 1.3 million cells. We also highlight and compare the computing performance of mbkmeans against the standard implementation of k-means and other popular single-cell clustering methods. Our software package is available in Bioconductor at https://bioconductor.org/packages/ mbkmeans.

AB - Single-cell RNA-Sequencing (scRNA-seq) is the most widely used high-throughput technology to measure genome-wide gene expression at the single-cell level. One of the most common analyses of scRNA-seq data detects distinct subpopulations of cells through the use of unsupervised clustering algorithms. However, recent advances in scRNA-seq technologies result in current datasets ranging from thousands to millions of cells. Popular clustering algorithms, such as k-means, typically require the data to be loaded entirely into memory and therefore can be slow or impossible to run with large datasets. To address this problem, we developed the mbkmeans R/Bioconductor package, an open-source implementation of the mini-batch k-means algorithm. Our package allows for on-disk data representations, such as the common HDF5 file format widely used for single-cell data, that do not require all the data to be loaded into memory at one time. We demonstrate the performance of the mbkmeans package using large datasets, including one with 1.3 million cells. We also highlight and compare the computing performance of mbkmeans against the standard implementation of k-means and other popular single-cell clustering methods. Our software package is available in Bioconductor at https://bioconductor.org/packages/ mbkmeans.

UR - http://www.scopus.com/inward/record.url?scp=85101135232&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=85101135232&partnerID=8YFLogxK

U2 - 10.1371/JOURNAL.PCBI.1008625

DO - 10.1371/JOURNAL.PCBI.1008625

M3 - Article

C2 - 33497379

AN - SCOPUS:85101135232

SN - 1553-734X

VL - 17

JO - PLoS computational biology

JF - PLoS computational biology

IS - 1

M1 - e1008625

ER -

Mbkmeans: Fast clustering for single cell data using mini-batch k-means

Abstract

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this