Gene length and detection bias in single cell RNA sequencing protocols [version 1; referees: 2 approved]

Background: Single cell RNA sequencing (scRNA-seq) has rapidly gained popularity for profiling transcriptomes of hundreds to thousands of single cells. This technology has led to the discovery of novel cell types and revealed insights into the development of complex tissues. However, many technical...

Full description

Bibliographic Details
Main Authors:	Belinda Phipson, Luke Zappia, Alicia Oshlack
Format:	Article
Language:	English
Published:	F1000 Research Ltd 2017-04-01
Series:	F1000Research
Subjects:	Bioinformatics Genomics Methods for Diagnostic & Therapeutic Studies
Online Access:	https://f1000research.com/articles/6-595/v1

id	doaj-4134471d22ac4476b762f880a5fbcd36
record_format	Article
spelling	doaj-4134471d22ac4476b762f880a5fbcd362020-11-25T03:19:56ZengF1000 Research LtdF1000Research2046-14022017-04-01610.12688/f1000research.11290.112181Gene length and detection bias in single cell RNA sequencing protocols [version 1; referees: 2 approved]Belinda Phipson0Luke Zappia1Alicia Oshlack2Murdoch Childrens Research Institute, Parkville, Victoria, 3052, AustraliaSchool of Biosciences, University of Melbourne, Parkville, Victoria, 3010, AustraliaSchool of Biosciences, University of Melbourne, Parkville, Victoria, 3010, AustraliaBackground: Single cell RNA sequencing (scRNA-seq) has rapidly gained popularity for profiling transcriptomes of hundreds to thousands of single cells. This technology has led to the discovery of novel cell types and revealed insights into the development of complex tissues. However, many technical challenges need to be overcome during data generation. Due to minute amounts of starting material, samples undergo extensive amplification, increasing technical variability. A solution for mitigating amplification biases is to include unique molecular identifiers (UMIs), which tag individual molecules. Transcript abundances are then estimated from the number of unique UMIs aligning to a specific gene, with PCR duplicates resulting in copies of the UMI not included in expression estimates. Methods: Here we investigate the effect of gene length bias in scRNA-Seq across a variety of datasets that differ in terms of capture technology, library preparation, cell types and species. Results: We find that scRNA-seq datasets that have been sequenced using a full-length transcript protocol exhibit gene length bias akin to bulk RNA-seq data. Specifically, shorter genes tend to have lower counts and a higher rate of dropout. In contrast, protocols that include UMIs do not exhibit gene length bias, with a mostly uniform rate of dropout across genes of varying length. Across four different scRNA-Seq datasets profiling mouse embryonic stem cells (mESCs), we found the subset of genes that are only detected in the UMI datasets tended to be shorter, while the subset of genes detected only in the full-length datasets tended to be longer. Conclusions: We find that the choice of scRNA-seq protocol influences the detection rate of genes, and that full-length datasets exhibit gene-length bias. In addition, despite clear differences between UMI and full-length transcript data, we illustrate that full-length and UMI data can be combined to reveal the underlying biology influencing expression of mESCs.https://f1000research.com/articles/6-595/v1BioinformaticsGenomicsMethods for Diagnostic & Therapeutic Studies
collection	DOAJ
language	English
format	Article
sources	DOAJ
author	Belinda Phipson Luke Zappia Alicia Oshlack
spellingShingle	Belinda Phipson Luke Zappia Alicia Oshlack Gene length and detection bias in single cell RNA sequencing protocols [version 1; referees: 2 approved] F1000Research Bioinformatics Genomics Methods for Diagnostic & Therapeutic Studies
author_facet	Belinda Phipson Luke Zappia Alicia Oshlack
author_sort	Belinda Phipson
title	Gene length and detection bias in single cell RNA sequencing protocols [version 1; referees: 2 approved]
title_short	Gene length and detection bias in single cell RNA sequencing protocols [version 1; referees: 2 approved]
title_full	Gene length and detection bias in single cell RNA sequencing protocols [version 1; referees: 2 approved]
title_fullStr	Gene length and detection bias in single cell RNA sequencing protocols [version 1; referees: 2 approved]
title_full_unstemmed	Gene length and detection bias in single cell RNA sequencing protocols [version 1; referees: 2 approved]
title_sort	gene length and detection bias in single cell rna sequencing protocols [version 1; referees: 2 approved]
publisher	F1000 Research Ltd
series	F1000Research
issn	2046-1402
publishDate	2017-04-01
description	Background: Single cell RNA sequencing (scRNA-seq) has rapidly gained popularity for profiling transcriptomes of hundreds to thousands of single cells. This technology has led to the discovery of novel cell types and revealed insights into the development of complex tissues. However, many technical challenges need to be overcome during data generation. Due to minute amounts of starting material, samples undergo extensive amplification, increasing technical variability. A solution for mitigating amplification biases is to include unique molecular identifiers (UMIs), which tag individual molecules. Transcript abundances are then estimated from the number of unique UMIs aligning to a specific gene, with PCR duplicates resulting in copies of the UMI not included in expression estimates. Methods: Here we investigate the effect of gene length bias in scRNA-Seq across a variety of datasets that differ in terms of capture technology, library preparation, cell types and species. Results: We find that scRNA-seq datasets that have been sequenced using a full-length transcript protocol exhibit gene length bias akin to bulk RNA-seq data. Specifically, shorter genes tend to have lower counts and a higher rate of dropout. In contrast, protocols that include UMIs do not exhibit gene length bias, with a mostly uniform rate of dropout across genes of varying length. Across four different scRNA-Seq datasets profiling mouse embryonic stem cells (mESCs), we found the subset of genes that are only detected in the UMI datasets tended to be shorter, while the subset of genes detected only in the full-length datasets tended to be longer. Conclusions: We find that the choice of scRNA-seq protocol influences the detection rate of genes, and that full-length datasets exhibit gene-length bias. In addition, despite clear differences between UMI and full-length transcript data, we illustrate that full-length and UMI data can be combined to reveal the underlying biology influencing expression of mESCs.
topic	Bioinformatics Genomics Methods for Diagnostic & Therapeutic Studies
url	https://f1000research.com/articles/6-595/v1
work_keys_str_mv	AT belindaphipson genelengthanddetectionbiasinsinglecellrnasequencingprotocolsversion1referees2approved AT lukezappia genelengthanddetectionbiasinsinglecellrnasequencingprotocolsversion1referees2approved AT aliciaoshlack genelengthanddetectionbiasinsinglecellrnasequencingprotocolsversion1referees2approved
_version_	1724620150845472768

Gene length and detection bias in single cell RNA sequencing protocols [version 1; referees: 2 approved]

Similar Items