TY - JOUR
T1 - Gapless assembly of complete human and plant chromosomes using only nanopore sequencing
AU - Koren, Sergey
AU - Bao, Zhigui
AU - Guarracino, Andrea
AU - Ou, Shujun
AU - Goodwin, Sara
AU - Jenike, Katharine M.
AU - Lucas, Julian
AU - McNulty, Brandy
AU - Park, Jimin
AU - Rautiainen, Mikko
AU - Rhie, Arang
AU - Roelofs, Dick
AU - Schneiders, Harrie
AU - Vrijenhoek, Ilse
AU - Nijbroek, Koen
AU - Nordesjo, Olle
AU - Nurk, Sergey
AU - Vella, Mike
AU - Lawrence, Katherine R.
AU - Ware, Doreen
AU - Schatz, Michael C.
AU - Garrison, Erik
AU - Huang, Sanwen
AU - McCombie, William Richard
AU - Miga, Karen H.
AU - Wittenberg, Alexander H.J.
AU - Phillippy, Adam M.
N1 - Publisher Copyright:
© 2024 Koren et al.
PY - 2024/11
Y1 - 2024/11
N2 - The combination of ultra-long (UL) Oxford Nanopore Technologies (ONT) sequencing reads with long, accurate Pacific Bioscience (PacBio) High Fidelity (HiFi) reads has enabled the completion of a human genome and spurred similar efforts to complete the genomes of many other species. However, this approach for complete, “telomere-to-telomere” genome assembly relies on multiple sequencing platforms, limiting its accessibility. ONT “Duplex” sequencing reads, where both strands of the DNA are read to improve quality, promise high per-base accuracy. To evaluate this new data type, we generated ONT Duplex data for three widely studied genomes: human HG002, Solanum lycopersicum Heinz 1706 (tomato), and Zea mays B73 (maize). For the diploid, heterozygous HG002 genome, we also used “Pore-C” chromatin contact mapping to completely phase the haplotypes. We found the accuracy of Duplex data to be similar to HiFi sequencing, but with read lengths tens of kilobases longer, and the Pore-C data to be compatible with existing diploid assembly algorithms. This combination of read length and accuracy enables the construction of a high-quality initial assembly, which can then be further resolved using the UL reads, and finally phased into chromosome-scale haplotypes with Pore-C. The resulting assemblies have a base accuracy exceeding 99.999% (Q50) and near-perfect continuity, with most chromosomes assembled as single contigs. We conclude that ONT sequencing is a viable alternative to HiFi sequencing for de novo genome assembly, and provides a multirun single-instrument solution for the reconstruction of complete genomes.
AB - The combination of ultra-long (UL) Oxford Nanopore Technologies (ONT) sequencing reads with long, accurate Pacific Bioscience (PacBio) High Fidelity (HiFi) reads has enabled the completion of a human genome and spurred similar efforts to complete the genomes of many other species. However, this approach for complete, “telomere-to-telomere” genome assembly relies on multiple sequencing platforms, limiting its accessibility. ONT “Duplex” sequencing reads, where both strands of the DNA are read to improve quality, promise high per-base accuracy. To evaluate this new data type, we generated ONT Duplex data for three widely studied genomes: human HG002, Solanum lycopersicum Heinz 1706 (tomato), and Zea mays B73 (maize). For the diploid, heterozygous HG002 genome, we also used “Pore-C” chromatin contact mapping to completely phase the haplotypes. We found the accuracy of Duplex data to be similar to HiFi sequencing, but with read lengths tens of kilobases longer, and the Pore-C data to be compatible with existing diploid assembly algorithms. This combination of read length and accuracy enables the construction of a high-quality initial assembly, which can then be further resolved using the UL reads, and finally phased into chromosome-scale haplotypes with Pore-C. The resulting assemblies have a base accuracy exceeding 99.999% (Q50) and near-perfect continuity, with most chromosomes assembled as single contigs. We conclude that ONT sequencing is a viable alternative to HiFi sequencing for de novo genome assembly, and provides a multirun single-instrument solution for the reconstruction of complete genomes.
UR - https://www.scopus.com/pages/publications/85209754190
UR - https://www.scopus.com/pages/publications/85209754190#tab=citedBy
U2 - 10.1101/gr.279334.124
DO - 10.1101/gr.279334.124
M3 - Article
C2 - 39505490
AN - SCOPUS:85209754190
SN - 1088-9051
VL - 34
SP - 1919
EP - 1930
JO - Genome Research
JF - Genome Research
IS - 11
ER -