Rxivist logo

Rxivist combines preprints from bioRxiv with data from Twitter to help you find the papers being discussed in your field. Currently indexing 62,963 bioRxiv papers from 279,321 authors.

Haplotype-aware genotyping from noisy long reads

By Jana Ebler, Marina Haukness, Trevor Pesout, Tobias Marschall, Benedict Paten

Posted 03 Apr 2018
bioRxiv DOI: 10.1101/293944

Motivation: Current genotyping approaches for single nucleotide variations (SNVs) rely on short, relatively accurate reads from second generation sequencing devices. Presently, third generation sequencing platforms able to generate much longer reads are becoming more widespread. These platforms come with the significant drawback of higher sequencing error rates, which makes them ill-suited to current genotyping algorithms. However, the longer reads make more of the genome unambiguously mappable and typically provide linkage information between neighboring variants. Results: In this paper we introduce a novel approach for haplotype-aware genotyping from noisy long reads. We do this by considering bipartitions of the sequencing reads, corresponding to the two haplotypes. We formalize the computational problem in terms of a Hidden Markov Model and compute posterior genotype probabilities using the forward-backward algorithm. Genotype predictions can then be made by picking the most likely genotype at each site. Our experiments indicate that longer reads allow significantly more of the genome to potentially be accurately genotyped. Further, we are able to use both Oxford Nanopore and Pacific Biosciences sequencing data to independently validate millions of variants previously identified by short-read technologies in the reference NA12878 sample, including hundreds of thousands of variants that were not previously included in the high-confidence reference set.

Download data

  • Downloaded 2,372 times
  • Download rankings, all-time:
    • Site-wide: 1,652 out of 62,963
    • In bioinformatics: 332 out of 6,253
  • Year to date:
    • Site-wide: 1,921 out of 62,963
  • Since beginning of last month:
    • Site-wide: 10,122 out of 62,963

Altmetric data


Downloads over time

Distribution of downloads per paper, site-wide


Sign up for the Rxivist weekly newsletter! (Click here for more details.)


News