DNASTAR, INC. — Department of Health and Human Services SBIR Phase II: 400

DNASTAR, INC. — SBIR Phase II award from Department of Health and Human Services.

Amount
$1,566,575
Agency
Department of Health and Human Services · National Institutes of Health
Program / Phase
SBIR · Phase II
Topic
400
Solicitation
PA18-591
NAICS
Place of performance
WI
Period
2017-03-01 → 2020-02-29

Description

Despite the tremendous success of short read next generation sequencingNGStechnologiestheir inherent inability to establish long range connectivity makes fundamental tasks such as genome closurehaplotype phasing and alternatively spliced transcript characterization all but impossibleNowtwo long read sequencing providersPacific BiosciencesPacBioand Oxford Nanopore TechnologiesONTare producing data that can overcome these critical shortcomingsPacBio is capable of producingkb reads and has seen increased adoption for closing microbial genomes in particularbut also for eurkaryotic genomics and transcriptomicsONT s MinION device is a portable real time sequencing platform capable of producingkb reads and has already been successfully applied to microbial sequencing and pathogen identificationONT s new high throughput instrumentthe PromethIONis being released inand will have sufficient output for human genome scale experimentsThe tremendous potential of both technologies is currently hampered by high error rateswhich makes assembly and consensus calling extremely computationally challengingVarious command line software programs have been developed to tackle these challengesbut they typically require substantial bioinformatic expertise and computing resources savvy and do not address the critical hurdles associated with diploid genomesWith long read sequencing poised to become a major resource for genomicsthere is clearly an urgent need for integrated easy to use assembly and analysis software that can handle and exploit the unique aspects of this dataToward that endwe have developed a prototype de novo assembler based on our patented Disk Sort AlignmentDSAalgorithm that can assemble an uncorrected bacterial genome data set into a single contig with andgtbase accuracy on a standard desktop computer in less thanhoursThe assembler uses DSA determined read overlaps to construct an assembly string graph from which a layout is fed to a novel consensus generator designed to maximize accuracy from this error prone dataThe overall goal of this direct to Phase II proposal is to transform the prototype into a fully scalable long read de novo assembler for both haploid and diploid genomesWe will first optimize the performance of the assembler componentsbuilding a solid foundation from which to incorporate the essential diploid aware capabilities ofidentifying large structural variation between two sister chromosomesadapting the consensus base caller to handle heterozygous SNVs and small indels andexploiting the long range connectivity of the data to properly phase the variants and produce accurate haplotype sequencesFinallywe will leverage these tools to identify alternatively spliced transcripts and allelespecific expression from long read RNA Seq dataConsistent with DNASTAR syear history of delivering easy to use expert level softwarethis assembler will give any user access to these revolutionary long read sequencing technologies and those to come Public Health Relevance Emerginglong readtechnologies have the ability to sequence DNA molecules one thousand times longer than currentnext generationinstrumentsThis remarkable advance has tremendous implications for genomic sciencesincluding supporting enhanced understanding of the causes of and cures for human diseaseIn this projectwe will develop the software needed to accurately convert this new data into truly complete genome sequences for any individual