Assembling haplotypes of a highly heterozygous gene cluster with canu

drosophilid

Junior Member

Join Date: Nov 2019

Posts: 1
- Share
- Tweet
#1

Assembling haplotypes of a highly heterozygous gene cluster with canu

11-25-2019, 06:47 AM

I am trying to assemble haplotypes for a peculiar region of the human genome that (1) has high heterozygosity, (2) has variation in presence or absence of entire genes, and (3) encompasses a gene cluster of highly similar paralogues. This is obviously making assembly difficult, since a paralogue on the same chromosome may have only 10% divergence from its duplicate, while the homologue on the other chromosome has 5% differences due to segregating polymorphism at that locus. Currently, I have nanopore and short reads from this region, both at approximately 30X coverage. I would like to use canu to assemble the nanopore reads, then short reads to polish, but I am getting nowhere near the full assembly. My command is

Code:

canu -p prefix -d canu_run genomeSize=250k correctedErrorRate=0.144 minOverlapLength=500 -nanopore-raw sample.fastq

Here sample.fastq are reads filtered for my region of interest, so its a fairly small total assembly. So far, I have tried varying the corrected error rate between 0.1 and 0.2, and the minOverlapLength between 500 and 1000 with no luck. Using BLAST, I can see large chunks of my genes of interest in the prefix.unassembled.fasta file. It seems varying error rates should help find a sweet spot of expected divergence between reads from the same allele at a locus, reads from different alleles at a locus, and reads from different paralogous loci - I'm wondering, is there any other parameters I can vary to try and get a more complete assembly? Is there any preprocessing I can do with the more accurate short reads to lead to a more complete assembly? Ideally, I eventually want phased haplotype information.
Tags: assembly, canu, heterozygous, nanopore, phase
SNPsaurus

Registered Vendor

Join Date: May 2013

Posts: 525
- Share
- Tweet
#2

11-25-2019, 10:15 AM

This isn't so helpful to solve the asked problem, but PacBio HiFi reads might be a better choice since you can generate 15kb sequences at 1% or 0.1% error rates. They might more obviously segregate into 4 bins (2 haplotypes at a locus, 2 at paralogous locus).

Providing nextRAD genotyping and PacBio sequencing services. http://snpsaurus.com
Comment

Previous template Next

Essential Discoveries and Tools in Epitranscriptomics

by seqadmin

The field of epigenetics has traditionally concentrated more on DNA and how changes like methylation and phosphorylation of histones impact gene expression and regulation. However, our increased understanding of RNA modifications and their importance in cellular processes has led to a rise in epitranscriptomics research. “Epitranscriptomics brings together the concepts of epigenetics and gene expression,” explained Adrien Leger, PhD, Principal Research Scientist on Modified Bases...
- Channel: Articles
Yesterday, 07:01 AM
Current Approaches to Protein Sequencing

by seqadmin

Proteins are often described as the workhorses of the cell, and identifying their sequences is key to understanding their role in biological processes and disease. Currently, the most common technique used to determine protein sequences is mass spectrometry. While still a valuable tool, mass spectrometry faces several limitations and requires a highly experienced scientist familiar with the equipment to operate it. Additionally, other proteomic methods, like affinity assays, are constrained...
- Channel: Articles
04-04-2024, 04:25 PM

Topics	Statistics	Last Post
Cancer Metastasis: A Deep Dive into Cellular Plasticity by seqadmin Started by seqadmin, 04-11-2024, 12:08 PM	0 responses 40 views 0 likes	Last Post by seqadmin 04-11-2024, 12:08 PM
Proteogenomic Profiles Offer New Clues in Prostate Cancer by seqadmin Started by seqadmin, 04-10-2024, 10:19 PM	0 responses 41 views 0 likes	Last Post by seqadmin 04-10-2024, 10:19 PM
Novel Diagnostic Assay Enhances Ovarian Cancer Detection by seqadmin Started by seqadmin, 04-10-2024, 09:21 AM	0 responses 36 views 0 likes	Last Post by seqadmin 04-10-2024, 09:21 AM
Evolutionary Dynamics of Centromeres: A Comparative Genomic Analysis by seqadmin Started by seqadmin, 04-04-2024, 09:00 AM	0 responses 55 views 0 likes	Last Post by seqadmin 04-04-2024, 09:00 AM

Seqanswers Leaderboard Ad

Announcement

Assembling haplotypes of a highly heterozygous gene cluster with canu

Comment

Latest Articles

ad_right_rmr

News