Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • trouble shooting PB assembly generating a larger than expected contig

    So I've got a newly assembled genome that was sequenced using PacBio sequencing and obtained >100x coverage. It was assembled with HGAP.3 and for the most part looks great except for an exceptionally large contig (the largest actually). it's estimated that our genomes largest chromosome is around 3.5Mb (based on electrokaryograph) but this one large contig is around 5Mb and has low complexity across the entire contig. Blasting this contig doesn't result in any matches of note at NCBI or to our reference genome.

    Any suggestions on how I could adjust the assembly to get rid of this large contig that I am fairly sure is not real?

  • #2
    Why don't you try gene-prediction on it to see what you get? Also, based on mapping, what kind of coverage does this contig have? Note, also, that it could be a bacterial symbiont, so you might want to try prokaryotic gene-calling. Bacteria don't usually have low complexity, though.

    Comment


    • #3
      Looking at the coverage when remapping all the raw data would be the most telling, does it have 100x coverage, is the coverage even?
      I'm actually really intrigued, I've done a lot of HGAP.3 assemblies, but have never seen a 'junk' contig get anywhere near that big. It's possible its just all the low complexity repeats getting overlapped together, but normally this would generate at max 10's of kb of sequence. You can also look at the overlap graph to see what the origin of the contig is https://gist.github.com/rhallPB/2d962e700d83270b0109 .

      Comment


      • #4
        Do you have Illumina reads that you can map to this contig in local mode (e.g. 'MagicBLAST' or 'Bowtie2 --local')? If not, you could try digitally fragmenting your PacBio reads into short reads and mapping.

        Map only to the single contig. As rhall has said, you should get a somewhat even coverage across this contig. Any big jumps in coverage indicate something that needs further investigation. A shift from one coverage level to another might indicate a misassembly, while a blip of extremely high coverage suggests transposon sequence that may be interfering with assembly.

        Comment

        Latest Articles

        Collapse

        • seqadmin
          Recent Advances in Sequencing Analysis Tools
          by seqadmin


          The sequencing world is rapidly changing due to declining costs, enhanced accuracies, and the advent of newer, cutting-edge instruments. Equally important to these developments are improvements in sequencing analysis, a process that converts vast amounts of raw data into a comprehensible and meaningful form. This complex task requires expertise and the right analysis tools. In this article, we highlight the progress and innovation in sequencing analysis by reviewing several of the...
          Today, 07:48 AM
        • seqadmin
          Essential Discoveries and Tools in Epitranscriptomics
          by seqadmin




          The field of epigenetics has traditionally concentrated more on DNA and how changes like methylation and phosphorylation of histones impact gene expression and regulation. However, our increased understanding of RNA modifications and their importance in cellular processes has led to a rise in epitranscriptomics research. “Epitranscriptomics brings together the concepts of epigenetics and gene expression,” explained Adrien Leger, PhD, Principal Research Scientist...
          04-22-2024, 07:01 AM

        ad_right_rmr

        Collapse

        News

        Collapse

        Topics Statistics Last Post
        Started by seqadmin, Today, 07:17 AM
        0 responses
        11 views
        0 likes
        Last Post seqadmin  
        Started by seqadmin, 05-02-2024, 08:06 AM
        0 responses
        19 views
        0 likes
        Last Post seqadmin  
        Started by seqadmin, 04-30-2024, 12:17 PM
        0 responses
        20 views
        0 likes
        Last Post seqadmin  
        Started by seqadmin, 04-29-2024, 10:49 AM
        0 responses
        28 views
        0 likes
        Last Post seqadmin  
        Working...
        X