Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • gnatnog
    Junior Member
    • Mar 2013
    • 2

    20kb Plasmid Assembly

    Hello all,

    I've very new to this type of analysis and was hoping I could get some help. I am trying to assemble the sequences of a 20kb plasmid that I had sequenced. They are short single reads that when all added up have over 4000x coverage. What I have done so far is I have trimmed off the first 4 and last 5 bases of the reads, since they were the lowest values when I quality checked them. I am trying to reduce the number of overall reads by quality filtering then down to get somewhere around 50x coverage.

    My plan is to use velvet to assemble, but I am a bit confused about what to do with the output. I have tested it a couple of times and can get the output files, but I have no idea what the next step should be. How do I decide what contigs are good and what are bad? I know that my plasmid is about 20kb, so should I just dismiss anything larger? The contigs file has a lot of different sequences, and I am not sure how to narrow it down from there. Any help would be greatly appreciated!
  • krobison
    Senior Member
    • Nov 2007
    • 734

    #2
    You might want to try using MUSKET or one of the other k-mer based tools out there to correct errors. This may be more effective than your trimming, and you might want to try assembling both trimmed & untrimmed data

    How long are your reads? Longer is better -- which is why paired end is far better than single end.

    Do you know anything about your plasmid backbone? What host was the plasmid prepared in? Screening out contigs which match the host (e.g. E.coli DH10B) would be a valuable next step. Screening out contigs corresponding to the center of the backbone may be useful -- as well as identifying the vector-insert junctions.

    Comment

    • gnatnog
      Junior Member
      • Mar 2013
      • 2

      #3
      I'll look into MUSKET for sure. What I'm starting to wonder is if I should not trim to reduce my coverage, but take a percentage of the reads. The problem is I have not a clue how to do that.

      My reads are 100bp. I trimmed them to 90bp.

      I can't find much on the plasmid backbone, that is the point of the sequencing. I'm trying to find the backbone of a plasmid I am developing. I do know about 3kb of the sequence, since I have inserted it myself. I'm estimating the plasmid to be about 20kb, as that is what makes sense from some restriction digests. I can take portions of the 3kb I inserted and find them in different contigs that velvet spits out, so at least I know that my plasmid is in there.

      Comment

      • mchaisso
        Member
        • Apr 2008
        • 84

        #4
        If you have an additional few hundred dollars to commit to the project, why not just run PacBio sequencing? Since the reads are O(length of the plasmid), it becomes MSA rather than assembly...


        Originally posted by gnatnog View Post
        Hello all,

        I've very new to this type of analysis and was hoping I could get some help. I am trying to assemble the sequences of a 20kb plasmid that I had sequenced. They are short single reads that when all added up have over 4000x coverage. What I have done so far is I have trimmed off the first 4 and last 5 bases of the reads, since they were the lowest values when I quality checked them. I am trying to reduce the number of overall reads by quality filtering then down to get somewhere around 50x coverage.

        My plan is to use velvet to assemble, but I am a bit confused about what to do with the output. I have tested it a couple of times and can get the output files, but I have no idea what the next step should be. How do I decide what contigs are good and what are bad? I know that my plasmid is about 20kb, so should I just dismiss anything larger? The contigs file has a lot of different sequences, and I am not sure how to narrow it down from there. Any help would be greatly appreciated!

        Comment

        • krobison
          Senior Member
          • Nov 2007
          • 734

          #5
          Originally posted by mchaisso View Post
          If you have an additional few hundred dollars to commit to the project, why not just run PacBio sequencing? Since the reads are O(length of the plasmid), it becomes MSA rather than assembly...

          That's a bit of a stretch; you won't have many reads quite that long, and they probably won't survive error correction.

          On the other hand, even one SMRT cell on such a small genome -- barring rampant host contamination -- should give a number of high quality long reads to easily assemble the genome

          Comment

          • krobison
            Senior Member
            • Nov 2007
            • 734

            #6
            Originally posted by gnatnog View Post
            I can't find much on the plasmid backbone, that is the point of the sequencing. I'm trying to find the backbone of a plasmid I am developing. I do know about 3kb of the sequence, since I have inserted it myself. I'm estimating the plasmid to be about 20kb, as that is what makes sense from some restriction digests. I can take portions of the 3kb I inserted and find them in different contigs that velvet spits out, so at least I know that my plasmid is in there.
            Try mapping the reads back to your backbone with Bowtie2. That would give you an estimate of what coverage you actually achieved for your plasmid. Also worth taking your longest contigs and BLASTing them against all known sequences; sometimes that can be an eye opener.

            I'm not a big fan of trimming, but the truth is I haven't played with it much. One challenge is that many assemblers have logic to deal with these issues, but not always clear how much -- so there is a complex interaction between various pre-processing tools & assemblers, and only empirically can you really work out the best strategy.

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              Today, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM
            • SEQadmin2
              Cancer Drug Resistance: The Lingering Barrier to Rising Survival
              by SEQadmin2



              Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

              There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
              07-08-2026, 05:17 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, Today, 11:10 AM
            0 responses
            7 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-13-2026, 10:26 AM
            0 responses
            29 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-09-2026, 10:04 AM
            0 responses
            38 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-08-2026, 10:08 AM
            0 responses
            25 views
            0 reactions
            Last Post SEQadmin2  
            Working...