Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • ewilbanks
    Member
    • Mar 2009
    • 83

    #1

    Metagenomics w/ 454 tips?

    Hi yall,

    Quick question for you. Does anyone have tips, tricks or recommendations for metagenomic assembly and binning programs? I'm working with two 200K read (100-350bp) datasets from microbial communities that are relatively simple (predicted to have fewer than 100 taxa, with a handful of dominant organisms). What are your favorites? Any pitfalls to avoid?

    Cheers,
    Lizzy
  • raw937
    Member
    • Aug 2010
    • 18

    #2
    assembly for 454 data

    454 data is a mess, but its the only long read technology as of today.
    Before, you try assembly be strict on your front end cleaning of you data. You must screen your reads hardcore (if you barcoded any samples) use tag cleaner to remove tags. Also, a removal of Ns and low quality scores would be helpful. You could try a de noising program if it is amplicon but I have not tried it for metas.
    Once you have removed all the homopolymers etc.
    Then forge or mira would be good start for your assembly.
    What percentage of your reads are 100 bp?
    If 50% then try abyss or velvet.

    More details would be help?

    Comment

    • rwenang
      Member
      • Jan 2009
      • 31

      #3
      you can try QIIME to process the data. http://qiime.sourceforge.net/index.html

      Comment

      • raw937
        Member
        • Aug 2010
        • 18

        #4
        Qiime!!! Is not for metas!!!

        Not for metas!

        Comment

        • rwenang
          Member
          • Jan 2009
          • 31

          #5
          Ah yes, its not 16s metagenomics. Definitely need another cup of coffee

          Comment

          • cliffbeall
            Senior Member
            • Jan 2010
            • 144

            #6
            I have been following the literature and it seems a new metagenome binning or taxonomy program comes out every month. It would be nice to see a comparison.

            I have used MEGAN, I think that it is one of the more used tools. It parses BLASTx results using the NCBI taxonomy, SEED, and KEGG. The BLASTx search is computationally intensive - A 275 megabase Illumina data set took about 1600 hours of computer time on our local cluster.

            Comment

            • MadsAlbertsen
              Member
              • Aug 2010
              • 26

              #7
              I would not try to assemble the data at all. 200k 454 reads seems very low to get any decent assembly even in very simple communities (or even in single genomes).

              200.000 reads x 250 bp read length = 50 Mb of sequence.

              50 Mb of sequence = 10x coverage of 1 genome.

              The easy way is to upload your data to the MG-RAST server (http://metagenomics.anl.gov/).

              It automatically annotates your sample to various databases and allows for comparison with a lot of public metagenomes.

              In addition to MG-RAST i've been using MEGAN and I very much like the reasoning behind the apporach. But if you do not have a reasonable computer cluster available it will take too long to BLASTX 200k reads against e.g. NCBI nr..

              rgds
              Mads

              Comment

              • raw937
                Member
                • Aug 2010
                • 18

                #8
                Metagenomic binning?

                Originally posted by cliffbeall View Post
                I have been following the literature and it seems a new metagenome binning or taxonomy program comes out every month. It would be nice to see a comparison.

                I have used MEGAN, I think that it is one of the more used tools. It parses BLASTx results using the NCBI taxonomy, SEED, and KEGG. The BLASTx search is computationally intensive - A 275 megabase Illumina data set took about 1600 hours of computer time on our local cluster.
                Cliff, did you assemble the illumina data set with abyss or velvet first?
                BlastX has a hard time with 76 bp or 100 bp read lengths.
                Meta Velvet looks like a sexy new way to assemble short read meta data.
                MEGAN is a good one and is most used, it does have a HIGH false positive rate. For microbes, SOrt-items and various IMER binning programs are around. Provide is great for viral metas. However, many others. Would be interested in a program that can input both 454 and illumina data without the flowgrams from 454.
                The other way I have thought about it is assemble then find protein orfs, then use blastp to compare various binning programs.
                BlastX takes forever!!

                Comment

                • cliffbeall
                  Senior Member
                  • Jan 2010
                  • 144

                  #9
                  Originally posted by raw937 View Post
                  Cliff, did you assemble the illumina data set with abyss or velvet first?
                  BlastX has a hard time with 76 bp or 100 bp read lengths.
                  Meta Velvet looks like a sexy new way to assemble short read meta data.
                  MEGAN is a good one and is most used, it does have a HIGH false positive rate. For microbes, SOrt-items and various IMER binning programs are around. Provide is great for viral metas. However, many others. Would be interested in a program that can input both 454 and illumina data without the flowgrams from 454.
                  The other way I have thought about it is assemble then find protein orfs, then use blastp to compare various binning programs.
                  BlastX takes forever!!
                  In the example I was quoting I didn't assemble first. I have done assembly with SOAP denovo but I didn't have enough coverage except for the most abundant sequences. Fortunately I get free time on the cluster (way to go, Ohio!).

                  Comment

                  • cliffbeall
                    Senior Member
                    • Jan 2010
                    • 144

                    #10
                    To add a data point, I did a quick benchmark with USEARCH. In my hands it is about 10X faster than blastx for searching Illumina reads against nr.

                    The drawbacks are that it uses more memory than blast so I had to split the database, and the results are not directly importable into MEGAN, though that should be doable with some work.

                    Comment

                    Latest Articles

                    Collapse

                    • SEQadmin2
                      Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                      by SEQadmin2



                      CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                      Despite this, “CRISPR helped turn genome editing from a specialized technique into
                      ...
                      07-31-2026, 11:01 AM
                    • SEQadmin2
                      Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                      by SEQadmin2


                      Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                      The systematic characterization of the human proteome has
                      ...
                      07-20-2026, 11:48 AM
                    • SEQadmin2
                      Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                      by SEQadmin2



                      Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                      ...
                      07-09-2026, 11:10 AM

                    ad_right_rmr

                    Collapse

                    News

                    Collapse

                    Topics Statistics Last Post
                    Started by SEQadmin2, Today, 10:13 AM
                    0 responses
                    13 views
                    0 reactions
                    Last Post SEQadmin2  
                    Started by SEQadmin2, 07-31-2026, 02:55 AM
                    0 responses
                    23 views
                    0 reactions
                    Last Post SEQadmin2  
                    Started by SEQadmin2, 07-24-2026, 12:17 PM
                    0 responses
                    19 views
                    0 reactions
                    Last Post SEQadmin2  
                    Started by SEQadmin2, 07-23-2026, 11:41 AM
                    0 responses
                    18 views
                    0 reactions
                    Last Post SEQadmin2  
                    Working...