Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • MadsAlbertsen
    Member
    • Aug 2010
    • 26

    #1

    De novo assembly utilizing abundance information from multiple samples

    In my lab we do loads of metagenome sequencing and assembly to recover complete genomes from environmental samples. To bin the genomes we use abundance information from multiple related samples and in general it is very easy to extract near complete genomes IF the assembly is decent.

    The big problem is micro-diversity, e.g. closely related species, which prevents nice assemblies. The problematic similarity between species seem to be approximately 94-99% average nucleotide identity (dependent on kmer choice etc..). Anything else we can get decent assemblies from if we just have enough coverage.

    Some metagenome assemblers try to deal with this during assembly (e.g. IDBA-UD). But I havn't really tried anything that has decent performance yet.

    Do anyone know of de novo assemblers that would be able to use the abundance information from multiple related metagenome samples directly in the assembly stage?

    Or alternatively something like khmer (https://khmer.readthedocs.org/en/latest/) that would allow read splitting prior to assembly - BUT using kmer abundance information from multiple samples?

    rgds
    Mads Albertsen
  • seb567
    Senior Member
    • Jul 2008
    • 260

    #2
    Re: assemble multiple related metagenome samples

    Originally posted by MadsAlbertsen View Post
    In my lab we do loads of metagenome sequencing and assembly to recover complete genomes from environmental samples. To bin the genomes we use abundance information from multiple related samples and in general it is very easy to extract near complete genomes IF the assembly is decent.

    The big problem is micro-diversity, e.g. closely related species, which prevents nice assemblies. The problematic similarity between species seem to be approximately 94-99% average nucleotide identity (dependent on kmer choice etc..). Anything else we can get decent assemblies from if we just have enough coverage.

    Some metagenome assemblers try to deal with this during assembly (e.g. IDBA-UD). But I havn't really tried anything that has decent performance yet.

    Do anyone know of de novo assemblers that would be able to use the abundance information from multiple related metagenome samples directly in the assembly stage?
    Hello,


    Why do you want to assemble multiple related metagenome samples together ?

    If you look at Cortex ( http://www.nature.com/ng/journal/v44...l/ng.1028.html ),
    each vertex in their de Bruijn subgraph has many coverage depth channels.

    For metagenomics samples, one sample already contains a mix of different genomes in various abundances. Mixing these samples would lead, I believe, to a meta-metagenome (whatever that means).


    If you want to assemble (and possibly profile) many samples individually,
    you should try Ray (the "Ray Meta" workflow) for denovo genome assembly of metagenomes.

    To download Ray (v2.2.0): http://denovoassembler.sourceforge.net/
    See the paper also: http://genomebiology.com/2012/13/12/R122


    Regarding multi-sample metagenomic analyses, we have a project in progress called
    "Ray Surveyor" which basically build a distributed de Bruijn subgraph for many samples
    like Cortex, but Ray engine is distributed (message passing) and this "Ray Surveyor" uses
    the actor model (this is new too !).


    Best luck to you in your research !

    Originally posted by MadsAlbertsen View Post
    Or alternatively something like khmer (https://khmer.readthedocs.org/en/latest/) that would allow read splitting prior to assembly - BUT using kmer abundance information from multiple samples?

    rgds
    Mads Albertsen

    Comment

    • MadsAlbertsen
      Member
      • Aug 2010
      • 26

      #3
      Thanks for the suggestions. I have been trying some of your Ray related projects . Cortex seem to be in the direction I was thinking of.

      Originally posted by seb567 View Post
      Why do you want to assemble multiple related metagenome samples together ?

      For metagenomics samples, one sample already contains a mix of different genomes in various abundances. Mixing these samples would lead, I believe, to a meta-metagenome (whatever that means).
      I want to use multiple related metagenome samples to untangle individual species from metagenomes. I'm not interested in the metagenome as is. Only in extracting complete genomes.

      It's already common to use abundance profiles for assembled scaffolds to do binning (scaffolds with similar abundance patterns originate from the same species) and also in clustering genes with similar expression profiles.

      Hence, if the abundance information could be used in the assembly process directly it might lead to decent assembly of species with many closely related strains, which currently is impossible.

      rgds
      Mads
      Last edited by MadsAlbertsen; 10-24-2013, 03:18 AM.

      Comment

      • seb567
        Senior Member
        • Jul 2008
        • 260

        #4
        Originally posted by MadsAlbertsen View Post
        Thanks for the suggestions. I have been trying some of your Ray related projects . Cortex seem to be in the direction I was thinking of.



        I want to use multiple related metagenome samples to untangle individual species from metagenomes. I'm not interested in the metagenome as is. Only in extracting complete genomes.

        It's already common to use abundance profiles for assembled scaffolds to do binning (scaffolds with similar abundance patterns originate from the same species) and also in clustering genes with similar expression profiles.

        Hence, if the abundance information could be used in the assembly process directly it might lead to decent assembly of species with many closely related strains, which currently is impossible.

        rgds
        Mads
        That's a good idea, but the coverage depth for each sample kmers needs to normalized with the number of reads since sample A may have twice the number of reads present in sample B.

        Coverage depth would be measured in X per reads, for example.

        Comment

        Latest Articles

        Collapse

        • SEQadmin2
          Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
          by SEQadmin2



          CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

          Despite this, “CRISPR helped turn genome editing from a specialized technique into
          ...
          07-31-2026, 11:01 AM
        • SEQadmin2
          Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
          by SEQadmin2


          Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

          The systematic characterization of the human proteome has
          ...
          07-20-2026, 11:48 AM

        ad_right_rmr

        Collapse

        News

        Collapse

        Topics Statistics Last Post
        Started by SEQadmin2, 08-06-2026, 07:41 AM
        0 responses
        14 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 08-03-2026, 10:13 AM
        0 responses
        31 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-31-2026, 02:55 AM
        0 responses
        40 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-24-2026, 12:17 PM
        0 responses
        26 views
        0 reactions
        Last Post SEQadmin2  
        Working...