Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • bloosnail
    Member
    • Jul 2015
    • 17

    Minimum amount of data needed for reliable results?

    We are trying to do analysis for whole genome metagenomic data taken from the surface of the eye. Each sample has millions of reads generated, but of those reads at most only 1-2% are bacterial reads. We are wondering if there is some information/resources about the amount of data available related to the reliability of the results eg. finding out the taxonomic information for bacteria down to species level that are greater than 1% relative abundance. Currently we are aligning the data to whole genome bacterial sequences, but there are many multi-mapping locations which many of which may be false positives. We have tried using Metaphlan2 to do alignment which uses a custom catalog of unique markers for different clades, but usually only several hundred reads will be mapped back -- many of the samples report very low/no species present. Specifically, we are wondering methods to do analysis for whole genome metagenomic sequences where the amount of data is very low. Any help is greatly appreciated.

    Daniel
  • Brian Bushnell
    Super Moderator
    • Jan 2014
    • 2709

    #2
    You might try removing human sequence, then assembling the rest and BLASTing the contigs against nt/nr/RefSeq microbial. Assuming the contigs are longer than read length, they will give you more reliable hits. What kind of depth do you have for the bacteria? You can find that out with a kmer-frequency histogram, after human reads are removed.

    Comment

    • bloosnail
      Member
      • Jul 2015
      • 17

      #3
      Thank you for the quick response. The idea of assembling the reads into contigs before alignment makes sense, I will let me supervisor know. Do you know of good software to do this? I have tried Velvet in the past but did not use it extensively.

      I forgot to mention that we have removed human sequences, although the revised reference genome that you created seems like it would be especially useful for us.

      Could you give more information on how to estimate the depth of the bacteria? There is generally less than 100,000 bacterial reads per sample out of 20-30 million initial reads (before any trimming/contaminant removal).

      Comment

      • Brian Bushnell
        Super Moderator
        • Jan 2014
        • 2709

        #4
        I suggest Spades or Megahit for metagenome assembly. 100k is not many reads; you might not have sufficient depth for assembly. But in that case, you may get a better assembly by combining all bacterial reads from all samples and assembling together. Then you can quantify by mapping to the combined assembly.

        For human removal, the raw human genome is fine in your case (bacteria). The masked version is mainly to allow decontamination of eukaryotes, which have shared sequence with human; bacteria basically don't.

        Comment

        • gringer
          David Eccles (gringer)
          • May 2011
          • 845

          #5
          You can create rarefaction curves to see if what you have is likely sufficient to describe the metagenomic profile.

          The basic process is to remove reads and see if your calculation of the species diversity is similar. A low complexity sample will plateau at a low coverage, while the diversity of a high complexity sample will just keep increasing substantially with more reads.

          Comment

          • dhtaft
            Junior Member
            • Dec 2016
            • 3

            #6
            I had some luck using IMSA in a similar situation to the one you describe, but only after human read removal

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              07-20-2026, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM
            • SEQadmin2
              Cancer Drug Resistance: The Lingering Barrier to Rising Survival
              by SEQadmin2



              Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

              There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
              07-08-2026, 05:17 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, 07-24-2026, 12:17 PM
            0 responses
            16 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-23-2026, 11:41 AM
            0 responses
            17 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-20-2026, 11:10 AM
            0 responses
            23 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-13-2026, 10:26 AM
            0 responses
            37 views
            0 reactions
            Last Post SEQadmin2  
            Working...