Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • lg36
    Member
    • Mar 2012
    • 12

    #1

    Which Pipeline is Correct??

    Dear All,

    I'd be very grateful for any advice on which (if any) of the following processes to take for SNP discovery? I find a different number of SNPs with each approach with the same SNP filters applied throughout.

    Raw FastQ File --> Direct Mapping to Reference --> SNP discovered 1,200

    Raw FastQ File --> De Novo Assembly --> Extract Paired Contigs --> Map Paired Contigs to Reference --> SNPs discovered 1,089

    Raw FastQ File --> De Novo Assembly --> Extract All Contigs --> Map Contigs to Reference --> SNPs discovered 1,383

    How can I know which SNPs are correct. Would it be useful to test the quality of the corresponding consensus sequence (for sets of known nucleotides?) or would that not help to judge which is most accurate.

    Best wishes lg36
  • westerman
    Rick Westerman
    • Jun 2008
    • 1104

    #2
    It does depend on the quality of your reference but I would say the first. I see no reason to do a de-novo assembly, especially for SNP discovery, when you can do an mapping instead. De-novo will always give worse and more questionable results.

    As for how to determine which ones are correct ... back to the lab you go! Independent verification via different methodology is the definitive proof. Oh, you can also take your sequencing results and apply statistical filters to it and list the SNPs that fall into, say, p<0.05 but where is the fun in that?

    Comment

    • lh3
      Senior Member
      • Feb 2008
      • 686

      #3
      Mapping based approach may be confused by long indels or large-scale changes, which leads to false SNPs. This is not that infrequent for human SNP discovery. Assembly can do better in such cases as it more effectively takes advantage of between-read information.

      On the other hand, although I believe for small genomes, assembly based approach is advantageous in theory, many existing assemblers and contig aligners are not fine tuned for assembly based SNP discovery. On Illumina data, for which the tool chain is relatively complete and mature, the overall accuracy of mapping based calls is likely to be better unless you are very careful about the assembly.

      Anyway, which is better highly depends on how you did the analysis and the divergence from your reference. We cannot just tell from the numbers. I recommend you look at calls unique to one set in IGV/tview and get a sense by yourself. This is the cheapest yet very effective way to answer your own question.

      Comment

      • lg36
        Member
        • Mar 2012
        • 12

        #4
        Thanks both so much for your help. Just to give you some more information, the second method gives me a consensus sequence which is most representative of the consensus sequence we have already PCR'd in the lab via a different method. Does this make method 2 more correct than method one or three.

        Comment

        • milo0615
          Member
          • Dec 2012
          • 39

          #5
          hi g36,

          Which tools did you use for mapping and SNP discovery?

          Originally posted by lg36 View Post
          Dear All,

          I'd be very grateful for any advice on which (if any) of the following processes to take for SNP discovery? I find a different number of SNPs with each approach with the same SNP filters applied throughout.

          Raw FastQ File --> Direct Mapping to Reference --> SNP discovered 1,200

          Raw FastQ File --> De Novo Assembly --> Extract Paired Contigs --> Map Paired Contigs to Reference --> SNPs discovered 1,089

          Raw FastQ File --> De Novo Assembly --> Extract All Contigs --> Map Contigs to Reference --> SNPs discovered 1,383

          How can I know which SNPs are correct. Would it be useful to test the quality of the corresponding consensus sequence (for sets of known nucleotides?) or would that not help to judge which is most accurate.

          Best wishes lg36

          Comment

          • jorge-bariloche
            Junior Member
            • Oct 2014
            • 4

            #6
            Hi lg36
            I'm sort of having the same doubt. How did you solve this issue?
            best wishes
            Jorge

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
              by SEQadmin2



              CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

              Despite this, “CRISPR helped turn genome editing from a specialized technique into
              ...
              07-31-2026, 11:01 AM
            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              07-20-2026, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, 08-03-2026, 10:13 AM
            0 responses
            17 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-31-2026, 02:55 AM
            0 responses
            33 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-24-2026, 12:17 PM
            0 responses
            23 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-23-2026, 11:41 AM
            0 responses
            21 views
            0 reactions
            Last Post SEQadmin2  
            Working...