Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • cormicp
    Junior Member
    • Jul 2009
    • 5

    Calling SNPS following Tophat alignment of RNA-SEQ reads

    Hi all,

    We have been assessing different methods of calling SNPs from our RNA-seq runs ranging from using the GA pipeline itself, to using MAQ to call the SNPs (using MAQ aligner to align the reference genome as well as a set of set of splice junction sequences derived from all known genes). Both these methods gave broadly similar results, which was great!

    However I have been interested in using Tophat to align our sequences as it would allow us to investigate novel transcripts or splice junctions not in the currently annotated gene set. The alignment using Tophat appeared to work fine, aside from seeming to output the quality scores for each aligned read in Solexa rather than the phred format I was expecting (using the --solexa1.3-quals option). Using cufflinks to quantitate transcript abundance gave results again broadly similar to the GA analysis pipeline, which again was super.

    However using Samtool to call SNPs in the Tophat aligned reads resulted in over four times as many SNPs being called as by the other two methods. Just eyeballing the alignments indicates that most of the new Samtools-called SNPs are called at the beginning and end of exons, where fewer reads have aligned back or where reads have been split between two adjacent exons. Using Samtools to call SNPs on the alignments previously generated by MAQ resulted in only marginally more SNPs called than by the MAQ SNP caller itself, leading us to believe that the problem is with the Tophat alignment rather than the SNP caller.

    All this may be a result of something stupid I am doing wrong but I figured there is no harm in finding out if others out there had encountered similar problems and if you have, how you dealt with them? I don't simply want to discard SNPs called at the beginning or end of exons for obvious reasons!

    Thanks for any help anyone can give me on this
  • shurjo
    Senior Member
    • Jan 2009
    • 132

    #2
    Perhaps it might help if you increased the anchor length specification (-a/--min-anchor-length) for Tophat during the initial alignment? From my experience, Illumina RNA-Seq reads have increased error rates for the first 10-12 bases and then again towards the end, so using a longer anchor (and being relatively strict for the number of mismatches in the -m/--splice-mismatches setting) may help to keep the splice junction alignments more "real" and result in lesser but truer SNP calls.

    Comment

    • bioinfosm
      Senior Member
      • Jan 2008
      • 483

      #3
      Why would you think that RNA-Seq reads have higher error rates for first 10-12 bases?
      --
      bioinfosm

      Comment

      • shurjo
        Senior Member
        • Jan 2009
        • 132

        #4
        Based on the error plots generated by our installation of the Illumina pipeline, that is what it looks like (I am attaching two pictures from a RNA and DNA lane: the RNA graph is the one with the bump at the beginning and then again around cycle 50 at the beginning of the second read).
        Attached Files

        Comment

        • edge
          Senior Member
          • Sep 2009
          • 199

          #5
          Hi cormicp,

          Do you figure out the solution for your doubt?
          Currently I'm facing the same problem as well.
          I have a Illumina RNA-seq pair-end read, reference transcriptome.
          However, I have no idea how to get the SNP result from my data set.
          Thanks for any advice.
          Last edited by edge; 05-30-2012, 09:19 AM. Reason: typo error

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
            by SEQadmin2


            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

            The systematic characterization of the human proteome has
            ...
            07-20-2026, 11:48 AM
          • SEQadmin2
            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
            by SEQadmin2



            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
            ...
            07-09-2026, 11:10 AM
          • SEQadmin2
            Cancer Drug Resistance: The Lingering Barrier to Rising Survival
            by SEQadmin2



            Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

            There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
            07-08-2026, 05:17 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, 07-24-2026, 12:17 PM
          0 responses
          26 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-23-2026, 11:41 AM
          0 responses
          20 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-20-2026, 11:10 AM
          0 responses
          28 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-13-2026, 10:26 AM
          0 responses
          38 views
          0 reactions
          Last Post SEQadmin2  
          Working...