Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • gfmgfm
    Member
    • Jun 2010
    • 64

    #1

    TopHat and cufflinks of non-strand-specific reads

    Hi all,

    I have questions regarding the strand info after TopHat and Cufflinks in NON-strand-specific experiments. I tried to look for the answers in the forum and in the web – but did not manage, so I am posting it here.

    I have Illumina RNA seq reads, from a library that was prepared with NON-strand specific protocol.

    1. I saw in tophat manual that TopHat will treat reads as strand specific. What option should I use when running TopHat if my reads are not strand specific?

    2. After TopHat – in the bam files – Will I get strand for all reads? Or only for junction read?
    In the bam file – will I have a strand according to the splice junction orientation or according the actual strand the read was mapped to?

    3. Cufflinks gives “a guess” to the strand of the transcript. How is this guess made?

    Thanks a lot.
  • RFo
    Junior Member
    • May 2012
    • 5

    #2
    Hi gfmgfm,
    did you manage to get an answer yet to your question? I'm especially interested in number one... I use Cufflinks for expression analysis and I'm not sure, whether it will use reads from both strands to calculate the FPKM value or just the reads which are on the sense-strand...

    Thanks

    Comment

    • dvanic
      Member
      • Jan 2012
      • 61

      #3
      From the manual
      --library-type TopHat will treat the reads as strand specific. Every read alignment will have an XS attribute tag. Consider supplying library type options below to select the correct RNA-seq protocol.
      If you look below at the possible flags:
      Library Type Examples Description
      fr-unstranded Standard Illumina Reads from the left-most end of the fragment (in transcript coordinates) map to the transcript strand, and the right-most end maps to the opposite strand.
      fr-firststrand dUTP, NSR, NNSR Same as above except we enforce the rule that the right-most end of the fragment (in transcript coordinates) is the first sequenced (or only sequenced for single-end reads). Equivalently, it is assumed that only the strand generated during first strand synthesis is sequenced.
      fr-secondstrand Ligation, Standard SOLiD Same as above except we enforce the rule that the left-most end of the fragment (in transcript coordinates) is the first sequenced (or only sequenced for single-end reads). Equivalently, it is assumed that only the strand generated during second strand synthesis is sequenced.
      I.e. even though the manual is saying that Tophat will treat all reads as stranded, the first option is giving you a strand-unaware alignment.

      After TopHat – in the bam files – Will I get strand for all reads? Or only for junction read?
      You will get a strand for all reads. Even if a read does not go over a splice junction, it can usually be uniquely mapped to the + or - strand of DNA that the RNA came from, i.e.

      DNA
      (+) ATGCCGAGAGAGAGTTCAGAGAGATTCG
      (-) TACGGCTCTCTCTCAAGTCTCTCTAAGC

      Read
      GCCGAGAGAG

      Will map to + strand only, and be reported to you as +
      In the bam file – will I have a strand according to the splice junction orientation or according the actual strand the read was mapped to?
      Hopefully this will be the same, since the splice site should be in the same "direction" as your actual read.

      3. Cufflinks gives “a guess” to the strand of the transcript. How is this guess made?
      Look at the information on the cufflinks manual/how it works page. But, as mentioned a lot on this forum, cufflinks is very far from perfect, and I would strongly recommend you run it giving it a reference annotation.

      Comment

      • RFo
        Junior Member
        • May 2012
        • 5

        #4
        Thanks dvanic for your reply.
        So is it right that the '--library-type' option is used only in tophat for the junction finding and has no effect on cufflinks and cuffdiff (even though they have the same option)?

        I analyzed my strand-specific SOLiD reads (2 conditions having 3 replicates each, total around 100M paired-end reads of a fungi) once using fr-secondstrand and once using fr-unstranded, and the only difference I got are some minor variations on the junction positions.

        I would have expected that using the "fr-secondstrand" option, only reads which align to the sense strand of the gene/transcript would be taken into account for the expression calculation. However, this doesn't seem to be the case because, because there are no major differences between "fr-secondstrand" and "fr-unstranded". This is true even for certain genes which have only anti-sense reads aligned to them.

        It just seems that there is not much use in using stranded information, or am I wrong?

        Comment

        Latest Articles

        Collapse

        • SEQadmin2
          Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
          by SEQadmin2



          CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

          Despite this, “CRISPR helped turn genome editing from a specialized technique into
          ...
          07-31-2026, 11:01 AM
        • SEQadmin2
          Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
          by SEQadmin2


          Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

          The systematic characterization of the human proteome has
          ...
          07-20-2026, 11:48 AM

        ad_right_rmr

        Collapse

        News

        Collapse

        Topics Statistics Last Post
        Started by SEQadmin2, 08-13-2026, 12:22 PM
        0 responses
        23 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 08-11-2026, 10:35 AM
        0 responses
        19 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 08-06-2026, 07:41 AM
        0 responses
        33 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 08-03-2026, 10:13 AM
        0 responses
        51 views
        0 reactions
        Last Post SEQadmin2  
        Working...