Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • tokikake
    Member
    • Nov 2011
    • 24

    #1

    strange mapping bias inside outside genes in prokayotic RNASeq

    Hi everyone,

    as it is mentioned in the title I have encountered a strange problem in a recent RNASeq experiment which I have never seen before.

    Some background:
    stranded RNASeq of prokaryotic mRNA; three samples (1 and 2 same condition; diff. harvest timepoint) with three replicates each; total RNA prep, DNAse digestion; rRNA depletion; library prep with dUTP method; Illumina sequencing; trimming; filtering of remaining tRNA/rRNA reads; mapping with bowtie2.

    After read counting, normalization and diff. expression analysis with Bayseq I had a strange bias towards one sample, which seems to dominate the rest. In the end I found out, that the summarized readcounts (of mapped reads inside genes) are in 6 out of 9 samples significantly lower and thus affecting the results of the diff. Expression analysis. I never had such a problem before and checked randomly ~6 other RNASeq projects and even if the input read numbers differ, the nomalized reads counts are equally distributed.

    To find out, what is the problem I counted all reads (sense and antisense reads inside and outside genes; see attachment). To my surprise there is a heavy bias towards the outside mapped genes in the problematic samples (2 and 3, all replicates).

    Does anyone has a clue what is the reason for this?

    I thought about DNA contamination, but then I would assume a equally distribution of reads sense and anti-sense. Or the dUTP library prep, because it is not sufficient for high GC organisms, but again there should be a equally distribution over the genome and of course over all samples. All samples, except 3.1 and 3.2, were treated (DNAse, depletion, library prep) in parallel.

    I would really appreciate any hint or suggestion!

    Tokikake
    Attached Files
  • bastianwur
    Member
    • Feb 2014
    • 98

    #2
    No clue what the biological cause could be.
    What are your genes in this case? Protein coding sequences, or also noncoding RNA? If noncoding is included, does it also include noncoding besides tRNA and rRNA?
    I ask because the rRNA read removal (under the assumption you don't do that by mapping do your reference, but with e.g. sortMeRNA) is not fully reliable (depending on how close your organism to anything in the database is), and you might still end up with a half ton of rRNA/tRNA. And if that's not the case, also tmRNA might have a very high expression value.

    And what does the fastQC report say?

    Comment

    • tokikake
      Member
      • Nov 2011
      • 24

      #3
      Hi bastian,

      thank you for your fast answer!
      The rRNA/tRNA reads were filtered via mapping against the reference (and not with sortMeRNA). But the tmRNA was a good point. I added it to my annotations (so genes in general not only protein coding ones) and found up to 13% more reads mapping into genes. I try now to annotate the RNAs with rfam and infernal and hope to reduce the amount of outside mapped reads.
      Nevertheless it is strange and I cannot see the same mapping bias in other projects, where I also didn't annotate the tmRNA (or any RNA prior to mapping, except tRNAs/rrNAs of course). Aber: Man lernt niemals aus!

      Have anyone else ever tested the amount of mapped reads inside/outside genes? It would be interesting to know!

      Comment

      Latest Articles

      Collapse

      • SEQadmin2
        Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
        by SEQadmin2



        CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

        Despite this, “CRISPR helped turn genome editing from a specialized technique into
        ...
        07-31-2026, 11:01 AM
      • SEQadmin2
        Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
        by SEQadmin2


        Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

        The systematic characterization of the human proteome has
        ...
        07-20-2026, 11:48 AM
      • SEQadmin2
        Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
        by SEQadmin2



        Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
        ...
        07-09-2026, 11:10 AM

      ad_right_rmr

      Collapse

      News

      Collapse

      Topics Statistics Last Post
      Started by SEQadmin2, 07-31-2026, 02:55 AM
      0 responses
      18 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-24-2026, 12:17 PM
      0 responses
      16 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-23-2026, 11:41 AM
      0 responses
      16 views
      0 reactions
      Last Post SEQadmin2  
      Started by SEQadmin2, 07-20-2026, 11:10 AM
      0 responses
      26 views
      0 reactions
      Last Post SEQadmin2  
      Working...