Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • Gazaldeep
    Junior Member
    • Nov 2016
    • 6

    Illumina paired end adapter contamination problem

    Hello everyone!

    I have rna-seq Illumina paired end reads and want to proceed with adapter trimming.
    I have some confusions:

    1. Does the 5' end of both the forward and reverse reads start from the first base of the insert? Or could there be some adapter contamination also at 5' end?
    From whatever I have read online, there shouldn't be any adapter present at 5' end. But, the data I am analyzing has around 75 reads (out of 7 million for forward read file) with adapter at 5' end. 75 sequences isn't much, but I want to know what causes this..

    2. For the forward reads, some 3' ends may have indexed adapter. In cases where this indexed adapter occurs within the sequence, I should delete the adapter and the following sequence, right? Even if the indexed primer is present at 5' end?? In which case the whole read should be deleted. (Because this was due to absence of insert between two adapters)

    3. Do the 5' ends of reverse reads have barcode sequences or any part of the indexed adapter?? I have 12,399 reads (out of 7 million) that have complete or a part of indexed adapter at 5' end, with a few of them within the reads.


    I am new to rna-seq data analysis, and have gone through lots of tutorials and explanations online, but everything seems to be really confusing at this moment.

    My main concern is: where to expect adapters in illumina forward and reverse reads respectively, and what to do upon encountering unexpected adapters.
  • GenoMax
    Senior Member
    • Feb 2008
    • 7142

    #2
    Originally posted by Gazaldeep View Post
    Hello everyone!

    I have rna-seq Illumina paired end reads and want to proceed with adapter trimming.
    I have some confusions:

    1. Does the 5' end of both the forward and reverse reads start from the first base of the insert? Or could there be some adapter contamination also at 5' end?
    From whatever I have read online, there shouldn't be any adapter present at 5' end. But, the data I am analyzing has around 75 reads (out of 7 million for forward read file) with adapter at 5' end. 75 sequences isn't much, but I want to know what causes this..
    There should be no contamination on 5'-end if you are using standard Illumina kits.

    2. For the forward reads, some 3' ends may have indexed adapter. In cases where this indexed adapter occurs within the sequence, I should delete the adapter and the following sequence, right? Even if the indexed primer is present at 5' end?? In which case the whole read should be deleted. (Because this was due to absence of insert between two adapters)
    Barcodes/Tag reads are never part of the actual read in Illumina sequencing. If you have tags in your sequence then there is something wrong. If you have some reads with no inserts they should be taken care of during trimming.

    My main concern is: where to expect adapters in illumina forward and reverse reads respectively, and what to do upon encountering unexpected adapters.
    Use bbduk from BBMap suite. Search for that thread here. It is straight forward to use and @Brian includes all commercially used adapters in a file included in the package. Just point bbduk to that file and scan/trim your data.

    Comment

    • Gazaldeep
      Junior Member
      • Nov 2016
      • 6

      #3
      Thanks for your reply!!

      Originally posted by GenoMax View Post
      There should be no contamination on 5'-end if you are using standard Illumina kits.
      So, the 72 reads with 5' adapter contamination should be deleted, right?

      Originally posted by GenoMax View Post
      Barcodes/Tag reads are never part of the actual read in Illumina sequencing. If you have tags in your sequence then there is something wrong. If you have some reads with no inserts they should be taken care of during trimming.
      The paired-end data I am trying to analyze was downloaded from DDBJ.

      After searching online and through your answer, I'm sure that I should delete the reads that have any adapter at 5' end (be it the 5' adapter or 3' adapter), and perform trimming for reads with adapter at 3' end or within the read.

      But I'm actually a bit confused about the Illumina sequencing steps.

      Are the barcodes removed after sorting the reads into different files based on different barcodes?? So the files we get in the end cannot have the barcodes, but may they have the constant part of the indexed adapter (which occurs before/after the barcode) or are the constant parts also removed with the barcodes?
      I want to be clear about the process.

      Comment

      • Gazaldeep
        Junior Member
        • Nov 2016
        • 6

        #4
        I could just use a tool for trimming, but before that, I want to be clear about what's happening. Maybe I've got it all wrong?

        Comment

        • GenoMax
          Senior Member
          • Feb 2008
          • 7142

          #5
          Originally posted by Gazaldeep View Post
          Thanks for your reply!!

          But I'm actually a bit confused about the Illumina sequencing steps.
          Check this video out for clarification: https://www.youtube.com/watch?v=HMyCqWhwB8E

          Are the barcodes removed after sorting the reads into different files based on different barcodes?? So the files we get in the end cannot have the barcodes, but may they have the constant part of the indexed adapter (which occurs before/after the barcode) or are the constant parts also removed with the barcodes?
          I want to be clear about the process.
          Illumina sequencing actually proceeds in four separate steps (for 2D barcodes, 3 for 1 D barcodes).

          Code:
          R1 --> R2 (index 1) --> R3 (index 2) --> R4.
          Illumina software keeps tracks of every cluster over R1 through R4. During base calling (conversion of BCL to FASTQ) index read sequences are extracted from R2 (and R3) and are transferred to the header of the FASTQ record to complete demultiplexing (you thus end up with R1/R2 files).

          It is possible to generate files with index reads in individual files so you end up with 4 files per sample. This is only needed for some applications (e.g. QIIME).

          Comment

          • Gazaldeep
            Junior Member
            • Nov 2016
            • 6

            #6
            Thanks!!! Really helpful!!

            In my reads, I have 5' end contaminated with 5' adapter (75 reads). Also, in 12,000 reads out of 7 million, 5' adapter is present with the reads.. what do you suggest? Should I delete those reads? Or should I just trim the adapter and the sequence preceeding it at 5'?? I'm using Cutadapt at present. But in any adapter removal tool, I will have to specify if I want to trim these reads and in what way..

            Sorry if my questions are naive!

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              07-20-2026, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM
            • SEQadmin2
              Cancer Drug Resistance: The Lingering Barrier to Rising Survival
              by SEQadmin2



              Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

              There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
              07-08-2026, 05:17 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, Yesterday, 12:17 PM
            0 responses
            13 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-23-2026, 11:41 AM
            0 responses
            14 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-20-2026, 11:10 AM
            0 responses
            23 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-13-2026, 10:26 AM
            0 responses
            37 views
            0 reactions
            Last Post SEQadmin2  
            Working...