Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • andreiafonseca
    Junior Member
    • Mar 2008
    • 9

    #1

    expected percentage of mapped reads in chip-seq experiment

    Dear all,


    I am starting to work with Chip-seq data to identify binding sites of transcription factors and we have just received a test run from the company in order to decide to go ahead with the sequencing. We are using Hiseq. For a test run we got 100,000 reads of each library. We have sequenced for each sample the correspondent input. For the samples 40-50% reads mapped to the human genome (allowing max 2 mismaches) whereas the input had a percentage of mapped reads ~90% . I was expecting to obtain a lower percentage of mapped reads in the input and not in the sample. I used bowtie to make the alignment. I would like to know if someone as had similar percentages of mapped reads in Chip-seq experiments.
    Thanks

    Andreia
  • kopi-o
    Senior Member
    • Feb 2008
    • 319

    #2
    This could depend on a lot of factors but 40-50% sounds low on the face of it. How does the quality score distribution of the ChIPped sequences look compared to the input? What read lengths are you using?

    Comment

    • andreiafonseca
      Junior Member
      • Mar 2008
      • 9

      #3
      Thanks for answering! The read length is 49 bp. What do you mean by the quality score distibution? I have checked the average of the QS per position on the read and the quality score distribution is similar between the sample and the input. In the 5' end we have an average ~38 and at the 3'end ~34.

      Comment

      • andreiafonseca
        Junior Member
        • Mar 2008
        • 9

        #4
        one more detail is single-end sequencing

        Comment

        • kopi-o
          Senior Member
          • Feb 2008
          • 319

          #5
          I was thinking that maybe the quality scores would be lower at the 3' end for the sample, but that doesn't appear to be the case. I'm not sure what the explanation could be then, maybe something in the sample preparation?

          Comment

          • andreiafonseca
            Junior Member
            • Mar 2008
            • 9

            #6
            it could happen that the enrichment was far from perfect, but why does the input has such a high proportion of mapped reads? shouldn't the input have a lower proportion because it is genomic DNA, so it has repeats, telomeric regions which will map to many locations?

            Comment

            • kopi-o
              Senior Member
              • Feb 2008
              • 319

              #7
              Well, maybe ... but input DNA sequencing also does not give an unbiased representation of the genome; open-chromatin regions like TSS are overrepresented there too, see e g


              Comment

              • ttnguyen
                Member
                • Mar 2010
                • 41

                #8
                Just wondering "mapped reads" are unique (I mean set -m 1 in Bowtie)? If so, probably 90% of mapped reads (only 49 bp in length) seems so high as ~50% of the human genome are masked by RepeatMasker?

                BTW, It would be helpful to see the distribution of sequence quality by using fastQC:

                Comment

                • andreiafonseca
                  Junior Member
                  • Mar 2008
                  • 9

                  #9
                  thanks for your message. I took a bit because I was checking how many reads were uniquely mapped. In samples ~40-50% were uniquely mapped and in input ~80%. In attachment you can see the distribution of the QS, Lib1 is a sample and Lib2 is the corresponding input.
                  thanks for the help

                  Comment

                  • ttnguyen
                    Member
                    • Mar 2010
                    • 41

                    #10
                    I could not find the attachment . I think 40-50% in samples is normal, but 80% in input is quite high. Would it be worth marking PCR duplicate, or comparing your mapping rates with the public datasets (e.g. ENCODE - I think they had Input as well)?

                    Comment

                    • andreiafonseca
                      Junior Member
                      • Mar 2008
                      • 9

                      #11
                      sorry just noticed there was a problem in the attachment
                      Attached Files

                      Comment

                      • andreiafonseca
                        Junior Member
                        • Mar 2008
                        • 9

                        #12
                        this image is for the sample the previous one was for input
                        Attached Files

                        Comment

                        • andreiafonseca
                          Junior Member
                          • Mar 2008
                          • 9

                          #13
                          can you explain me what do you mean by marking PCR duplicate?

                          Comment

                          • ttnguyen
                            Member
                            • Mar 2010
                            • 41

                            #14
                            This is very good QS I think as the average is high and the variation is low (even though at 3' end). Could you show your Bowtie command?

                            Comment

                            • andreiafonseca
                              Junior Member
                              • Mar 2008
                              • 9

                              #15
                              -f -a --best --strata -v 2 hg19 fasta

                              then I selected from these the unique alignments

                              can you tell me what do you mean by marking PCR duplicate?

                              Comment

                              Latest Articles

                              Collapse

                              • SEQadmin2
                                Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                                by SEQadmin2



                                CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                                Despite this, “CRISPR helped turn genome editing from a specialized technique into
                                ...
                                07-31-2026, 11:01 AM
                              • SEQadmin2
                                Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                                by SEQadmin2


                                Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                                The systematic characterization of the human proteome has
                                ...
                                07-20-2026, 11:48 AM
                              • SEQadmin2
                                Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                                by SEQadmin2



                                Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                                ...
                                07-09-2026, 11:10 AM

                              ad_right_rmr

                              Collapse

                              News

                              Collapse

                              Topics Statistics Last Post
                              Started by SEQadmin2, 08-03-2026, 10:13 AM
                              0 responses
                              21 views
                              0 reactions
                              Last Post SEQadmin2  
                              Started by SEQadmin2, 07-31-2026, 02:55 AM
                              0 responses
                              34 views
                              0 reactions
                              Last Post SEQadmin2  
                              Started by SEQadmin2, 07-24-2026, 12:17 PM
                              0 responses
                              25 views
                              0 reactions
                              Last Post SEQadmin2  
                              Started by SEQadmin2, 07-23-2026, 11:41 AM
                              0 responses
                              21 views
                              0 reactions
                              Last Post SEQadmin2  
                              Working...