Hi,
Just got my first testdata from a RRBS library prepared with MspI digestion with the NEXTFLEX® Bisulfite Library Prep Kit for Illumina, and I am have some questions.
The library was sequenced paired-end on Illumina Miseq. I was expecting that most of my reads had remnants of the MspI cutsite at the start, so I wrote a little script to check that. The expectation is that forward reads should start with 'CGG' or 'TGG', if the first 'C' was methylated or not, respectively. For the reverse reads I was expecting mainly 'CAA', but also a few 'CGA' (unmethylated cytosines were used for end-repair, and I am assuming that bisulfite conversion rate is not 100%), as well as a few 'CAG' or 'CGG' (theoretically possible).
So, ideally I should have the following combinations (forward - reverse):
CGG-CAA
TGG-CAA
CGG-CGA
TGG-CGA
CGG-CAG
TGG-CAG
CGG-CGG
TGG-CGG
My expections were met in the way that CGG-CAA and TGG-CAA was by far the most common, ~20% of read pairs and ~6% of read pairs, respectively. The other combinations account for less than 1% in total. What puzzles me though is that I only get ~30% read pairs that have any of the above patterns.
I figured that perhaps some of the fragments I get after the MspI digestion are somehow sheared during the library prep and loose the cutsite on one side, so I relaxed the search criteria to also count read pairs that have the expected pattern only in the forward read OR in the reverse read. An additional ~28% of the read pairs fell into this category.
Which leaves me with >40% of read pairs that don't have any of the expected patterns, neither in the forward nor the reverse read.
So, my question: Is this normal? Sure there will be some sequencing error, but the read quality is very good overall and the error rate should be very low at the start of the reads, so that can't affect 40% reads. Am I missing something here?
I'd really appreciate if anyone could share their experience!
cheers,
Christoph
Just got my first testdata from a RRBS library prepared with MspI digestion with the NEXTFLEX® Bisulfite Library Prep Kit for Illumina, and I am have some questions.
The library was sequenced paired-end on Illumina Miseq. I was expecting that most of my reads had remnants of the MspI cutsite at the start, so I wrote a little script to check that. The expectation is that forward reads should start with 'CGG' or 'TGG', if the first 'C' was methylated or not, respectively. For the reverse reads I was expecting mainly 'CAA', but also a few 'CGA' (unmethylated cytosines were used for end-repair, and I am assuming that bisulfite conversion rate is not 100%), as well as a few 'CAG' or 'CGG' (theoretically possible).
So, ideally I should have the following combinations (forward - reverse):
CGG-CAA
TGG-CAA
CGG-CGA
TGG-CGA
CGG-CAG
TGG-CAG
CGG-CGG
TGG-CGG
My expections were met in the way that CGG-CAA and TGG-CAA was by far the most common, ~20% of read pairs and ~6% of read pairs, respectively. The other combinations account for less than 1% in total. What puzzles me though is that I only get ~30% read pairs that have any of the above patterns.
I figured that perhaps some of the fragments I get after the MspI digestion are somehow sheared during the library prep and loose the cutsite on one side, so I relaxed the search criteria to also count read pairs that have the expected pattern only in the forward read OR in the reverse read. An additional ~28% of the read pairs fell into this category.
Which leaves me with >40% of read pairs that don't have any of the expected patterns, neither in the forward nor the reverse read.
So, my question: Is this normal? Sure there will be some sequencing error, but the read quality is very good overall and the error rate should be very low at the start of the reads, so that can't affect 40% reads. Am I missing something here?
I'd really appreciate if anyone could share their experience!
cheers,
Christoph
Comment