Seqanswers Leaderboard Ad

**Jon_Keats** · 01-26-2011, 03:40 PM

How do you define the inner-mate-distance? I think the properly paired is based on the data you provide so (r- 125 with default STD of 20bp) if the inner mate distance is really 180bp with STD of 32 bp few of your reads will be called as "properly paired" as their true inner-mate-distance is outside of the mean/STD range you provided. I've been using a 5 million read subset against a transcriptome reference with BWA to get a data based range I feed into tophat.

**AdamB** · 01-27-2011, 03:08 AM

I have previously mapped the reads using Bioscope (ABI), and that gave me the mate pair distance stats. I just tried TopHat using 125±50 bp, which I thought more accurately reflected the spread, but the number of 'properly paired' reads is almost the same. From my Bioscope mapping stats, there are approximately 4 millions reads with a mate pair distance of 75-175 bp, yet only 5,602 (0.14%) are 'properly paired'. It seems there is a glaring error somewhere?

**AdamB** · 01-27-2011, 04:00 AM

Also if anyone has tips on how to extract a subset of matching F3 and F5 reads from .csfasta and .qual files, please let me know.

**AdamB** · 01-31-2011, 03:54 AM

If anyone has an answer to this question it would be much appreciated, thanks.

**AdamB** · 02-03-2011, 07:53 AM

An update:

I realised there are not the same number of F3 reads as F5-BC reads:
F3 = 23,077,379
F5-BC = 26,929,508

So, for a sample of 250,000 reads, I extracted those reads for which there was both F3 and F5-BC present, and mapped with TopHat.

250,000 reads > samtools flagstat

Code:

186066 in total
0 QC failure
0 duplicates
186066 mapped (100.00%)
186066 paired in sequencing
70366 read1
115700 read2
96 properly paired (0.05%)
7190 with itself and mate mapped
178876 singletons (96.14%)
0 with mate mapped to a different chr
0 with mate mapped to a different chr (mapQ>=5)

230,999 reads (present in F3 and F5-BC) > samtools flagstat

Code:

212529 in total
0 QC failure
0 duplicates
212529 mapped (100.00%)
212529 paired in sequencing
85982 read1
126547 read2
55096 properly paired (25.92%)
69488 with itself and mate mapped
143041 singletons (67.30%)
0 with mate mapped to a different chr
0 with mate mapped to a different chr (mapQ>=5)

There are now 26% 'properly paired' reads, and more mapped reads.

**chenyao** · 08-17-2011, 02:59 AM

Originally posted by AdamB View Post

An update:

I realised there are not the same number of F3 reads as F5-BC reads:
F3 = 23,077,379
F5-BC = 26,929,508

So, for a sample of 250,000 reads, I extracted those reads for which there was both F3 and F5-BC present, and mapped with TopHat.

250,000 reads > samtools flagstat

Code:

186066 in total
0 QC failure
0 duplicates
186066 mapped (100.00%)
186066 paired in sequencing
70366 read1
115700 read2
96 properly paired (0.05%)
7190 with itself and mate mapped
178876 singletons (96.14%)
0 with mate mapped to a different chr
0 with mate mapped to a different chr (mapQ>=5)

230,999 reads (present in F3 and F5-BC) > samtools flagstat

Code:

212529 in total
0 QC failure
0 duplicates
212529 mapped (100.00%)
212529 paired in sequencing
[COLOR="DarkOrange"]85982 read1
126547 read2[/COLOR]
55096 properly paired (25.92%)
69488 with itself and mate mapped
143041 singletons (67.30%)
0 with mate mapped to a different chr
0 with mate mapped to a different chr (mapQ>=5)

There are now 26% 'properly paired' reads, and more mapped reads.

Adam,Do you find the problem ,I have the same problem. Did you notice the number of read1 and read2 differ a lot?

**marcowanger** · 08-17-2011, 04:01 AM

Just random thought,

Is it transcriptome reads?

How divergent is your genome reference to the sample you used for SOLiD?

Is the ABI inner mate distance the externa or internall insert size? Top hat defines inser size as the inner part encompassed by 2 reads.

**chenyao** · 08-17-2011, 04:16 AM

Originally posted by marcowanger View Post

Just random thought,

Is it transcriptome reads?

How divergent is your genome reference to the sample you used for SOLiD?

Is the ABI inner mate distance the externa or internall insert size? Top hat defines inser size as the inner part encompassed by 2 reads.

The reference and my sample are all mouse.

the detail of my problem is here,

http://seqanswers.com/forums/showthread.php?t=13407,

Do you have any suggestion.

**tintin306** · 03-08-2012, 08:30 AM

you exp is strand specific rna-seq since you use --library-type?..

Topics	Statistics	Last Post
Cancer Metastasis: A Deep Dive into Cellular Plasticity by seqadmin Started by seqadmin, 04-11-2024, 12:08 PM	0 responses 22 views 0 likes	Last Post by seqadmin 04-11-2024, 12:08 PM
Proteogenomic Profiles Offer New Clues in Prostate Cancer by seqadmin Started by seqadmin, 04-10-2024, 10:19 PM	0 responses 24 views 0 likes	Last Post by seqadmin 04-10-2024, 10:19 PM
Novel Diagnostic Assay Enhances Ovarian Cancer Detection by seqadmin Started by seqadmin, 04-10-2024, 09:21 AM	0 responses 19 views 0 likes	Last Post by seqadmin 04-10-2024, 09:21 AM
Evolutionary Dynamics of Centromeres: A Comparative Genomic Analysis by seqadmin Started by seqadmin, 04-04-2024, 09:00 AM	0 responses 50 views 0 likes	Last Post by seqadmin 04-04-2024, 09:00 AM

Seqanswers Leaderboard Ad

Announcement

'Properly paired' reads in sam flag from TopHat mapping

Comment

Comment

Comment

Comment

Comment

Comment

Comment

Comment

Comment

Latest Articles

ad_right_rmr

News