Seqanswers Leaderboard Ad

**clk** · 09-28-2012, 06:45 AM

Hi,
looking for an answer to the same question I came across your post... Did you ever find an answer to this question? The discrepancies I see between IGV and samtools mpileup are huge in my data, which is RNAseq data...

Thanks!

**liu_xt005** · 10-01-2012, 11:56 AM

Sorry I did not find a good solution.
But my study showed that SAMtools tend to keep longer tails for indels than GATK and others. For example, SAMtools gives TAAAA:TAAA (REF:ALT), while GATK gives TA:T.

**clk** · 10-01-2012, 12:07 PM

Thanks for responding. I finally found a solution to this, so I'll post it here in case is useful for others.

I found out that samtools filters reads before including them in the pileup; it reads the flag field in the bam file and discards reads that
a) are not paired
b) not properly mapped
c) mate is not mapped
d) alignment is not primary
e) reads fail quality control of vendor
f) is marked as PCR duplicates.

If the filters (a) and (c) are not desired, you can use the parameter -A.
In addition, samtools performs realignment unless the parameter -B is used, and discards low quality reads unless -Q0 is used. Finally, it stops at a certain number of reads unless the -d parameter is invoqued.

I really needed a good quantification of the reads at each position, so I needed to make sure I could trust the pileup (or vcf) files generated by samtools. So I wrote a small script that parses the bam file, reading the flag field, and removes specific reads from the alignment. In this way, I was finally able to produce an alignment that gave me the exact same counts with IGV and samtools pileup.

It would be really great if all these little details were more clear in the documentation, but in the end, the filtering criteria used by samtools was adequate for my needs. Except for the "anomalous read pairs" parameter (-A), which not very appropriate for RNA-seq data.

Hope that helps somebody!

**zyxue** · 01-19-2015, 10:14 AM

Hi clk, where did you find the information? Did you analyze the source code for samtools mpileup? Thanks!

Topics	Statistics	Last Post
Cancer Metastasis: A Deep Dive into Cellular Plasticity by seqadmin Started by seqadmin, 04-11-2024, 12:08 PM	0 responses 25 views 0 likes	Last Post by seqadmin 04-11-2024, 12:08 PM
Proteogenomic Profiles Offer New Clues in Prostate Cancer by seqadmin Started by seqadmin, 04-10-2024, 10:19 PM	0 responses 28 views 0 likes	Last Post by seqadmin 04-10-2024, 10:19 PM
Novel Diagnostic Assay Enhances Ovarian Cancer Detection by seqadmin Started by seqadmin, 04-10-2024, 09:21 AM	0 responses 24 views 0 likes	Last Post by seqadmin 04-10-2024, 09:21 AM
Evolutionary Dynamics of Centromeres: A Comparative Genomic Analysis by seqadmin Started by seqadmin, 04-04-2024, 09:00 AM	0 responses 52 views 0 likes	Last Post by seqadmin 04-04-2024, 09:00 AM

Seqanswers Leaderboard Ad

Announcement

Pileup / extract information from BAM/SAM files

Comment

Comment

Comment

Comment

Latest Articles

ad_right_rmr

News