UCSC Encode files issue with MAPQ and Flags

Jandropu

Junior Member

Join Date: Apr 2013
Posts: 3

UCSC Encode files issue with MAPQ and Flags

11-27-2017, 02:24 PM

Hi,

I have been reading Seqanswers for quite some time, but this my first question.

Does someone has experience with BAM files from UCSC-encode project? I am interested in a small RNASeq dataset from CSHL and I found some issues in the BAM files hosted at UCSC.

It seems there are multiple alignments (2) for the same read, but not flagged as primary and secondary. In the first example bellow, same read name appears twice with same flag (0) and same alignment score (254). In the second example the duplicated read name is the reverse complement sequence.

Code:

# Number of multiple alignments in a sample
samtools view ftp://hgdownload.cse.ucsc.edu/goldenPath/hg19/encodeDCC/wgEncodeCshlShortRnaSeq//wgEncodeCshlShortRnaSeqA549CellShorttotalTapAlnRep1.bam | head -n 10000| awk '{print $1}' | sort | uniq -c | awk '{print $1}' | sort | uniq -c
   9958 1
     21 2

# Example 1
view ftp://hgdownload.cse.ucsc.edu/goldenPath/hg19/encodeDCC/wgEncodeCshlShortRnaSeq/wgEncodeCshlShortRnaSeqA549CellShorttotalTapAlnRep1.bam | grep -m 2 'TUPAC_0038:7:110:15042:19226#0/1'
TUPAC_0038:7:110:15042:19226#0/1        0       chr1    45243516        254     34M2S   *       0       0       CTCGTGATGAAAACTTTGTCCAGTTCTGCTACTGAA Ycacc\KW_RSLWSVMTTYT]a\a_`_KXK\Z\RQ_
TUPAC_0038:7:110:15042:19226#0/1        0       chr1    45244067        254     3S31M2S *       0       0       CTCGTGATGAAAACTTTGTCCAGTTCTGCTACTGAA Ycacc\KW_RSLWSVMTTYT]a\a_`_KXK\Z\RQ_

# Example 2
samtools view ftp://hgdownload.cse.ucsc.edu/goldenPath/hg19/encodeDCC/wgEncodeCshlShortRnaSeq/wgEncodeCshlShortRnaSeqA549CellShorttotalTapAlnRep1.bam | grep -m 2 'TUPAC_0038:7:88:9054:14403#0/1'
TUPAC_0038:7:88:9054:14403#0/1  16      chr1    234792  246     18S18M  *       0       0       CCGATCTTTTTTTTTTTCTAAGGACATCCTAAAGGA    ghfeecchhhhfhhhhhhghhhhhhhhhhhhhhhhh
TUPAC_0038:7:88:9054:14403#0/1  0       chr1    464897  246     18M18S  *       0       0       TCCTTTAGGATGTCCTTAGAAAAAAAAAAAGATCGG    hhhhhhhhhhhhhhhhhghhhhhhfhhhhcceefhg

Related to it, the alignment scores are also different from what I have normally seen in BAM files. It ranges from 246 to 255:

Code:

samtools view ftp://hgdownload.cse.ucsc.edu/goldenPath/hg19/encodeDCC/wgEncodeCshlShortRnaSeq/wgEncodeCshlShortRnaSeqA549CellShorttotalTapAlnRep1.bam | head -n 10000 | awk '{print $5}' | sort | uniq -c
     13 246
      5 247
    149 248
     11 249
     19 250
     26 251
     93 252
   9587 253
     77 254
     20 255

Is this normal? Am I missing something? It seems to happen only for a small percentage of reads. Would you just ignore those with multiple alignments for further analysis? If so, how would you filter them out?

Thanks!!

Tags: bam, encode, flags, mapq, ucsc

Previous template Next

Essential Discoveries and Tools in Epitranscriptomics

by seqadmin

The field of epigenetics has traditionally concentrated more on DNA and how changes like methylation and phosphorylation of histones impact gene expression and regulation. However, our increased understanding of RNA modifications and their importance in cellular processes has led to a rise in epitranscriptomics research. “Epitranscriptomics brings together the concepts of epigenetics and gene expression,” explained Adrien Leger, PhD, Principal Research Scientist...
- Channel: Articles
04-22-2024, 07:01 AM
Current Approaches to Protein Sequencing

by seqadmin

Proteins are often described as the workhorses of the cell, and identifying their sequences is key to understanding their role in biological processes and disease. Currently, the most common technique used to determine protein sequences is mass spectrometry. While still a valuable tool, mass spectrometry faces several limitations and requires a highly experienced scientist familiar with the equipment to operate it. Additionally, other proteomic methods, like affinity assays, are constrained...
- Channel: Articles
04-04-2024, 04:25 PM

Topics	Statistics	Last Post
Cancer Metastasis: A Deep Dive into Cellular Plasticity by seqadmin Started by seqadmin, 04-11-2024, 12:08 PM	0 responses 59 views 0 likes	Last Post by seqadmin 04-11-2024, 12:08 PM
Proteogenomic Profiles Offer New Clues in Prostate Cancer by seqadmin Started by seqadmin, 04-10-2024, 10:19 PM	0 responses 57 views 0 likes	Last Post by seqadmin 04-10-2024, 10:19 PM
Novel Diagnostic Assay Enhances Ovarian Cancer Detection by seqadmin Started by seqadmin, 04-10-2024, 09:21 AM	0 responses 51 views 0 likes	Last Post by seqadmin 04-10-2024, 09:21 AM
Evolutionary Dynamics of Centromeres: A Comparative Genomic Analysis by seqadmin Started by seqadmin, 04-04-2024, 09:00 AM	0 responses 56 views 0 likes	Last Post by seqadmin 04-04-2024, 09:00 AM

Seqanswers Leaderboard Ad

Announcement

UCSC Encode files issue with MAPQ and Flags

Latest Articles

ad_right_rmr

News