Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • biterbilen
    Junior Member
    • Jun 2009
    • 6

    bwa MD and cigar fields inconsistency

    Hello,

    I have sequencing data where the position and frequency of mismatches play important role in the downstream analysis. I generated short read mappings using BWA. In the samse output file I have inconsistent MD and cigar fields. As far as I saw MD field is generated from CIGAR field and should be consistent with it. Did anyone have the same problem?

    >less uniq_part_001.fastq.sam | cut -f 1-6,10,19 |grep -v "*"| head
    seqAAA_0 16 chr12 50113894 25 36M CGCCATCTGTTTTTTTTTTTTTTTTTTTTTTTTTTT MD:Z:1A34
    seqAAA_1 0 chr4 178545809 0 36M AAAAAAAAAAAAAAAAAAAAAAAAAAACATATGCCT MD:Z:34T1
    seqAAA_3 16 chr10 6381930 25 36M CAGAAGACGTTTTTTTTTTTTTTTTTTTTTTTTTTT MD:Z:8T27
    seqAAA_4 0 chr4 39577867 25 36M AAAAAAAAAAAAAAAAAAAAAAAAAAACGTTTGCCC MD:Z:27A8
    seqAAA_7 0 chr9 117092453 25 36M AAAAAAAAAAAAAAAAAAAAAAAAAAACTTATGCCC MD:Z:26C9
    seqAAA_8 16 chr20 17734163 25 36M CGGAATAAGTTTTTTTTTTTTTTTTTTTTTTTTTTT MD:Z:0A35
    seqAAA_10 0 chr8 112121071 0 36M AAAAAAAAAAAAAAAAAAAAAAAAAAACTTTAGCCG MD:Z:28A7
    seqAAA_11 16 chr2 66418968 0 36M GCATACCTCTTTTTTTTTTTTTTTTTTTTTTTTTTT MD:Z:36
    seqAAA_12 0 chr16 73425684 0 36M AAAAAAAAAAAAAAAAAAAAAAAAAAAGCTGGGACC MD:Z:34G1
    seqAAA_13 16 chr3 22762342 0 36M GGGGCATCCTTTTTTTTTTTTTTTTTTTTTTTTTTT MD:Z:1T34

    Biter Bilen
  • SillyPoint
    Member
    • May 2008
    • 39

    #2
    According to the latest (I think) SAM format spec, at http://samtools.sourceforge.net/SAM1.pdf, an M in the Cigar field is either "Match or mismatch" (sec. 2.2.3) -- as differentiated from an insertion or deletion. Only in the MD field do you get told whether it was a match or mismatch at a given position (see footnote 3 to the table in sec 2.2.4).

    So for your first read, the reference has an A in the 2nd position, where the read has a G. All other bases match.

    Not sure why Cigar works that way -- historical reasons, probably. lh3 may know more.

    SillyPoint

    Comment

    • nilshomer
      Nils Homer
      • Nov 2008
      • 1283

      #3
      Originally posted by SillyPoint View Post
      According to the latest (I think) SAM format spec, at http://samtools.sourceforge.net/SAM1.pdf, an M in the Cigar field is either "Match or mismatch" (sec. 2.2.3) -- as differentiated from an insertion or deletion. Only in the MD field do you get told whether it was a match or mismatch at a given position (see footnote 3 to the table in sec 2.2.4).

      So for your first read, the reference has an A in the 2nd position, where the read has a G. All other bases match.

      Not sure why Cigar works that way -- historical reasons, probably. lh3 may know more.

      SillyPoint
      I believe protein encoded cigar strings (the originator of CIGAR or run length encoding in sequence formats) differentiate between match and mismatch, although there is discussion of changing the SAM format to do this as well (among many other things). For the SAM format, you could check out the active discussion the sourceforge SAMtools developer emailing list for more information as this is an active area of discussion. Please commenting on the format here and on the mailing lists so the community and developers can respond to your needs in any upcoming SAM format.

      Comment

      • lh3
        Senior Member
        • Feb 2008
        • 686

        #4
        Someone told me CIGAR was first introduced in Ensembl/exonerate and was designed for nucleotide alignment. The original CIGAR only contain three operations: M/I/D where M stands for alignment match and can be a sequence match or mismatch. SAM's CIGAR is an extension and so keeps M. We are in the middle of adding new operations to differentiate sequence match and mismatch.

        Comment

        • nilshomer
          Nils Homer
          • Nov 2008
          • 1283

          #5
          Originally posted by lh3 View Post
          Someone told me CIGAR was first introduced in Ensembl/exonerate and was designed for nucleotide alignment. The original CIGAR only contain three operations: M/I/D where M stands for alignment match and can be a sequence match or mismatch. SAM's CIGAR is an extension and so keeps M. We are in the middle of adding new operations to differentiate sequence match and mismatch.
          I guess I should question my sources. I think Guy Slater first introduced CIGAR for Exonerate? Looking at exonerate, you are right it is only M/I/D. Anyone else know more?

          Comment

          Latest Articles

          Collapse

          • SEQadmin2
            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
            by SEQadmin2


            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

            The systematic characterization of the human proteome has
            ...
            07-20-2026, 11:48 AM
          • SEQadmin2
            Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
            by SEQadmin2



            Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
            ...
            07-09-2026, 11:10 AM
          • SEQadmin2
            Cancer Drug Resistance: The Lingering Barrier to Rising Survival
            by SEQadmin2



            Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

            There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
            07-08-2026, 05:17 AM

          ad_right_rmr

          Collapse

          News

          Collapse

          Topics Statistics Last Post
          Started by SEQadmin2, Today, 12:17 PM
          0 responses
          10 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, Yesterday, 11:41 AM
          0 responses
          11 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-20-2026, 11:10 AM
          0 responses
          23 views
          0 reactions
          Last Post SEQadmin2  
          Started by SEQadmin2, 07-13-2026, 10:26 AM
          0 responses
          37 views
          0 reactions
          Last Post SEQadmin2  
          Working...