Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • attilav
    Junior Member
    • Mar 2011
    • 7

    weird BWA SAM (samse) output

    Hi all!

    I've done an alignment job using BWA, and as a result, I've got myself an SAM file.
    I read through the SAM format specification on samtools.sourceforge.net, I also checked out the example sam file, that comes with the example library of Samtools.
    However, my SAM file seems to be different from what it should be ( or i'm just stupid, which is highly likely).
    First there is the header part, that seems to be alright. The alignment part is, where is gets messy, it looks like this:

    NG-5232_4_1_1033_2620#0 4 * 0 0 * * 0 0 CGTTACGGTGTCGGTCTCGTAGAGATATGAACCCTCGTCCCCATGGATTCATGCCAGTTCGTTTATCGCTCGGCATACCTCGCATTCCGTCCTCTGTATTANNNNNNN ).,33<B>A<AAAAAAAAA@@=84@###################################################################################

    So basically, In the first line, I get the FastQ shortread identifier, after that an 8 character long code, that, as far as I can tell, only includes 4-s, 0-s and *-s in every case. Then the second line consists of the sequence, that was supposed to be aligned, and then in the third line the Phred scores.
    And all I got is 3 such lines for every sequence.

    Can you guys tell me, how can I interpret this result, and what may be the cause of me not getting the standard 11 mandatory fields per alignment output, that the format specification mentions?

    Thanks, Attila
    Last edited by attilav; 12-21-2011, 11:59 AM.
  • aggp11
    Member
    • Jun 2011
    • 87

    #2
    Attilav,

    if I am reading this right, then everything except the quality of the read is fine here.

    A sam file is basically tab delimited and it seems like you have a tab-delimited file.

    NG-5232_4_1_1033_2620#0 - I think would be the read name (first column), confirm it in the Fastq file. Everything afterwards represents a different column according to the sam header.

    BWA lists only one record for each read (sequence) and that's why "All you got were 3 such lines for every sequence".

    I think your alignment worked just fine.

    The only thing I would be worried about is the quality of the bases in this particular read. A # represents a q-score of 2 which is really low and in the case of this read almost 75% of the read has q-score "2" bases.

    I hope this helps.

    Praful

    Comment

    • swbarnes2
      Senior Member
      • May 2008
      • 910

      #3
      I don't think there's anythign wrong. You did single end alignment. That second column can only have 3 different outputs in single end data: 0, 4, and 16. The line you posted has a 4, meaning it didn't align anywhere, and with no paired end mate, there's nothing more a .sam file can say about an unmapped read.

      Some of those other columns have information about where the mate aligned, but since you have no mates, of course those will be empty.

      Comment

      • Richard Finney
        Senior Member
        • Feb 2009
        • 701

        #4
        I think your editor is breaking long lines for you. The output (if on one line) looks good. Do a "head -1" from the command line.

        Comment

        Latest Articles

        Collapse

        • SEQadmin2
          Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
          by SEQadmin2


          Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

          The systematic characterization of the human proteome has
          ...
          07-20-2026, 11:48 AM
        • SEQadmin2
          Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
          by SEQadmin2



          Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
          ...
          07-09-2026, 11:10 AM
        • SEQadmin2
          Cancer Drug Resistance: The Lingering Barrier to Rising Survival
          by SEQadmin2



          Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

          There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
          07-08-2026, 05:17 AM

        ad_right_rmr

        Collapse

        News

        Collapse

        Topics Statistics Last Post
        Started by SEQadmin2, 07-24-2026, 12:17 PM
        0 responses
        19 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-23-2026, 11:41 AM
        0 responses
        19 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-20-2026, 11:10 AM
        0 responses
        25 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 07-13-2026, 10:26 AM
        0 responses
        38 views
        0 reactions
        Last Post SEQadmin2  
        Working...