Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • jms1223
    Junior Member
    • Feb 2009
    • 2

    #1

    MAQ and short read length (DGE)

    We are currently looking into the viability of Digital Gene Expression (DGE) or mRNA-seq as a possible replacement for expression microarrays in our breast cancer studies. DGE generates reads that are only 17 bases in length, and thus allowing for even 1 mismatch is a little questionable when aligning against the human genome. MAQ doesn't seem to allow you to specify the -n flag as anything less than 1 - is this something that can be altered easily? I would love to align my short reads via MAQ but only keep those that align perfectly.

    Along those lines, if a read maps to more than 1 location, MAQ will randomly pick one of those locations for the placement of that read. Is there any way to customize this function so that it checks against a coordinate file or something like that so we can at least have MAQ select a location for that read that is only in the transcriptome to raise our chances of the placement being 'correct'?

    Thank you for your help
  • bioinfosm
    Senior Member
    • Jan 2008
    • 483

    #2
    I have looked at DGE data, and even with 16/17 bp, more than 90% map to the tag sequences (all possible 16mers with the enzyme specificity).
    I am curious to see how MAQ can be modified as well.. quite a few other tools have specific tag algorithm to take care of such aspects..
    --
    bioinfosm

    Comment

    • jms1223
      Junior Member
      • Feb 2009
      • 2

      #3
      You really see >90% mapping to "canonical" regions?
      I've been aligning with MAQ with -n set to 1, and map >99% to the genome. I then extend all reads 4bp off the 5' end and only keep reads that contain CATG (we cut with NlaIII) - we're only keeping 50% of our mapped reads at this step. Then after that we check to see the overlap with genic regions, and it is certainly not as high as you report. What do you do differently?

      Comment

      • kmcarr
        Senior Member
        • May 2008
        • 1181

        #4
        Technically you should not be trying to align your DGE reads to the genome. The tags may not exist as contiguous sequence in the genome; they may span splice sites or polyadenylation sites. To properly interpret DGE data you should first generate a complete set of predicted tags from the genome and transcriptome and then attempt to align your reads to that. To do this you need a well annotated genome. Please see this thread linked below for the software stack created by Ariel Paulson at the Stowers Institute for creating these tag tables and then scripts to interpret the Eland alignments.

        Discussion of next-gen sequencing related bioinformatics: resources, algorithms, open source efforts, etc


        I have used this pipeline for a couple of DGE projects. In one project with Arabidopsis I was able to map 97% of my filtered reads to predicted tags. This was allowing for up to 2 mismatches in the alignment. Counting only perfect matches the hit rate was ~90%. Not all of these were mapped to annotated genes though. Roughly 63% were mapped to genes, the remainder were to intergenic or repetitive regions.
        Last edited by kmcarr; 02-23-2009, 03:29 PM. Reason: correct spelling error

        Comment

        • Torst
          Senior Member
          • Apr 2008
          • 275

          #5
          Originally posted by jms1223 View Post
          Along those lines, if a read maps to more than 1 location, MAQ will randomly pick one of those locations for the placement of that read. Is there any way to customize this function so that it checks against a coordinate file or something like that so we can at least have MAQ select a location for that read that is only in the transcriptome to raise our chances of the placement being 'correct'?
          If you only want to have reads mapped to your transcriptome, perhaps just make your reference sequences the transcripts themselves, rather than the genome sequence?

          --Torst

          Comment

          • bioinfosm
            Senior Member
            • Jan 2008
            • 483

            #6
            kmarr and Torst answered that for me jms1223
            --
            bioinfosm

            Comment

            Latest Articles

            Collapse

            • SEQadmin2
              Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
              by SEQadmin2



              CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

              Despite this, “CRISPR helped turn genome editing from a specialized technique into
              ...
              07-31-2026, 11:01 AM
            • SEQadmin2
              Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
              by SEQadmin2


              Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

              The systematic characterization of the human proteome has
              ...
              07-20-2026, 11:48 AM
            • SEQadmin2
              Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
              by SEQadmin2



              Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
              ...
              07-09-2026, 11:10 AM

            ad_right_rmr

            Collapse

            News

            Collapse

            Topics Statistics Last Post
            Started by SEQadmin2, Yesterday, 07:41 AM
            0 responses
            12 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 08-03-2026, 10:13 AM
            0 responses
            27 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-31-2026, 02:55 AM
            0 responses
            39 views
            0 reactions
            Last Post SEQadmin2  
            Started by SEQadmin2, 07-24-2026, 12:17 PM
            0 responses
            25 views
            0 reactions
            Last Post SEQadmin2  
            Working...