Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • bioinfosm
    Senior Member
    • Jan 2008
    • 483

    short reads missed by aligners

    Anyone looking into the No Match eland reads, or reads that come off solexa that are not mapped to the reference?
    Any other kind of contamination control like eColi, etc?

    I was looking into blat on the entire nt, but would love to hear what people are using.

    sm
    --
    bioinfosm
  • swbarnes2
    Senior Member
    • May 2008
    • 910

    #2
    Try using something like velvet to align all the unaligned reads to each other, then BLAST those contigs against nr. If they are crummy reads, they won't align to each other.

    We tested an in-house clone collection, and I found a fair bit of e.coli contamination. And I've found vector-looking things in microbial samples...stuff like that. If your reference has a biggish deletion compared to what you really sequenced, you might find it this way.

    Comment

    • acnoll
      Member
      • Mar 2008
      • 14

      #3
      Originally posted by bioinfosm View Post
      Anyone looking into the No Match eland reads, or reads that come off solexa that are not mapped to the reference?
      Any other kind of contamination control like eColi, etc?

      I was looking into blat on the entire nt, but would love to hear what people are using.

      sm
      One approach that takes a while but exhaustively looks at all the NMs is to do a blat on the genome of interest to kick out gapped hits and take what is left and then blast to nr to find contaminants. I was thinking to then take the top couple contaminants and look at the matching hits to see if there is any overlap since maybe reads from the contaminant intersect with those mapped to the genome of interest. This might be most important for SNP calling.

      Comment

      • Mr. Gunn
        Member
        • Dec 2007
        • 10

        #4
        Here's a nice comparison of the various short-read aligners, including eland.

        This month I’ve come across some interesting statistics on the performance of Maq, Eland, and other short-read alignment tools as applied to Illumina/Solexa data. I took note because these pr…

        Comment

        • bioinfosm
          Senior Member
          • Jan 2008
          • 483

          #5
          thanks for your inputs...

          Edena and velvet - 2 de novo assemblers using short read data gave so different outputs!

          Velvet gave 2 contigs that pointed to a fragment that was supposedly deleted out and should not have been sequenced

          edena on the other hand gave 10 or so contigs 100-120 bp long, that align perfectly to the eColi K-12!
          --
          bioinfosm

          Comment

          • zee
            NGS specialist
            • Apr 2008
            • 249

            #6
            Reads that aren't matched by Eland are interesting because we would suppose that they're not repeats because Eland reports the matches with multiple locations.
            I would say that gaps in a read would probably be missed by Eland, so use a short read aligner that can find gaps on these reads. I've been using novoalign (www.novocraft.com) and it can find up to 7/8 gaps in a 36bp read matching to a reference sequence, and fast on large ones. I've even tested it on simulated data with mutation rates in excess of 15% and it still finds them. Use a very high threshold e.g. -t 200 to find potentially all permutations for your read.
            I'd be interested to know how much more you may be able to match out of your Eland NM reads.

            Comment

            • kmay
              Member
              • Aug 2008
              • 29

              #7
              Just a note from my side:

              As you know from other threads, we can map from 10bp onwards, with gaps and PMs. However, before tweaking the unmapped reads into the reference genome, look at viral genomes, vectors etc.
              We found numerous perfect matches there. Especially when working on specific cell lines, check the history of that line, how it was immortalized etc. You´ll be surprised how many good old retroviral friends you find!

              Cheers

              Klaus

              Comment

              • Chipper
                Senior Member
                • Mar 2008
                • 323

                #8
                Interresting note, have you looked also at if you can remap the retroviral sequences with mismatches to human and if it seems to be a source of background in alignments?

                Comment

                • kmay
                  Member
                  • Aug 2008
                  • 29

                  #9
                  Chipper,

                  more on that with HEK cells and SV40 and Adenovirus is described in our paper

                  Klaus

                  Comment

                  • zee
                    NGS specialist
                    • Apr 2008
                    • 249

                    #10
                    I just read the Sultan paper Kmay, nice work

                    However, I am a little confused because it says that reads were mapped with ELAND, " Illumina deep sequencing was used to generate 27-bp reads from replicate samples for each cell line. Reads were mapped to the human genome (hg18, NCBI build 36.1) using the Eland software, allowing up to two mismatches (see SOM). Of the total reads, 50% matched to unique genomic locations," (http://www.sciencemag.org/cgi/content/full/1160342/DC1)

                    And the actual read data is unavailable . So I'm assuming that you'll used the proprietary genomatix mapper in a separate study?? Where can we get this read data?

                    Comment

                    • kmay
                      Member
                      • Aug 2008
                      • 29

                      #11
                      zee,

                      you are right. The original data were mapped with ELAND. At those days our GMS was under development. Later we looked at the ELAND non mapped reads and ran those over the viral genomes with our GMS. The actual data reads are deposited at the GEO.

                      Comment

                      Latest Articles

                      Collapse

                      • SEQadmin2
                        Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                        by SEQadmin2


                        Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                        The systematic characterization of the human proteome has
                        ...
                        07-20-2026, 11:48 AM
                      • SEQadmin2
                        Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                        by SEQadmin2



                        Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                        ...
                        07-09-2026, 11:10 AM
                      • SEQadmin2
                        Cancer Drug Resistance: The Lingering Barrier to Rising Survival
                        by SEQadmin2



                        Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

                        There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
                        07-08-2026, 05:17 AM

                      ad_right_rmr

                      Collapse

                      News

                      Collapse

                      Topics Statistics Last Post
                      Started by SEQadmin2, Yesterday, 12:17 PM
                      0 responses
                      11 views
                      0 reactions
                      Last Post SEQadmin2  
                      Started by SEQadmin2, 07-23-2026, 11:41 AM
                      0 responses
                      11 views
                      0 reactions
                      Last Post SEQadmin2  
                      Started by SEQadmin2, 07-20-2026, 11:10 AM
                      0 responses
                      23 views
                      0 reactions
                      Last Post SEQadmin2  
                      Started by SEQadmin2, 07-13-2026, 10:26 AM
                      0 responses
                      37 views
                      0 reactions
                      Last Post SEQadmin2  
                      Working...