Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • gsgs
    Senior Member
    • Oct 2009
    • 139

    #1

    where to get aligned datasets ?

    I usually download the data from genbank

    but it's tedious to align the sequences, filter out the possible errors
    or incorrect insertions or just distant not well matching strains.

    Others must have done the same thing ...

    It should be useful to provide the aligned data to others,
    so they needn't redo it.
    But I didn't find it. Genbank doesn't seem interested
    to provide it or to store it and make it available from other's uploads
  • GenoMax
    Senior Member
    • Feb 2008
    • 7142

    #2
    What datasets are you referring to?

    If you are looking for gene level pre-compiled alignments then "Homologene" is the place you want to visit. Here is an example: http://www.ncbi.nlm.nih.gov/homologene/?term=brca2

    UCSC provides alignments. Look in the alignments section: http://hgdownload.soe.ucsc.edu/downloads.html#human

    Ensembl also has similar information available: http://www.ensembl.org/info/website/...s/compara.html

    Genome level alignments are also at Ensembl: http://www.ensembl.org/info/genome/c.../analyses.html

    Comment

    • gsgs
      Senior Member
      • Oct 2009
      • 139

      #3
      I'm mainly doing influenza sequencing.

      So, I need aligned datasets of ~10000 sequences of length 838-2280 nucleotides
      for avian influenza of the 8 segments and 15 different strains for the HA and 9 for the
      NA and each of these probably divided into an Eurasian and North American lineage.

      Earlier here I had mitochondrial human DNA, 15000 sequences of length 16680
      I also (occasionally) did Dengue, the 4 groups, Ebola etc.
      Today I was trying helicobacter pylori ...

      it's always the same problem, takes hours to generate suitable aligned datasets

      Comment

      • GenoMax
        Senior Member
        • Feb 2008
        • 7142

        #4
        A search brought this up. You must have seen this already: http://www.ncbi.nlm.nih.gov/genomes/FLU/FLU.html

        Then there is http://www.fludb.org/brc/home.spg?decorator=influenza

        For Mitochondria: http://www.ncbi.nlm.nih.gov/genome/organelle/

        As you know first hand, it takes time/effort to create meaningful MSA's. I am going to speculate that NCBI creates those for genes of model organisms/common genomes using the limited resources they have.

        You should consider making your own alignments available since that would save someone else some frustration.

        Comment

        • gsgs
          Senior Member
          • Oct 2009
          • 139

          #5
          For influenza, I think the best is to download all the ~400000 unaligned genbank sequences
          in fasta-format, which they provide in one file of ~650MB.
          But then you must filter for segments, groups, align, sort etc.
          I'm doing this regularly ~1-2 times per year for the ~130000 avian sequences
          into 5+2+9+16 aligned files. Takes 10-20hours.
          If only one person in the world would be doing the same ...it would save much time.

          Ideally you would have ~100 files with aligned sequences for the strains with an index from each.
          And the files sorted by best neighbor match. From these you can extract and filter whatever you want.
          flugenome.org did something like this, but is no longer being updated.


          flu comes from birds , whenever
          it jumps to new hosts you want to know where it came from,
          the genome and each of the 8 segments separately, how it evolved,
          whether/where there is pandemic danger.

          And then the human and swine sequences for special types less regularly,
          when the flu-season starts and there are new variants or such.

          I assume it's similar for other organisms : the data should be provided
          in filtered,sorted,aligned form.

          I could easily make my files available from my HD, where to put them so other will find it ?
          Best to send them on micro-SD



          what's MSA
          Last edited by gsgs; 11-26-2015, 05:55 AM.

          Comment

          • GenoMax
            Senior Member
            • Feb 2008
            • 7142

            #6
            MSA = Multiple sequence alignment

            Isn't NCBI allowing you to do something similar to flugenome here (it is limited to 1000 genomes): http://www.ncbi.nlm.nih.gov/genomes/...i?go=alignment

            That said, I agree with you that the analysis you are doing would be a useful resource for the flu community. But since the number of people working on flu must be relatively small can't you propose this internally (at a relevant meeting/working group) that a resource such as this be created and then hosted by the group.

            Or you could write to NCBI and the group that manages the flu database and see if they would be interested in presenting the data the way you are proposing.

            Comment

            • gsgs
              Senior Member
              • Oct 2009
              • 139

              #7
              it's not just the MSA, you must remove/separate errors and nonmatches and
              single-nucleotide insertions (==> probably error) , pseudo-recombinations , wrong segments,
              wrong or missing strain-classifications, and such.
              And then sort the sequences. And these are typically 10000 sequences.
              It can be done, but takes some time (or tedious automization...)

              I've been talking with the genbank flu expert in emails since 2006.
              They are not interested. Genbank-flu has improved since
              2006, though. More features, more uniform=computer friendly,

              I could upload it somewhere, but noone will find it.

              the flu-community may be small (and I'm not a member with meetings or writing papers
              or professional=being paid or such) but this problem in general should apply to all sequencing.
              It's just my amateur pandemic concern, that started with H5N1 in 2005

              They may have somehow "solved" it in the human community (?)

              Comment

              Latest Articles

              Collapse

              • SEQadmin2
                Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                by SEQadmin2



                CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                Despite this, “CRISPR helped turn genome editing from a specialized technique into
                ...
                Yesterday, 11:01 AM
              • SEQadmin2
                Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                by SEQadmin2


                Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                The systematic characterization of the human proteome has
                ...
                07-20-2026, 11:48 AM
              • SEQadmin2
                Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                by SEQadmin2



                Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                ...
                07-09-2026, 11:10 AM

              ad_right_rmr

              Collapse

              News

              Collapse

              Topics Statistics Last Post
              Started by SEQadmin2, Yesterday, 02:55 AM
              0 responses
              9 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-24-2026, 12:17 PM
              0 responses
              12 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-23-2026, 11:41 AM
              0 responses
              12 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-20-2026, 11:10 AM
              0 responses
              24 views
              0 reactions
              Last Post SEQadmin2  
              Working...