Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • Tsuyoshi
    Member
    • Sep 2012
    • 24

    #1

    How to retrieve an organism's whole proteome from NCBI

    HI.
    I am suffering from a problem from retrieving the whole proteome dataset from NCBI for a while. Now I only have the taxonomic id of the organism (txid684364), and when I use the batch entrez of NCBI (http://www.ncbi.nlm.nih.gov/protein/...anism:noexp%5D), only part of the protein dataset was downloaded to the local computer. However, previously it worked well when I retrieved several other genome proteome.

    Would anyone please give any solution to resolve this problem? I tried using efetch, however, I am confused of the command lines, Could anyone please teach me how to use the efetch by taking this organism(http://www.ncbi.nlm.nih.gov/protein/?txid684364) as examples to retrieve its whole genome proteins data?
    Thanks!
  • GenoMax
    Senior Member
    • Feb 2008
    • 7142

    #2
    After you search with the txid on the protein page http://www.ncbi.nlm.nih.gov/protein

    Go to "Display Settings" drop-down, choose "FASTA" or format you need.

    Then go to "Send to" drop-down on the right and then choose "Destination" as "File". Finally click on "create file".

    I can see 8706 items.
    Attached Files

    Comment

    • Tsuyoshi
      Member
      • Sep 2012
      • 24

      #3
      Originally posted by GenoMax View Post
      After you search with the txid on the protein page http://www.ncbi.nlm.nih.gov/protein

      Go to "Display Settings" drop-down, choose "FASTA" or format you need.

      Then go to "Send to" drop-down on the right and then choose "Destination" as "File". Finally click on "create file".

      I can see 8706 items.
      Thank you GenoMax!
      I did that but after clicking on "create file", an empty sequence.fasta file will be automatically downloaded. And there was an sentence said "Your session has expired. Please repeat your search" inside the file.
      Have you succeeded in getting the right fasta file?

      Comment

      • GenoMax
        Senior Member
        • Feb 2008
        • 7142

        #4
        Originally posted by Tsuyoshi View Post
        Thank you GenoMax!
        I did that but after clicking on "create file", an empty sequence.fasta file will be automatically downloaded. And there was an sentence said "Your session has expired. Please repeat your search" inside the file.
        Have you succeeded in getting the right fasta file?
        The first time around I had not done a complete download but after your post I did. I do get a FASTA file but it had only ~800 sequences in it (nowhere close to 8700 shown on the search page).

        I next tried Genepept format download. That got me a file with 5706 matches for "LOCUS". Still not 8706 items but closer.

        You may want to contact NCBI help desk if the genpept download is not adequate for your needs.

        Comment

        • Tsuyoshi
          Member
          • Sep 2012
          • 24

          #5
          Originally posted by GenoMax View Post
          The first time around I had not done a complete download but after your post I did. I do get a FASTA file but it had only ~800 sequences in it (nowhere close to 8700 shown on the search page).

          I next tried Genepept format download. That got me a file with 5706 matches for "LOCUS". Still not 8706 items but closer.

          You may want to contact NCBI help desk if the genpept download is not adequate for your needs.
          Yes, Thanks GenoMax, I neither got the full 8706 sequences. I am going to try another methods. Thank you again.

          Comment

          • d1antho
            Member
            • Mar 2012
            • 15

            #6
            You could also directly access the ftp site: ftp://ftp.ncbi.nih.gov/genomes/

            From there, you can go to the folder for your organism and look for the the protein information folder/file and retrieve the protein.fa

            You can point and click to this page or you can use command line tools [in unix/linux or mac] such as wget or cURL to retrieve the file.

            Additionally, you could use ensembl (either the ftp site; ftp://ftp.ensembl.org/pub/ or use bioMart to retrieve the information; http://www.ensembl.org/biomart/martview/)

            Comment

            • GenoMax
              Senior Member
              • Feb 2008
              • 7142

              #7
              Originally posted by d1antho View Post
              You could also directly access the ftp site: ftp://ftp.ncbi.nih.gov/genomes/

              From there, you can go to the folder for your organism and look for the the protein information folder/file and retrieve the protein.fa

              You can point and click to this page or you can use command line tools [in unix/linux or mac] such as wget or cURL to retrieve the file.

              Additionally, you could use ensembl (either the ftp site; ftp://ftp.ensembl.org/pub/ or use bioMart to retrieve the information; http://www.ensembl.org/biomart/martview/)
              The organism (Batrachochytrium dendrobatidis JAM81) Tsuyoshi is looking for is not available at NCBI genomes site. It sounds like a chytrid so it may not be on main ensembl site either.

              Comment

              • GenoMax
                Senior Member
                • Feb 2008
                • 7142

                #8
                Originally posted by Tsuyoshi View Post
                Yes, Thanks GenoMax, I neither got the full 8706 sequences. I am going to try another methods. Thank you again.
                Looks like this Genome was sequenced by JGI.

                You can find their protein set here: ftp://ftp.jgi-psf.org/pub/JGI_data/B...teins.fasta.gz

                Parent page for the data for this genome is at: http://genome.jgi-psf.org/Batde5/Bat...nload.ftp.html

                Comment

                • Tsuyoshi
                  Member
                  • Sep 2012
                  • 24

                  #9
                  Originally posted by d1antho View Post
                  You could also directly access the ftp site: ftp://ftp.ncbi.nih.gov/genomes/

                  From there, you can go to the folder for your organism and look for the the protein information folder/file and retrieve the protein.fa

                  You can point and click to this page or you can use command line tools [in unix/linux or mac] such as wget or cURL to retrieve the file.

                  Additionally, you could use ensembl (either the ftp site; ftp://ftp.ensembl.org/pub/ or use bioMart to retrieve the information; http://www.ensembl.org/biomart/martview/)
                  Thank you very much d1antho. I would like to try your method for retrieving other proteomes dataset.

                  Comment

                  • Tsuyoshi
                    Member
                    • Sep 2012
                    • 24

                    #10
                    Originally posted by GenoMax View Post
                    Looks like this Genome was sequenced by JGI.

                    You can find their protein set here: ftp://ftp.jgi-psf.org/pub/JGI_data/B...teins.fasta.gz

                    Parent page for the data for this genome is at: http://genome.jgi-psf.org/Batde5/Bat...nload.ftp.html
                    Thank you so much GenoMax, and yes I downloaded the protein dataset of Batrachochytrium dendrobatidis from JGI. The fasta file contains the sequences, however, the title of each sequence begins with jgi format, which would bring problems for the BLASTP step, since I want to compare the protein datasets between my own proteomics data and Batrachochytrium dendrobatidis proteomes.

                    Anyway, I figured out an alternative method to retrieve the protein dataset from NCBI. By using the url (http://eutils.ncbi.nlm.nih.gov/entre...tmode=text&id=) and adding the GI list (maximum number is around 800 sequences for this method) after that url. Just paste the url into the web browser the corresponding sequences in fasta format will be automatically downloaded. Although it sounds time consuming, I finally got the dataset I wanted.

                    Thank you again for your kind reply.

                    Comment

                    • d1antho
                      Member
                      • Mar 2012
                      • 15

                      #11
                      Hi Tsuyoshi,
                      The broad institute have a genome for batrachochytrium_dendrobatidis: http://www.broadinstitute.org/annota...Downloads.html

                      Project and release information is here:


                      Probably a day late but I hope this helps anyway

                      Comment

                      • padmoo
                        Member
                        • Jun 2015
                        • 16

                        #12
                        Hi everyone,
                        I have a similar problem. I have transcript IDs from JGI but I need ensemble, entrez or GI IDs to run a analysis with KOBAS. I'd rather not search for all 13490 genes manually in the NCBI database and was wondering if someone knows an easy way to get the matching IDs. The organism I'm working with is Thalassiosira pseudonana. There are also KEGG IDs available but they are also in JGI format or EC numbers which KOBAS does not seem to support.

                        Does anyone know a neat way to solve my problem?

                        Thanks!

                        Comment

                        • GenoMax
                          Senior Member
                          • Feb 2008
                          • 7142

                          #13
                          Originally posted by padmoo View Post
                          Hi everyone,
                          I have a similar problem. I have transcript IDs from JGI but I need ensemble, entrez or GI IDs to run a analysis with KOBAS. I'd rather not search for all 13490 genes manually in the NCBI database and was wondering if someone knows an easy way to get the matching IDs. The organism I'm working with is Thalassiosira pseudonana. There are also KEGG IDs available but they are also in JGI format or EC numbers which KOBAS does not seem to support.

                          Does anyone know a neat way to solve my problem?

                          Thanks!
                          If JGI has not made the mappings available then there may be no easy way. NCBI does have a GFF file available (http://www.ncbi.nlm.nih.gov/genome/54) but you probably can't use it as is.

                          Comment

                          • padmoo
                            Member
                            • Jun 2015
                            • 16

                            #14
                            Hi GenoMax,

                            thanks for the link to the NCBI gff! I tried to find this but was unsuccessful.

                            I do have a gff file from JGI, so it shouldn't be a problem to match those with the NCBI file.

                            Comment

                            • GenoMax
                              Senior Member
                              • Feb 2008
                              • 7142

                              #15
                              Originally posted by padmoo View Post
                              Hi GenoMax,

                              thanks for the link to the NCBI gff! I tried to find this but was unsuccessful.

                              I do have a gff file from JGI, so it shouldn't be a problem to match those with the NCBI file.
                              Good. As long as you have a common "key" to anchor the two files you should be able to map the ID's.

                              Comment

                              Latest Articles

                              Collapse

                              • SEQadmin2
                                Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                                by SEQadmin2



                                CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                                Despite this, “CRISPR helped turn genome editing from a specialized technique into
                                ...
                                07-31-2026, 11:01 AM
                              • SEQadmin2
                                Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                                by SEQadmin2


                                Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                                The systematic characterization of the human proteome has
                                ...
                                07-20-2026, 11:48 AM

                              ad_right_rmr

                              Collapse

                              News

                              Collapse

                              Topics Statistics Last Post
                              Started by SEQadmin2, 08-13-2026, 12:22 PM
                              0 responses
                              29 views
                              0 reactions
                              Last Post SEQadmin2  
                              Started by SEQadmin2, 08-11-2026, 10:35 AM
                              0 responses
                              24 views
                              0 reactions
                              Last Post SEQadmin2  
                              Started by SEQadmin2, 08-06-2026, 07:41 AM
                              0 responses
                              38 views
                              0 reactions
                              Last Post SEQadmin2  
                              Started by SEQadmin2, 08-03-2026, 10:13 AM
                              0 responses
                              51 views
                              0 reactions
                              Last Post SEQadmin2  
                              Working...