Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • pmiguel
    Senior Member
    • Aug 2008
    • 2328

    #1

    Why is the growth rate of the SRA decreasing?

    Growth of the short read archive at EMBL appears to be plateauing:



    That is the doubling time is trending upwards:



    This is well below the doubling time for raw megabases/$ -- which is around 6 months.

    Is some other archive for raw data being used? Or is the raw data simply not being submitted to archives any longer?

    --
    Phillip
  • NicoBxl
    not just another member
    • Aug 2010
    • 264

    #2
    the scale is logarithmic

    Comment

    • GenoMax
      Senior Member
      • Feb 2008
      • 7142

      #3
      People are hitting ISP data caps trying to upload data

      Comment

      • mgogol
        Senior Member
        • Mar 2008
        • 197

        #4
        I think people are slowing down a little on generating data after they realized how much it takes to analyze it.

        Or maybe just not sharing.

        Comment

        • pmiguel
          Senior Member
          • Aug 2008
          • 2328

          #5
          Originally posted by NicoBxl View Post
          the scale is logarithmic
          Yeah, I know. But I would expect the doubling time to be similar to the doubling time for megabases/$--currently about 6 months. Instead is appears to be at 14 months and is trending upwards.

          --
          Phillip

          Comment

          • maubp
            Peter (Biopython etc)
            • Jul 2009
            • 1544

            #6
            Originally posted by GenoMax View Post
            People are hitting ISP data caps trying to upload data
            Or their Institute's bandwidth isn't up to it

            Some of our local sequencing providers can submit direct to the ENA/SRR on your behalf - the only tricky bit is providing the metadata.

            Comment

            • kmcarr
              Senior Member
              • May 2008
              • 1181

              #7
              Perhaps people are realizing there isn't sufficient value in archiving every scrap of raw sequence data produced to justify the cost. I think there is an argument to made that as the cost for each Gbp of sequence decreases so does the value. Not long ago when producing even a Mbp of sequence meant a substantial investment in both dollars and person hours you made sure that every bp of DNA you sequenced was meaningful and to protect that investment by having your data safely stored for posterity. Now one can produce hundreds of Gbp for orders of magnitude less effort and money so researchers are a somewhat less choosey about what and how much they sequence.

              Let's be honest, how much raw sequence is ever downloaded from the ENA or SRA for research purposes. I agree with the NCBI's current stance on submission of raw sequence to the SRA. They will accept submissions of raw sequence that are directly reported on in a publication or that correlate to an analyzed data set in some other repository at NCBI (e.g. GEO, Genome, etc.)

              Comment

              • james hadfield
                Moderator
                Cambridge, UK
                Community Forum
                • Feb 2008
                • 224

                #8
                We have been sequencing like this for three to four years (see the jump in 2008) and thats about as long as most PhDs and post-docs work on a project before moving on. Maybe everyone is enjoying a long summer after a crazy time in the lab and before writing all this data up!

                Comment

                • srasdk
                  Member
                  • Jun 2011
                  • 19

                  #9
                  A growing portion of sequencing capacity is occupied by human disease studies(cancer, diabetes,etc..) and private medical/pharma sequencing. The former is not exchanged between archives due to differences in privacy laws, the latter stays private.

                  Comment

                  • Richard Finney
                    Senior Member
                    • Feb 2009
                    • 701

                    #10
                    "differences in privacy laws".

                    How come there's not much public cancer data sets?

                    Are there laws preventing people from making their genome public? I imagine the motivation to help others suffering from a disease that is killing them might be pretty strong.

                    If ethics is the problem, perhaps the ethics needs to be over-hauled. A thousand eyeballs looking at some of the problems in cancer might bring a lot of solutions, particularly if patients are willing to let their genomes out.

                    Comment

                    • srasdk
                      Member
                      • Jun 2011
                      • 19

                      #11
                      A researcher can get access to cancer data through an application process, where he/she is effectively promising not to use it for non-consented research. The data is public, but with concent-based limitations.
                      What I was pointing out is that the data is not exchanged between archives due to different application processes which is due differences in privacy laws. So ENA does not count NCBI cancer data and vise versa. As a result, it is hard to calculate how much data is currently produced and archived.

                      Comment

                      • damiankao
                        Member
                        • Jan 2010
                        • 49

                        #12
                        People just aren't submitting to the SRA because its a pain in the ass honestly. Sequences are being generated in such huge volumes and speed, I think it's hard for users to keep up with submissions.

                        Comment

                        • samanta
                          Senior Member
                          • Feb 2010
                          • 108

                          #13
                          We saw the same trend (slowing down of SRA growth) -


                          There are three possibilities -

                          i) The exponential spike was due to US stimulus spending. Now we are seeing Tea Party decline.

                          ii) SRA scared people about shutting down early last year, and that may have forced some to change submission style.

                          iii) Everyone has too much data and they are down to analysis.
                          http://homolog.us

                          Comment

                          Latest Articles

                          Collapse

                          • SEQadmin2
                            Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                            by SEQadmin2



                            CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                            Despite this, “CRISPR helped turn genome editing from a specialized technique into
                            ...
                            07-31-2026, 11:01 AM
                          • SEQadmin2
                            Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                            by SEQadmin2


                            Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                            The systematic characterization of the human proteome has
                            ...
                            07-20-2026, 11:48 AM

                          ad_right_rmr

                          Collapse

                          News

                          Collapse

                          Topics Statistics Last Post
                          Started by SEQadmin2, 08-06-2026, 07:41 AM
                          0 responses
                          23 views
                          0 reactions
                          Last Post SEQadmin2  
                          Started by SEQadmin2, 08-03-2026, 10:13 AM
                          0 responses
                          37 views
                          0 reactions
                          Last Post SEQadmin2  
                          Started by SEQadmin2, 07-31-2026, 02:55 AM
                          0 responses
                          43 views
                          0 reactions
                          Last Post SEQadmin2  
                          Started by SEQadmin2, 07-24-2026, 12:17 PM
                          0 responses
                          26 views
                          0 reactions
                          Last Post SEQadmin2  
                          Working...