Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • dsturgill
    Junior Member
    • May 2008
    • 5

    #1

    Data retention

    Hi Folks -

    How about a survey to see what data people retain from sequencing experiments? We have completed ~ 10 runs on an Illumina GA I, and have stored complete raw data (images, raw intensities, etc... the complete run folder), on external 1TB disks. I don't think everyone stores images long term. We may not be able to keep it up forever.

    Anyone know if it is possible to retain images using GA II?
  • new300
    Member
    • Mar 2008
    • 50

    #2
    It's possible to retain images on GA2s of course, but they are a lot bigger... In general it's reasonable to retain the raw intensity and noise files as reprocessing those with new basecallers may be of interest.

    If look at the data retained by the NCBI in the short read archive they are currently storing Raw intensity and noise files, processed intensities, the 4 quality scores and the fastq files (the SRFs also have the settings used to generate the data I believe). I think the short read archive only contains PF (purity filtered data) at the moment.

    Personally I think that's overkill. I'd store the raw intensities (PF and non-PF), 4 quality scores and a basecall. Bare in mind that it is possible to regenerate everything from a complete set of raw intensities.

    Comment

    • Chipper
      Senior Member
      • Mar 2008
      • 323

      #3
      Storing the "raw data" as DNA in the freezer is likely going to be a more cost-effective option...

      What would you expect or hope to be able to achive by reanalysing the images?

      Comment

      • dsturgill
        Junior Member
        • May 2008
        • 5

        #4
        Originally posted by Chipper View Post
        What would you expect or hope to be able to achive by reanalysing the images?
        My thought was that future improvements in the image segmentation algorithm or intensity extraction could give you different results later on. For example, improvements might be better able to discern individual clusters when density is very high.

        In our microarray experiments, we always save raw images rather than just intensities, since we see variation whenever you do gridding or intensity extraction.

        You're right - saving sample to re-run later is a logical approach. Deleting what we see as raw data may be simply be a mental hurdle to get over.

        Comment

        • dvh
          Member
          • Jul 2008
          • 35

          #5
          My group has archived all bzip2'd image files to 2Tb USB disks. Currently they are £200 from western digital. Seemed cheap in comparison to reagents/lab time.

          Partly, we also thought better base calling algorithms than Bustard are already in development. e.g altacyclic, others. Can someone perhaps start a thread on this?

          dvh
          Last edited by dvh; 09-16-2008, 02:23 PM.

          Comment

          • ScottC
            Senior Member
            • Jan 2008
            • 244

            #6
            Originally posted by Chipper View Post
            What would you expect or hope to be able to achive by reanalysing the images?
            Well, if I understand correctly, the new pipeline that's on the way is supposed to increase the data output by 15-30% just through image analysis improvements alone. There's a lot of room for improvement in that area, apparently.

            That being said, we're keeping the images on our server only until we have no more room, then they'll be deleted as space is required. Individuals can keep the raw data on external disks if they want it.

            Scott.

            Comment

            • new300
              Member
              • Mar 2008
              • 50

              #7
              Originally posted by dvh View Post
              My group has archived all bzip2'd image files to 2Tb USB disks. Currently they are £200 from western digital. Seemed cheap in comparison to reagents/lab time.

              Partly, we also thought better base calling algorithms than Bustard are already in development. e.g altacyclic, others. Can someone perhaps start a thread on this?

              dvh
              With GA2 and read lengths approaching 100bp paired end this probably becomes infeasible. Storing the raw intensity I think gives you the biggest bang for your storage buck. All the new basecallers work from raw intensities, not images.

              It's true there's a lot to be gained back using improved image analysis, perhaps 10 or 20%, it's all a trade off.

              Comment

              Latest Articles

              Collapse

              • SEQadmin2
                Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                by SEQadmin2



                CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                Despite this, “CRISPR helped turn genome editing from a specialized technique into
                ...
                07-31-2026, 11:01 AM
              • SEQadmin2
                Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                by SEQadmin2


                Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                The systematic characterization of the human proteome has
                ...
                07-20-2026, 11:48 AM
              • SEQadmin2
                Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                by SEQadmin2



                Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                ...
                07-09-2026, 11:10 AM

              ad_right_rmr

              Collapse

              News

              Collapse

              Topics Statistics Last Post
              Started by SEQadmin2, 07-31-2026, 02:55 AM
              0 responses
              20 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-24-2026, 12:17 PM
              0 responses
              16 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-23-2026, 11:41 AM
              0 responses
              16 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-20-2026, 11:10 AM
              0 responses
              26 views
              0 reactions
              Last Post SEQadmin2  
              Working...