Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • yaximik
    Senior Member
    • Apr 2011
    • 199

    #1

    MiSeq PE run output files

    Hi,

    Can anyone advise me why two output files from a paired end run differ in size? The file for run 2 is about 2 times bigger than for read 1, so I thought it inlcudes both reads 1 and reads 2. Yet, the end strings of each read name (either 1:N:0:1 in read 1 or 2:N:0:1 in read 2) indicate this is not the case. So why the difference?
  • mcnelson.phd
    Senior Member
    • Jul 2011
    • 162

    #2
    Are you referring to the gzip compressed or the un-compressed fastq files?

    I haven't seen a case where one read files is twice as large as the other, but I have seen differences in file size for the gzipped files. Part of this is most likely due to the compression algorithm being able to compress one file better than the other. It's also possible that with adapter trimming turned on, depending on the quality of your data, you could have longer reads for read 2 because the data quality dropped enough that Reporter couldn't properly identify the adapter and thus didn't trim it.

    Either way, I wouldn't be concerned about it.

    Comment

    • yaximik
      Senior Member
      • Apr 2011
      • 199

      #3
      No, this is not compression as decompressed files are also about twice longer. The number of records as counted using Biopieces is the same, and clean&trim reduces file size to about the same as read 1. It is very likely much longer records with lots of Ns, which is very surprising. I guess somtheing is wrong with basecalling.

      Comment

      • kcchan
        Senior Member
        • Jul 2012
        • 186

        #4
        Are all of the reads the same length? Did you disable adapter trimming? This may contribute to unequal file sizes.

        Comment

        • yaximik
          Senior Member
          • Apr 2011
          • 199

          #5
          Looks like this indeed a read quality issue as quality filtering/adaptor removal levels file sizes, although read 2 file size now always become smaller. Say, from original 4.2 GB and 8 GB files shrink to 3.8 GB and 3.6 GB, and this is a common trend no matter what library sizes or run lengths are. I am bugging Illumina Tech Support with that. For example, read quality peaks at 100-120 cycles and sharply declines after that even with library size around 600 bp.

          Comment

          • GenoMax
            Senior Member
            • Feb 2008
            • 7142

            #6
            Originally posted by yaximik View Post
            I am bugging Illumina Tech Support with that. For example, read quality peaks at 100-120 cycles and sharply declines after that even with library size around 600 bp.
            Is this a "low complexity"/amplicon type library, if so this is a known issue.

            Are you running MCS v.2.1.1.13 on your MiSeq?

            Comment

            • yaximik
              Senior Member
              • Apr 2011
              • 199

              #7
              No these were all genomic libraries. I just recently upgraded to v.2.1.13.0, the majority of runs were done with whatever was the previous version.
              ILMN support thinks it is a matrix issue, so I was asked to make a full 2x250 bp phiX run to get kinda baseline output and then go from there. Also, they noticed that in the majority of runs the first nucleotide is C, which is consistent with the fact that most of the samples are from ancient archeological samples. From Paabo work on Neanderthal it is known that DNA breaks preferentially either after or before G.

              Comment

              Latest Articles

              Collapse

              • SEQadmin2
                Beyond CRISPR/Cas9: Understand, Choose, and Use the Right Genome Editing Tool
                by SEQadmin2



                CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).

                Despite this, “CRISPR helped turn genome editing from a specialized technique into
                ...
                07-31-2026, 11:01 AM
              • SEQadmin2
                Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                by SEQadmin2


                Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                The systematic characterization of the human proteome has
                ...
                07-20-2026, 11:48 AM

              ad_right_rmr

              Collapse

              News

              Collapse

              Topics Statistics Last Post
              Started by SEQadmin2, Yesterday, 10:35 AM
              0 responses
              7 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 08-06-2026, 07:41 AM
              0 responses
              25 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 08-03-2026, 10:13 AM
              0 responses
              45 views
              0 reactions
              Last Post SEQadmin2  
              Started by SEQadmin2, 07-31-2026, 02:55 AM
              0 responses
              48 views
              0 reactions
              Last Post SEQadmin2  
              Working...