1 is the lane, 2782 the x coordinate, 993 the y-coordinate, etc.. There's no standard for read names, you could name them anything you want (in fact, they could all be the same if you wanted).
Unconfigured Ad
Collapse
X
-
hmm I guess that depends on the version of illumina reads..I haven't seen five numerical...so I got confuse!!!Originally posted by dpryan View Post1 is the lane, 2782 the x coordinate, 993 the y-coordinate, etc.. There's no standard for read names, you could name them anything you want (in fact, they could all be the same if you wanted).
but it means this is fastq format that's for sure..
Comment
-
http://www.biostars.org/p/911/Originally posted by ron128 View PostDear Sir/Madam, I have no clue as to what platform was used for sequencing this data. The only thing which has been told to me is that this data is from NIH3T3 cell lines :/ the sequencing company has shut up shop and unfortunately there is no way to be sure about this. I already tried using the phred64 option in prinseq, but I was getting the same error which says that the input file is not in fastq format. Thanks for your thoughts on this
may be this will help
Comment
-
Thanks a ton for this! I tried solexaQA as well, but guess whatOriginally posted by paa6 View Posthttp://www.biostars.org/p/911/
may be this will help
another weird error
I have written in to solexaQA mailing list. awaiting a response.
Well I cannot know for certain whether the file is corrupt. My intuition is that it is not corrupt, as i can open the file perfectly fine in a text editor like emacs or vi.
Comment
-
Dear Mr Ryan, Thanks a ton for all your insights. I am trying to run the grep command and see if it is what you are suggesting. I should have some insights by tomorrow upon greping the data, once i get access to my server. thanks a ton! I will post back tomorrow what i get for this data.Originally posted by dpryan View PostActually, this is somewhere between illumina 1.4 and 1.7, since the multiplex tag is included. I should also mention that it's lane 5, not 1 (it's tile 1). I'm used to seeing the 1101+ tile numbers from the hiseq...
Comment
-
The mystery Deepens
@ Mr Ryan: Tried what you suggested. Turns out it IS pre 1.8
Tried out solexa QA as well. Returns Casava 1.3 as the pipeline used for generating the data.
Now here is where things start to get intriguing. I managed to run fastqc on this data, which is now telling me in the report that Casava 1.5 was used for generating the data. It is differing with the solexaQA. might be a minor difference between the 1.3 and 1.5 which both the tools are not able to pick up. What intrigues me is that, I still cannot run this stupid data using bowtie (which is format independant) OR Bwa. ANd the data is not corrupt, because I have managed to run fastqc. I am looking at a few things in here. WIll keep everyone posted. Is there a website someone knows where I can host these reads? Gdrive maybe? or is there something NGS specific? I am thinking this is a great dataset for ppl troubleshooting NGS programs and wanting to learn about NGS in general, to get started working upon
Comment
-
Google drive, dropbox, copy.com, there are a few option out there for sharing bigger files. What sort of error do you get when you run bowtie? It has an option to tell it how the phred scores were encoded (and anyway, one could always change them with a simple script).
Comment
-
For a bit more understanding of the different FASTQ formats, have a look at the Wikipedia page:
The example sequence in that section looks pretty close to your format.
Comment
Latest Articles
Collapse
-
by SEQadmin2
CRISPR/Cas9 sparked the gene editing revolution for both research and therapeutics.1 But this system still showed severe issues that limited its applications. The most prominent were the heavy reliance on PAM sequences, delivery limitations, double-stranded breaks that prompt unintended edits and cell death, and editing inefficiency (both in targeting and in knock-in reliability).
Despite this, “CRISPR helped turn genome editing from a specialized technique into...-
Channel: Articles
07-31-2026, 11:01 AM -
-
by SEQadmin2
Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.
The systematic characterization of the human proteome has...-
Channel: Articles
07-20-2026, 11:48 AM -
-
by SEQadmin2
Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
...-
Channel: Articles
07-09-2026, 11:10 AM -
ad_right_rmr
Collapse
News
Collapse
| Topics | Statistics | Last Post | ||
|---|---|---|---|---|
|
Started by SEQadmin2, 08-03-2026, 10:13 AM
|
0 responses
15 views
0 reactions
|
Last Post
by SEQadmin2
08-03-2026, 10:13 AM
|
||
|
Started by SEQadmin2, 07-31-2026, 02:55 AM
|
0 responses
32 views
0 reactions
|
Last Post
by SEQadmin2
07-31-2026, 02:55 AM
|
||
|
Started by SEQadmin2, 07-24-2026, 12:17 PM
|
0 responses
23 views
0 reactions
|
Last Post
by SEQadmin2
07-24-2026, 12:17 PM
|
||
|
Started by SEQadmin2, 07-23-2026, 11:41 AM
|
0 responses
21 views
0 reactions
|
Last Post
by SEQadmin2
07-23-2026, 11:41 AM
|
Comment