SEQanswers

Go Back   SEQanswers > Applications Forums > Genomic Resequencing



Similar Threads
Thread Thread Starter Forum Replies Last Post
Problemastic reads quality score distribution by dwgsim yjx1217 Bioinformatics 1 10-08-2012 01:28 PM
dwgsim to simulate Illumina reads gene coder Bioinformatics 1 07-07-2011 09:51 AM
Help me!!!!! low % of properly paired reads!!! Trudy Bioinformatics 1 05-24-2011 11:26 PM
Simulated reads mapping to same region - maq simulate & dwgsim gprakhar Bioinformatics 2 02-18-2011 11:12 PM
properly paired reads in TopHat output shurjo Bioinformatics 2 12-02-2010 11:35 AM

Reply
 
Thread Tools
Old 04-05-2014, 11:57 AM   #1
prasadg
Member
 
Location: Newark, Delaware, USA

Join Date: Mar 2012
Posts: 16
Default Dwgsim reads not generating header properly

Hello,

I have been using dwgsim to stimulate the reads for my masters thesis project.
reads generated by dwgsim are as follows:

@gi|29366675|ref|NC_000866.4|_148598_1_0_1_0_0_7:0:0_0:0:0_0/1
TCAATGTTTAAAATTTTTCTATAAGCCTTTAACTTTATAGAATTATTATTCCAGACTAAATTTTCAGTCTGTTCATCATGTTTATCAATGATATTTAAAATCGAATCAAGCAAGATAAACGTCTCAAACGAAATTATGTTCGATTGCAGAAGTTTAAAAATATAACTTGATTGAACTTTTGAATTATACTCAAAAATTTCTTTAAAAGCAGAAACTTCAACTTTTTTACTAAAATAATAAATGTTGCGAATATCTTCTTCAAACTTAAATTTAATTTGCTCTAAGCGTCCGATATATTCA
+
274243352322242112304101.421322033226312022224227224543221700163322322254222233202222323410323211235002421417.1212/32032144113/34243333142240265131237222310434142423502121331306222223125411125322651421222221301325123622224320232443510122.240104223//211227532272513317.654312223233312046223331322/3520
@gi|29366675|ref|NC_000866.4|_76814_1_1_0_0_0_6:0:0_0:0:0_1/1
TGGACAAGAATTTTTTATTGAAATAAAACCTAAAAAAGAAACACAACCACCAGTTAAACCAGCACATCTAACGACCGCAGCGAAGAAAAGATTTATGAATGAAATTTATACCTGGTCTGTGAACACTGACAAATGGAAAGCAGCACAATCTTTAGCTGAAAAGCGTGGAATAAAATTTAGAAGTCTAACAGAAGATGGATTAGGAGCTCTTGGCTTTAAGGGGGCATAATGGCTATTTTTTAAATAATTAATGAAAGCACTCCCCAATTTCCAAAGGTTAAGCAATCATTAAATGATAAG
+
22522022353212222523315421112234/62214021243134223,43/1223252232230232/3.223202222224233232313252311325252022243342030246412026523234522/22132/22/15341151224513230152142411053242222234-22431351303212121442/22222144322524220/251221301223332102013114322254132.252140334236327221301340224342321213223224
@gi|29366675|ref|NC_000866.4|_107901_1_0_1_0_0_5:0:0_0:0:0_2/1
TGCTCCTAAATTCTTGGAAGATGTGCTTGCAACTGAAATGGCAGATGAAATCAATAAAGACATTCTGCAGTCTATGATTACAGTGTCAAAACGCTATAAAGTTACAGGAATTACTGATAGTGGATTCATCGATTTGAGTTATGCATCTGCTCCTGAAGCTGGTCGTTCATTATACCGAATGGTATGTGAAATGGTTTCGAATATCCAAAAAGAATCAACTTATACAGCAACGTTCTGTGGTGCTTCAGCTCGTGCCGCTGCGATTCTTGCTGCATCAGGCTGGTTAAAACATAAACCAGA
+
000122222226252023235322.522002744/4354861323223512353233012456213104435423/2061742/2303222333236.531/6/22113.2322841222215253302242/04024543251222220232343020061613.32023222362232225221145475/032212202230/44451252162-0423101/3104/64221342214.15423122522032240254130331452422223422/022232323313226243
@gi|29366675|ref|NC_000866.4|_98258_1_1_0_0_0_2:1:0_0:0:0_3/1
TGTCCGGAGATAATAAAGTCATTTTTAATCCTCTTTAATATGCTTTAAAATATTTATACCATTGACATACCATGAGATACTGGAACATACTCAGCAGAATGAACCGAATCCACAAATATAACTGGCGCGTAGTCGTCGCTCATATCCTGAAGCTCTTTTGAAAATACTTCAGATGCTAATCGCATGTCATCTTTATCCGCATAATCAATAAATTTTGACTGCGTTGATAACCATCCAAAAATCACTAAAGACATTCCTAAATCGTCATGATAACCTTCTTCAGCCGCCCAAGACACGCCT
+
2211223513131223225222123224233522355451323223322312222200121/122332322306232222122122420532154611333313662/202536242223136710122332362322223533020225222221/3/42424432025020224312234203210.41425200660222222261/222523444326015421214235422/430203321424132223//52-2222.3321431123251222/25312263422323023

As per the documentation of dwgsim given on sourceforge: in reads name header after 2nd underscore there should be start end 2 (zero-based) but in my case its always 1. I dont know what is that happening. Have anybody faced problem like this? Any help is greatly appreciated!!
prasadg is offline   Reply With Quote
Old 04-05-2014, 12:51 PM   #2
mastal
Senior Member
 
Location: uk

Join Date: Mar 2009
Posts: 667
Default

What type of data are you trying to simulate?

According to source forge
"The FASTQs for BWA are split into two files, the first file for one end, the second file for the other. For paired end reads, this means that E1 is in the first file and E2 is in the second file."
mastal is offline   Reply With Quote
Old 04-05-2014, 01:03 PM   #3
prasadg
Member
 
Location: Newark, Delaware, USA

Join Date: Mar 2012
Posts: 16
Default

Quote:
Originally Posted by mastal View Post
What type of data are you trying to simulate?

According to source forge
"The FASTQs for BWA are split into two files, the first file for one end, the second file for the other. For paired end reads, this means that E1 is in the first file and E2 is in the second file."

Thanks alot for reply Mastel,

I want to create metagenomic reads. I want to create single end reads. So I was using following command:
dwgsim -C 10 -1 300 -2 0 test.fasta out

I think this is happening cause I am keepin -2 as 0. but now when i kept -2 an 300 it did give me start end 2.

I am new to dwgsim. Can you tell me how can i stimulate single end read and get header with start end 2?
prasadg is offline   Reply With Quote
Old 04-05-2014, 01:13 PM   #4
mastal
Senior Member
 
Location: uk

Join Date: Mar 2009
Posts: 667
Default

I have used other simulation software, but not dwgsim.

My understanding, after reading the source forge page, is that you will only get End2 values if you generate paired end reads or mate pair reads.
mastal is offline   Reply With Quote
Old 04-05-2014, 01:22 PM   #5
prasadg
Member
 
Location: Newark, Delaware, USA

Join Date: Mar 2012
Posts: 16
Default

Quote:
Originally Posted by mastal View Post
I have used other simulation software, but not dwgsim.

My understanding, after reading the source forge page, is that you will only get End2 values if you generate paired end reads or mate pair reads.
Oh ok .

Firstly i thought the end 2 values are where the sequence had ended. So when I am keeping length of first read and second read 300. I am getting following output

@gi|29366675|ref|NC_000866.4|_33409_33681_0_1_0_0_4:0:0_6:0:0_0/2
TAATATTAAAACCCTGCAGTCGTTGGCAAATGATATTCGCAATAAAAAGCAATCTCTGATCGCAGCAGTAGATAAAGCTAAAAAAGTTCAAGCGGCTATAGAAAAAGCATCTTCTGAGTTTATTGATCATGCTGATGAAATAGCACTGCTTCAAGAAGAACTTGATAAAATTGTTAAGACAAAAACTAATTTAGTAATGGAAAAATATCACCGAGGAATTTTGACTGATGTGCTCAAAGATTCTGGTATTAAAGGTGCTATTATTAAAGAGTACATTCCATTATTTAATAAGCAGATTAA
+
4205034434320-222263232022112521140114136152/232253122241252241522252243213222122126043226243422206623122445422422251/026242223222216442203261213213232/22123362225222321460230262462322252132132422302271331122331333644.22311342613242252273223540422321223122220224/232432424421300722-230425/32124026021
@gi|29366675|ref|NC_000866.4|_119906_120098_0_1_0_0_2:0:0_4:0:0_1/2
TTAGAAAATCTAGCAGCAAGTTCTTTTTTAACTGCCGGGGAATTATTTAAATCCGGGTCATCCATCCGTTTTTTAAGGTCTTCTTAGGCAGCTTCAACTGATTTAACCGTTGAGTCTTTACTCATATCAGCTGAATCGGCATATTTTTCAAAACGAATCATCGCAGCTCGAGCTTCATTAGCCTTCATTAAAGCATTTTTTCTTTCTTCCGGTGAAAGTTGCTCTAATTTTTCTTCTTCTGCCGCACGCTCTTCGTCGGTAGTCAGCGCTTCTTTATTATCTACACCACGAATCCAGTTA
+
6332230203154523124232221227314264042026331352032620321324531342242321315122063221566440250/164228223405221232201353/2222422241225221323221245322262433433110125120503/33215222111251435116204414233/24123124322525233322345224272232103126132511135642252/55333424642/1742222131312222132221247220224222143
@gi|29366675|ref|NC_000866.4|_35520_35694_0_1_0_0_3:1:0_8:1:0_2/2
GTCACTGGGATCTGAATGGATTTTATATTTATAAAGGAATGGAATCTCATGGTCTTGAACCCGATTTCCTTAAGACTTATAAAGAAGTGTGGTCTGGTCATTTCCATACTATTTCTGCGGCTGCAAACGTTAGATATATTGGGACACCATGGACACTAACCGCAGGTGACGAGAATGACCCTCGTGGGTTCTGGATGTTTGATACAGAAACAGAACGAACGGAATTTATCCCAAACAATACTACCTGGCATCGTAGAATTCATTATCCATTTAAAGGAAAAACTGACTATAAAGATTTTC
+
32433322324425213201033414534/3423212122342522520234343/45213242215522212224712202132043563/0332322603120513211238/11044233222233211234211/34

So perception was if end 1 is 33409(from 1 sequence) then if i am taking length of 1st is 300. so end value 2 should 33409+300 = 33709 but where it is 33681. I am confused about what is end 1 and end 2 value?

If you know then please can you explain me what is end 1 and end 2 values?
prasadg is offline   Reply With Quote
Old 04-05-2014, 01:31 PM   #6
mastal
Senior Member
 
Location: uk

Join Date: Mar 2009
Posts: 667
Default

What sequencing platform (e.g. Illumina, Ion Torrent, SOliD) are your simulated reads supposed to be from?

I think if you read a bit about the sequencing technology of whichever platform you are trying to simulate, you will understand what the reads should be like.
mastal is offline   Reply With Quote
Old 04-05-2014, 01:33 PM   #7
prasadg
Member
 
Location: Newark, Delaware, USA

Join Date: Mar 2012
Posts: 16
Default

Reads are stimulated from Illumina. I will look in to it.

Thanks alot for all your help!!
prasadg is offline   Reply With Quote
Reply

Tags
dwgsim

Thread Tools

Posting Rules
You may not post new threads
You may not post replies
You may not post attachments
You may not edit your posts

BB code is On
Smilies are On
[IMG] code is On
HTML code is Off




All times are GMT -8. The time now is 12:58 AM.


Powered by vBulletin® Version 3.8.9
Copyright ©2000 - 2019, vBulletin Solutions, Inc.
Single Sign On provided by vBSSO