Seqanswers Leaderboard Ad

Collapse

Announcement

Collapse
No announcement yet.
X
 
  • Filter
  • Time
  • Show
Clear All
new posts

  • Dwgsim reads not generating header properly

    Hello,

    I have been using dwgsim to stimulate the reads for my masters thesis project.
    reads generated by dwgsim are as follows:

    @gi|29366675|ref|NC_000866.4|_148598_1_0_1_0_0_7:0:0_0:0:0_0/1
    TCAATGTTTAAAATTTTTCTATAAGCCTTTAACTTTATAGAATTATTATTCCAGACTAAATTTTCAGTCTGTTCATCATGTTTATCAATGATATTTAAAATCGAATCAAGCAAGATAAACGTCTCAAACGAAATTATGTTCGATTGCAGAAGTTTAAAAATATAACTTGATTGAACTTTTGAATTATACTCAAAAATTTCTTTAAAAGCAGAAACTTCAACTTTTTTACTAAAATAATAAATGTTGCGAATATCTTCTTCAAACTTAAATTTAATTTGCTCTAAGCGTCCGATATATTCA
    +
    274243352322242112304101.421322033226312022224227224543221700163322322254222233202222323410323211235002421417.1212/32032144113/34243333142240265131237222310434142423502121331306222223125411125322651421222221301325123622224320232443510122.240104223//211227532272513317.654312223233312046223331322/3520
    @gi|29366675|ref|NC_000866.4|_76814_1_1_0_0_0_6:0:0_0:0:0_1/1
    TGGACAAGAATTTTTTATTGAAATAAAACCTAAAAAAGAAACACAACCACCAGTTAAACCAGCACATCTAACGACCGCAGCGAAGAAAAGATTTATGAATGAAATTTATACCTGGTCTGTGAACACTGACAAATGGAAAGCAGCACAATCTTTAGCTGAAAAGCGTGGAATAAAATTTAGAAGTCTAACAGAAGATGGATTAGGAGCTCTTGGCTTTAAGGGGGCATAATGGCTATTTTTTAAATAATTAATGAAAGCACTCCCCAATTTCCAAAGGTTAAGCAATCATTAAATGATAAG
    +
    22522022353212222523315421112234/62214021243134223,43/1223252232230232/3.223202222224233232313252311325252022243342030246412026523234522/22132/22/15341151224513230152142411053242222234-22431351303212121442/22222144322524220/251221301223332102013114322254132.252140334236327221301340224342321213223224
    @gi|29366675|ref|NC_000866.4|_107901_1_0_1_0_0_5:0:0_0:0:0_2/1
    TGCTCCTAAATTCTTGGAAGATGTGCTTGCAACTGAAATGGCAGATGAAATCAATAAAGACATTCTGCAGTCTATGATTACAGTGTCAAAACGCTATAAAGTTACAGGAATTACTGATAGTGGATTCATCGATTTGAGTTATGCATCTGCTCCTGAAGCTGGTCGTTCATTATACCGAATGGTATGTGAAATGGTTTCGAATATCCAAAAAGAATCAACTTATACAGCAACGTTCTGTGGTGCTTCAGCTCGTGCCGCTGCGATTCTTGCTGCATCAGGCTGGTTAAAACATAAACCAGA
    +
    000122222226252023235322.522002744/4354861323223512353233012456213104435423/2061742/2303222333236.531/6/22113.2322841222215253302242/04024543251222220232343020061613.32023222362232225221145475/032212202230/44451252162-0423101/3104/64221342214.15423122522032240254130331452422223422/022232323313226243
    @gi|29366675|ref|NC_000866.4|_98258_1_1_0_0_0_2:1:0_0:0:0_3/1
    TGTCCGGAGATAATAAAGTCATTTTTAATCCTCTTTAATATGCTTTAAAATATTTATACCATTGACATACCATGAGATACTGGAACATACTCAGCAGAATGAACCGAATCCACAAATATAACTGGCGCGTAGTCGTCGCTCATATCCTGAAGCTCTTTTGAAAATACTTCAGATGCTAATCGCATGTCATCTTTATCCGCATAATCAATAAATTTTGACTGCGTTGATAACCATCCAAAAATCACTAAAGACATTCCTAAATCGTCATGATAACCTTCTTCAGCCGCCCAAGACACGCCT
    +
    2211223513131223225222123224233522355451323223322312222200121/122332322306232222122122420532154611333313662/202536242223136710122332362322223533020225222221/3/42424432025020224312234203210.41425200660222222261/222523444326015421214235422/430203321424132223//52-2222.3321431123251222/25312263422323023

    As per the documentation of dwgsim given on sourceforge: in reads name header after 2nd underscore there should be start end 2 (zero-based) but in my case its always 1. I dont know what is that happening. Have anybody faced problem like this? Any help is greatly appreciated!!

  • #2
    What type of data are you trying to simulate?

    According to source forge
    "The FASTQs for BWA are split into two files, the first file for one end, the second file for the other. For paired end reads, this means that E1 is in the first file and E2 is in the second file."

    Comment


    • #3
      Originally posted by mastal View Post
      What type of data are you trying to simulate?

      According to source forge
      "The FASTQs for BWA are split into two files, the first file for one end, the second file for the other. For paired end reads, this means that E1 is in the first file and E2 is in the second file."

      Thanks alot for reply Mastel,

      I want to create metagenomic reads. I want to create single end reads. So I was using following command:
      dwgsim -C 10 -1 300 -2 0 test.fasta out

      I think this is happening cause I am keepin -2 as 0. but now when i kept -2 an 300 it did give me start end 2.

      I am new to dwgsim. Can you tell me how can i stimulate single end read and get header with start end 2?

      Comment


      • #4
        I have used other simulation software, but not dwgsim.

        My understanding, after reading the source forge page, is that you will only get End2 values if you generate paired end reads or mate pair reads.

        Comment


        • #5
          Originally posted by mastal View Post
          I have used other simulation software, but not dwgsim.

          My understanding, after reading the source forge page, is that you will only get End2 values if you generate paired end reads or mate pair reads.
          Oh ok .

          Firstly i thought the end 2 values are where the sequence had ended. So when I am keeping length of first read and second read 300. I am getting following output

          @gi|29366675|ref|NC_000866.4|_33409_33681_0_1_0_0_4:0:0_6:0:0_0/2
          TAATATTAAAACCCTGCAGTCGTTGGCAAATGATATTCGCAATAAAAAGCAATCTCTGATCGCAGCAGTAGATAAAGCTAAAAAAGTTCAAGCGGCTATAGAAAAAGCATCTTCTGAGTTTATTGATCATGCTGATGAAATAGCACTGCTTCAAGAAGAACTTGATAAAATTGTTAAGACAAAAACTAATTTAGTAATGGAAAAATATCACCGAGGAATTTTGACTGATGTGCTCAAAGATTCTGGTATTAAAGGTGCTATTATTAAAGAGTACATTCCATTATTTAATAAGCAGATTAA
          +
          4205034434320-222263232022112521140114136152/232253122241252241522252243213222122126043226243422206623122445422422251/026242223222216442203261213213232/22123362225222321460230262462322252132132422302271331122331333644.22311342613242252273223540422321223122220224/232432424421300722-230425/32124026021
          @gi|29366675|ref|NC_000866.4|_119906_120098_0_1_0_0_2:0:0_4:0:0_1/2
          TTAGAAAATCTAGCAGCAAGTTCTTTTTTAACTGCCGGGGAATTATTTAAATCCGGGTCATCCATCCGTTTTTTAAGGTCTTCTTAGGCAGCTTCAACTGATTTAACCGTTGAGTCTTTACTCATATCAGCTGAATCGGCATATTTTTCAAAACGAATCATCGCAGCTCGAGCTTCATTAGCCTTCATTAAAGCATTTTTTCTTTCTTCCGGTGAAAGTTGCTCTAATTTTTCTTCTTCTGCCGCACGCTCTTCGTCGGTAGTCAGCGCTTCTTTATTATCTACACCACGAATCCAGTTA
          +
          6332230203154523124232221227314264042026331352032620321324531342242321315122063221566440250/164228223405221232201353/2222422241225221323221245322262433433110125120503/33215222111251435116204414233/24123124322525233322345224272232103126132511135642252/55333424642/1742222131312222132221247220224222143
          @gi|29366675|ref|NC_000866.4|_35520_35694_0_1_0_0_3:1:0_8:1:0_2/2
          GTCACTGGGATCTGAATGGATTTTATATTTATAAAGGAATGGAATCTCATGGTCTTGAACCCGATTTCCTTAAGACTTATAAAGAAGTGTGGTCTGGTCATTTCCATACTATTTCTGCGGCTGCAAACGTTAGATATATTGGGACACCATGGACACTAACCGCAGGTGACGAGAATGACCCTCGTGGGTTCTGGATGTTTGATACAGAAACAGAACGAACGGAATTTATCCCAAACAATACTACCTGGCATCGTAGAATTCATTATCCATTTAAAGGAAAAACTGACTATAAAGATTTTC
          +
          32433322324425213201033414534/3423212122342522520234343/45213242215522212224712202132043563/0332322603120513211238/11044233222233211234211/34

          So perception was if end 1 is 33409(from 1 sequence) then if i am taking length of 1st is 300. so end value 2 should 33409+300 = 33709 but where it is 33681. I am confused about what is end 1 and end 2 value?

          If you know then please can you explain me what is end 1 and end 2 values?

          Comment


          • #6
            What sequencing platform (e.g. Illumina, Ion Torrent, SOliD) are your simulated reads supposed to be from?

            I think if you read a bit about the sequencing technology of whichever platform you are trying to simulate, you will understand what the reads should be like.

            Comment


            • #7
              Reads are stimulated from Illumina. I will look in to it.

              Thanks alot for all your help!!

              Comment

              Latest Articles

              Collapse

              • seqadmin
                Strategies for Sequencing Challenging Samples
                by seqadmin


                Despite advancements in sequencing platforms and related sample preparation technologies, certain sample types continue to present significant challenges that can compromise sequencing results. Pedro Echave, Senior Manager of the Global Business Segment at Revvity, explained that the success of a sequencing experiment ultimately depends on the amount and integrity of the nucleic acid template (RNA or DNA) obtained from a sample. “The better the quality of the nucleic acid isolated...
                03-22-2024, 06:39 AM
              • seqadmin
                Techniques and Challenges in Conservation Genomics
                by seqadmin



                The field of conservation genomics centers on applying genomics technologies in support of conservation efforts and the preservation of biodiversity. This article features interviews with two researchers who showcase their innovative work and highlight the current state and future of conservation genomics.

                Avian Conservation
                Matthew DeSaix, a recent doctoral graduate from Kristen Ruegg’s lab at The University of Colorado, shared that most of his research...
                03-08-2024, 10:41 AM

              ad_right_rmr

              Collapse

              News

              Collapse

              Topics Statistics Last Post
              Started by seqadmin, 03-27-2024, 06:37 PM
              0 responses
              12 views
              0 likes
              Last Post seqadmin  
              Started by seqadmin, 03-27-2024, 06:07 PM
              0 responses
              11 views
              0 likes
              Last Post seqadmin  
              Started by seqadmin, 03-22-2024, 10:03 AM
              0 responses
              53 views
              0 likes
              Last Post seqadmin  
              Started by seqadmin, 03-21-2024, 07:32 AM
              0 responses
              68 views
              0 likes
              Last Post seqadmin  
              Working...
              X