Unconfigured Ad

Collapse
X
 
  • Time
  • Show
Clear All
new posts
  • lingling huang
    Member
    • Mar 2016
    • 48

    #1

    Can use multiple CPUs or thread to speed up the running rate for tofu_wrap.py?

    Hi, there. I have 7 smrtcell data. It's really slow to run tofu_wrap.py, I want to know to weather I can use multiple CPUs or thread to speed up the running rate. Thank you for any tips!
  • Magdoll
    Member
    • Aug 2011
    • 31

    #2
    Hi,

    Short answer is yes. There are multiple ways to hack tofu_wrap.py -- will require additional monitoring of the parallel jobs. But I first need to ask what parameters did you use to call tofu_wrap.py? What are the size bins (aka what are the subdirectories in clusterOut/)? Do you have an SGE cluster? Do you have multiple nodes from which you can run parallel pbtranscript.py cluster jobs?

    Also, please consider joining the Iso-Seq google group: https://groups.google.com/forum/#!forum/smrt_isoseq

    Since tofu_wrap.py is on the cutting edge (it's not officially supported in SMRTAnalysis 2.x but is on the agenda for SMRTAnalysis 3.x), the google group is better suited

    --Liz

    Comment

    • lingling huang
      Member
      • Mar 2016
      • 48

      #3
      Originally posted by Magdoll View Post
      Hi,

      Short answer is yes. There are multiple ways to hack tofu_wrap.py -- will require additional monitoring of the parallel jobs. But I first need to ask what parameters did you use to call tofu_wrap.py? What are the size bins (aka what are the subdirectories in clusterOut/)? Do you have an SGE cluster? Do you have multiple nodes from which you can run parallel pbtranscript.py cluster jobs?

      Also, please consider joining the Iso-Seq google group: https://groups.google.com/forum/#!forum/smrt_isoseq

      Since tofu_wrap.py is on the cutting edge (it's not officially supported in SMRTAnalysis 2.x but is on the agenda for SMRTAnalysis 3.x), the google group is better suited

      --Liz
      Code:
      tofu_wrap.py --nfl_fa isoseq_nfl.fasta --ccs_fofn reads_of_insert.fofn --bas_fofn input.fofn -d clusterOut --quiver --bin_manual "(0,2,4,6,8,9,10,11,12,13,15,17,19,20,23)" --gmap_db /zs32/data-analysis/liucy_group/llhuang/Reflib/gmapdb --gmap_name gmapdb_h19 --output_seqid_prefix human isoseq_flnc.fasta final.consensus.fa
      The lab's sever is high-powered single-node computer and
      has no an SGE cluster. Can I use multiple CPUs to run command?

      Comment

      • Magdoll
        Member
        • Aug 2011
        • 31

        #4
        Without SGE, it will be slow anyway.

        But if you think that single node can handle it, one way is to run multiple instances of `pbtranscript.py cluster` on the different bins.

        ex: tofu_wrap.py always creates bins 0to1kb_part0, 1to2kb_part0, etc

        You can terminate tofu_wrap and keep the bins as they are. Then separately in each bin call a separate instance of cluster:

        pbtranscript.py cluster isoseq_flnc.fasta final.consensus.fa \
        --nfl_fa isoseq_nfl.fasta -d cluster --ccs_fofn reads_of_insert.fofn \
        --bas_fofn input.fofn --quiver --use_sge \
        --max_sge_jobs 40 --unique_id 300 --blasr_nproc 24 --quiver_nproc 8

        (for you, you would remove the --use_sge and --max_sge_jobs option)

        (see cluster tutorial here: https://github.com/PacificBioscience...-and-Quiver%29)

        My guess is this would make it a bit faster but still relatively slow since everything is running in serial instead of parallel, but may be better than nothing...

        Comment

        Latest Articles

        Collapse

        • SEQadmin2
          New Genomics Technologies Take Aim at Long-Standing Limits
          by SEQadmin2


          Researchers using sequencing and genomics tools often have to make trade-offs. They can choose between speed or scale, short reads or long-range information, or targeted panels or a view of the whole transcriptome. New technologies that have been released this year are built to address those tough choices.

          We asked six companies the same four questions to learn about their latest products. The new technologies bring a lot to the table, including rethinking sequencing
          ...
          Yesterday, 10:25 AM
        • SEQadmin2
          How Immunogenomics Decodes Immunity’s Genetic Blueprint
          by SEQadmin2




          The immune system’s power comes from its genetic diversity, allowing myriad threats to be neutralized through first recognizing foreign antigens. That diversity is also what makes the immune system so difficult to study. Recent advances in sequencing technology and computational biology, however, are giving researchers new tools to understand immune responses and immune-related diseases in greater detail.

          This convergence of genetics, immunology, and computation...
          09-01-2026, 05:41 AM

        ad_right_rmr

        Collapse

        News

        Collapse

        Topics Statistics Last Post
        Started by SEQadmin2, Today, 09:51 AM
        0 responses
        7 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 09-25-2026, 09:06 AM
        0 responses
        31 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 09-23-2026, 11:05 AM
        0 responses
        26 views
        0 reactions
        Last Post SEQadmin2  
        Started by SEQadmin2, 09-18-2026, 11:37 AM
        1 response
        46 views
        0 reactions
        Last Post pekgio
        by pekgio
         
        Working...