Unconfigured Ad

Collapse
X
 
  • Filter
  • Time
  • Show
Clear All
new posts
  • Shishir
    Member
    • Nov 2012
    • 22

    How to build combined library for repeatmasker?

    I want to to perform repeatmasking on ant genome. Ant is also included in the repeat library obtained from Repbase. The library RepeatMaskerLib.embl also contains my species. However, I also generated de novo repeat library by repeatmodeler (consensi.fa). How can I use both library for masking ant genome ? Should I concatenate RepeatMaskerLib.embl and consensi.fa or can I use both library with single command?
  • suryasaha
    Member
    • Mar 2011
    • 27

    #2
    You can go both ways. Have not done this myself but the simplest way would be to pull put ant repeats from RepeatMaskerLib.embl and merge them with your custom repeat library (as fasta). You could then use the -lib option in RepeatMasker to supply the updated custom repeat library and mask the ant genome.

    AFAIK, you cannot combine both libraries in a single command.

    Comment

    • HeyIamNuria
      Member
      • Dec 2012
      • 19

      #3
      Hey, I think this would be a great solution, but I have a problem...
      I want to mask the repeats, but I want to classify them as well.

      If I get the Repbase library in fasta format (to concatenate it with my own database and use it with -lib) then I lose the classificators (ex: #LINE/L1 --> ).

      [I don't know if anyone came out with a script to put the "#" and "/" where they should be, but I don't know how, because usually there's only the Superfamily but not the Order (ex L1 but not LINE)]

      I also tried to convert RepeatMaskerLib.embl to .fasta but I lose the classification info too.

      I guess it won't work to put my repeats, in .embl, into RepeatMaskerLib.embl, but I haven't tried yet.


      Did anyone find a way to solve this?
      Thank you for any help

      Comment

      • HeyIamNuria
        Member
        • Dec 2012
        • 19

        #4
        or maybe...

        Does anyone know how to use RepeatClassifier (part of RepeatModeler) only. Or make RepeatModeler classify a bunch of repeats already listed?

        Thanks

        Comment

        • rhubley
          Member
          • Sep 2012
          • 10

          #5
          RepeatClassifier can be run independently. If you hand it a fasta file it will generate a new file with a *.classified suffix. The output file will be a fasta file with all the id#class/subclass identifiers.

          Comment

          • rhubley
            Member
            • Sep 2012
            • 10

            #6
            Check out the util/buildRMLibFromEMBL.pl script included in the RepeatMasker distribution. This will convert the RepeatMaskerLib.embl file into a fasta file with the "#class/subclass" nomenclature.

            Comment

            • HeyIamNuria
              Member
              • Dec 2012
              • 19

              #7
              Thank you for your help.
              I manage to get my library classified, but I still have another doubt.

              When I use Repeatmasker with a custom library (-lib) the .tbl file seems to be made with the results of masking the genome with
              Homo sapies repeats alone ("The query species was assumed to be homo"). My guess is that I can't obtain a .tbl output with a custom library (even if this library is made by RepeatModeler). Is this true?

              Thank you again


              Nuria
              Last edited by HeyIamNuria; 06-21-2013, 06:18 AM. Reason: The problem was not clearly explained.

              Comment

              • HeyIamNuria
                Member
                • Dec 2012
                • 19

                #8
                If the answer to my previous question is yes (so I can't get a real .tbl output when using RepeatMasker with the -lib option) is there any script that can produce something similar to the .tbl output with RepeatMasker .out or .gff files?

                (My library is classified with #class/subclass identifiers).


                Thank you for your time

                Nuria

                Comment

                • rhubley
                  Member
                  • Sep 2012
                  • 10

                  #9
                  Yes. You can use the RepeatMasker/util/buildSummary.pl script to reprocess the *.out file and summarize the results. The buildSummary.pl creates a similar table to the *.tbl file but also includes tabulations of each individual repeat and each seq/chr of the input.

                  Comment

                  • HeyIamNuria
                    Member
                    • Dec 2012
                    • 19

                    #10
                    Thank you, buildSummary seems to be what I was looking for, but I have another question.

                    I masked the same genome with different libraries and I used the buildSummary script to extract the repeat information from the *.out files.

                    But every .tbl file considers a different genome length. It is a number between the original total length and the excluding N/X-runs value. In my case 10 to 20 Mb smaller that the original length.

                    Consequently the % of repeats have important differences.

                    Can anyone help?

                    Thank you very much

                    Nuria

                    Comment

                    • HeyIamNuria
                      Member
                      • Dec 2012
                      • 19

                      #11
                      ok, I only had to read the buildSummary.pl script.

                      If someone else has the same problem, I solved using:
                      perl buildSummary.pl -useAbsoluteGenomeSize -genome tablewithscaffoldsnamesandlength.tsv -species yourspecies RMaskeroutput.out

                      Nuria

                      Comment

                      • Pseudoknot
                        Junior Member
                        • May 2012
                        • 4

                        #12
                        Hi Nuria,

                        I'm having the same problem. Just to be clear, you masked your genome using the RepeatModeler results and then used only the buildSummary.pl script to classify the repeats?

                        Thanks!

                        Comment

                        Latest Articles

                        Collapse

                        • SEQadmin2
                          Proteomic Platforms: How to Choose the Right Analytical Strategy to Improve Detection and Clinical Applications
                          by SEQadmin2


                          Proteomics platforms are evolving rapidly, with advances in mass spectrometry and affinity-based approaches expanding what researchers can detect and at what scale. As the field moves toward deeper proteome coverage and clinical applications, scientists face an increasingly complex landscape of tools. This article will explore how researchers are navigating these choices to find the right platform for their work.

                          The systematic characterization of the human proteome has
                          ...
                          07-20-2026, 11:48 AM
                        • SEQadmin2
                          Advanced Sequencing Platforms Tackle Neuroscience’s Toughest Genomics Problems
                          by SEQadmin2



                          Genomics studies in neuroscience face a special challenge due to the brain’s complexity and scarcity of samples. Mapping changes in cell type and state using conventional next-generation sequencing methods remains challenging. Advances in technologies like single-cell sequencing, spatial transcriptomics, and long-read sequencing have opened the door to deeper studies of the brain and diseases like Alzheimer’s, amyotrophic lateral sclerosis (ALS), and schizophrenia.
                          ...
                          07-09-2026, 11:10 AM
                        • SEQadmin2
                          Cancer Drug Resistance: The Lingering Barrier to Rising Survival
                          by SEQadmin2



                          Cancer survival rates have significantly increased in the last few decades in the United States, reaching a combined 70% 5-year survival rate by 2021. Behind this number, there are years of research to find new therapies, drug targets, and early detection methods. But there is one core challenge that keeps slowing down these advances, and it’s about drug resistance.

                          There is no single reason why many patients don’t respond to treatment as expected. Cancer is...
                          07-08-2026, 05:17 AM

                        ad_right_rmr

                        Collapse

                        News

                        Collapse

                        Topics Statistics Last Post
                        Started by SEQadmin2, 07-24-2026, 12:17 PM
                        0 responses
                        31 views
                        0 reactions
                        Last Post SEQadmin2  
                        Started by SEQadmin2, 07-23-2026, 11:41 AM
                        0 responses
                        23 views
                        0 reactions
                        Last Post SEQadmin2  
                        Started by SEQadmin2, 07-20-2026, 11:10 AM
                        0 responses
                        215 views
                        0 reactions
                        Last Post SEQadmin2  
                        Started by SEQadmin2, 07-13-2026, 10:26 AM
                        0 responses
                        79 views
                        0 reactions
                        Last Post SEQadmin2  
                        Working...