RepeatExplorer2

All repeat classes Galaxy CLI

Graph-based identification and quantification of repeats from unassembled sequencing reads

Groups unassembled sequencing reads into clusters by sequence similarity, using a graph representation. Each cluster corresponds to a repeat family, and the number of reads it contains estimates that family’s share of the genome. No assembly is needed, and low coverage is sufficient.

Output

  • An HTML report with one entry per cluster
  • Cluster graph layouts, whose shape is diagnostic of the repeat type
  • Repeat classification based on similarity to known elements and to REXdb
  • Estimated genome proportion for each cluster and for each repeat class

How to cite

  • Novák P., Neumann P., Pech J., Steinhaisl J., Macas J. (2013) RepeatExplorer: a Galaxy-based web server for genome-wide characterization of eukaryotic repetitive elements from next-generation sequence reads. Bioinformatics 29: 792-793. doi:10.1093/bioinformatics/btt054
  • Novák P., Neumann P., Macas J. (2020) Global analysis of repetitive DNA from unassembled sequence reads using RepeatExplorer2. Nature Protocols 15: 3745-3776. doi:10.1038/s41596-020-0400-y
  • Novák P., Neumann P., Macas J. (2010) Graph-based clustering and characterization of repetitive sequences in next-generation sequencing data. BMC Bioinformatics 11. doi:10.1186/1471-2105-11-378