Seqanswers Leaderboard Ad

**loba17** · 10-17-2011, 07:39 AM

its mee again ... it seems not so many people use meta-idba?

I run the analysis with different kmin and max values and I compared, n50, n90, and number of contigs. It seems meta-idba works best for kmin=25 and kmax=75 in my case. It is, however, possible that I fused together reads that do not belong together. I guess I have to assembly the reads using the contigs as references and determine the level of variation.

**koadman** · 12-01-2011, 09:44 PM

Hi Loba,
I've used standard IDBA for isolate assembly and with values of k below 30 chimerism/misassembly becomes a problem even in isolates. I haven't examined the issue in metagenomes closely enough to say how much worse the problem becomes in that context, but my intuition says be careful. Unless whatever you're planning to do with the assembly is robust to chimerism, I would exercise great caution in setting k to anything below 35 or so. The bigger the better. Unfortunately you will be limited by read accuracy, since longer k means each k-mer is more likely to contain an error, and I have not heard of an error correction approach that is sensible in the context of metagenomics where low abundance k-mers are expected to be real and not just noise. But that issue hasn't stopped some people from attempting filtering of low-abundance kmers in metagenomic data:

404 Not Found

http://ivory.idyll.org/blog/jul-10/kmer-filtering

**seb567** · 04-05-2012, 12:09 PM

Greetings !

Originally posted by koadman View Post

Hi Loba,
I've used standard IDBA for isolate assembly and with values of k below 30 chimerism/misassembly becomes a problem even in isolates. I haven't examined the issue in metagenomes closely enough to say how much worse the problem becomes in that context, but my intuition says be careful. Unless whatever you're planning to do with the assembly is robust to chimerism, I would exercise great caution in setting k to anything below 35 or so. The bigger the better. Unfortunately you will be limited by read accuracy, since longer k means each k-mer is more likely to contain an error, and I have not heard of an error correction approach that is sensible in the context of metagenomics where low abundance k-mers are expected to be real and not just noise. But that issue hasn't stopped some people from attempting filtering of low-abundance kmers in metagenomic data:
http://ivory.idyll.org/blog/jul-10/kmer-filtering

Even with a k-mer length of 21, the chimeric contig rate is very low with the Ray assembler. We tested it on in-house data as well as on the 124 samples from Qin et al. 2010.

(We are preparing a paper about it.)

You can download Ray here.

We have a mailing list too.

How to use Ray:

HTML Code:

mpiexec -n 64 Ray \
 -k \
 31 \
 -p \
 Sample/ERR011142_1.fastq.gz \
 Sample/ERR011142_2.fastq.gz \
 -p \
 Sample/ERR011143_1.fastq.gz \
 Sample/ERR011143_2.fastq.gz \
 -o \
 Assembly

Where mpiexec is a program that coordinates 64 instances of Ray an where Ray is the MPI-compatible Ray assemble executable.

Sébastien Boisvert

Topics	Statistics	Last Post
Cancer Metastasis: A Deep Dive into Cellular Plasticity by seqadmin Started by seqadmin, 04-11-2024, 12:08 PM	0 responses 59 views 0 likes	Last Post by seqadmin 04-11-2024, 12:08 PM
Proteogenomic Profiles Offer New Clues in Prostate Cancer by seqadmin Started by seqadmin, 04-10-2024, 10:19 PM	0 responses 57 views 0 likes	Last Post by seqadmin 04-10-2024, 10:19 PM
Novel Diagnostic Assay Enhances Ovarian Cancer Detection by seqadmin Started by seqadmin, 04-10-2024, 09:21 AM	0 responses 53 views 0 likes	Last Post by seqadmin 04-10-2024, 09:21 AM
Evolutionary Dynamics of Centromeres: A Comparative Genomic Analysis by seqadmin Started by seqadmin, 04-04-2024, 09:00 AM	0 responses 56 views 0 likes	Last Post by seqadmin 04-04-2024, 09:00 AM

Seqanswers Leaderboard Ad

Announcement

k-range in Meta-IDBA

Comment

Comment

Comment

Latest Articles

ad_right_rmr

News