That's a good point. Some filtering is necessary to take care of pileup of reads due to biases. I do that for alignment and SNP discovery, but think twice about it during de novo assembly. If no underlying genome is known, it is hard to tell whether the duplicated reads come from error or real sequence.
|