Data Analysis question

lre1234

Senior Member

Join Date: Aug 2011

Posts: 110
- Share
- Tweet
#1

Data Analysis question

11-30-2016, 01:15 PM

Hi all,
I'm going through some exome-seq data (10 samples), and following a "standard" pipeline - BWA, HaplotypeCaller, GenotypeGVCF, etc.. nothing too fancy.

My question is on the VQSR step. From my understanding, this works best when you have lots of samples to look at, but since I only have 10 samples, would it be better to simply do some hard-filtering on the dataset? Alternatively, I was thinking of bringing in some additional samples (such as a bunch of 1K genome exomes) and add these into the mix. This would increase the number of samples. Does anyone have any thoughts on this? Also, how many samples would be truly need to do the variant filtering with the VQSR rather than hard-filtering? (I'm having a little bit of a hard time finding what the minimum number of samples for it to work reliably).

Thanks for any advice
Tags: exome analsys, haplotypecaller, vqsr
vivek_

PhD Student

Join Date: Jul 2012

Posts: 164
- Share
- Tweet
#2

12-01-2016, 07:32 AM

Adding 1000 genomes samples might be counterintuitive because you probably already use them as a training set to build the VQSR classifier. There shouldn't be an issue with using hard filtering but regarding sample size for VQSR, you might ask the developers on the GATK forum to get the right answer.
Comment

Previous template Next

Topics	Statistics	Last Post
The Role of Spliceosomes in RNA Splicing and Genome Evolution by seqadmin Started by seqadmin, Today, 07:03 AM	0 responses 10 views 0 likes	Last Post by seqadmin Today, 07:03 AM
A Closer Look at the Enigmatic Genomes of Oikopleura dioica by seqadmin Started by seqadmin, 05-10-2024, 06:35 AM	0 responses 31 views 0 likes	Last Post by seqadmin 05-10-2024, 06:35 AM
Advanced Epigenome Editing Platform Explores Gene Regulation Mechanisms by seqadmin Started by seqadmin, 05-09-2024, 02:46 PM	0 responses 41 views 0 likes	Last Post by seqadmin 05-09-2024, 02:46 PM
Telomere Maintenance by PARP1: A New Perspective in Cancer Research by seqadmin Started by seqadmin, 05-07-2024, 06:57 AM	0 responses 33 views 0 likes	Last Post by seqadmin 05-07-2024, 06:57 AM

Seqanswers Leaderboard Ad