Go Back   SEQanswers > Bioinformatics > Bioinformatics

Similar Threads
Thread Thread Starter Forum Replies Last Post
How does Bowtie identify paired end reads? Volklor RNA Sequencing 3 04-02-2012 12:00 PM
Samtools Flag to identify outward facing (mate-pair not paired end) reads lesander Bioinformatics 5 10-28-2011 01:24 PM
size of overhang after sonication? comingme Illumina/Solexa 10 10-14-2011 07:45 AM
how to identify paired end from qseq or fastq zhaowei Bioinformatics 1 02-02-2011 12:46 PM
Application of paired end short read library to identify promoter region Moondust Illumina/Solexa 1 07-01-2009 08:46 AM

Thread Tools
Old 04-03-2012, 06:16 AM   #1
Graham Etherington
Location: Norwich, England.

Join Date: Apr 2010
Posts: 22
Default How do I identify reads that overhang the end of a reference?

I have 76nt single-end reads and I want to map them to a reference and identify the reads that overhang the end of the reference.
What would be the best mapper to do the mapping with?
If I use BWA as the mapper, I can find reads that have up to about 9nt overhangs (by comparing the reads' unclipped start site to the length of the reference the read maps to), but what I'm looking for is some way to find reads that overhang by about 40 nts, i.e. the first 37nts of the read match the end of the reference, but (obviously) the last 40nts don't.
Do any next-gen mappers allow such mapping?
Graham Etherington is offline   Reply With Quote
Old 04-03-2012, 08:58 AM   #2
Peter (Biopython etc)
Location: Dundee, Scotland, UK

Join Date: Jul 2009
Posts: 1,543

Well the SAM/BAM file format itself doesn't allow this directly (the start/end co-ordinates must be within the length of the reference), so any reads which did map off the ends would probably have to be represented with a large insert or soft clip at the end.

I'm not sure of any specialized tool to do this. What I would consider is trying to map the first and last halves of your reads to the start/end of the reference. i.e. Some scripting to prepare the modified input files, then reuse a mapper like BWA, and interpret the results accordingly. Fiddly - so I hope someone else has done the hard work already!
maubp is offline   Reply With Quote
Old 06-30-2014, 09:38 PM   #3
Location: Belgium

Join Date: Nov 2008
Posts: 79

Dear all,

I want to re-open this discussion as I have a similar question: how can I detect and extract overhanging reads, and this at both ends of my reference?
I have overlapping MiSeq reads which I merged, resulting in reads with an average length of 300 - 350 bp. I'm only interested in what we call hybrid reads: reads partly mapping with the borders of my reference.
Any advice on a mapping tool? Even use Blast?
Thanks for the input.
strob is offline   Reply With Quote
Old 07-01-2014, 09:22 AM   #4
Brian Bushnell
Super Moderator
Location: Walnut Creek, CA

Join Date: Jan 2014
Posts: 2,707


If you run BBMap with default parameters, it will soft-clip ONLY reads that map off the ends of reference scaffolds. So, you would just have to filter the sam file for reads containing the "S" symbol in their cigar strings.

You can adjust the fraction of the read that must be on the reference with the "minratio" flag (default 0.56). I suggest something like this:

minratio=0.2 maxindel=0

...which will look for alignments with as little as 20% of the read over the reference and not look for alignments containing indels.
Brian Bushnell is offline   Reply With Quote

bwa, mapping, overhang, reference

Thread Tools

Posting Rules
You may not post new threads
You may not post replies
You may not post attachments
You may not edit your posts

BB code is On
Smilies are On
[IMG] code is On
HTML code is Off

All times are GMT -8. The time now is 12:31 PM.

Powered by vBulletin® Version 3.8.9
Copyright ©2000 - 2020, vBulletin Solutions, Inc.
Single Sign On provided by vBSSO