Discrepancy between HTSeq-count counts and total mapped reads

M_staats

Junior Member

Join Date: Mar 2013

Posts: 1
- Share
- Tweet
#1

Discrepancy between HTSeq-count counts and total mapped reads

03-21-2013, 05:13 AM

Hi All,

There seems to be a rather large discrepancy between the number of mapped Illumina reads and the total number of HTSeq counts (gene counts + counts with no_feature + ambiguous counts + too_low_aQual + not_aligned + alignment_not_unique).

For instance:
I used Tophat2 to map paired-end illumina reads. My two FASTQ files contain a total of 39785174 reads of which 4177881 were unmapped (unmapped.bam). A total of 35607293 reads were mapped (accepted_hits.bam), of which 31390707 were uniquely mapped (grep -w "NH:i:1" accepted_hits.sam | sort | uniq | wc -l).

My input HTSeq-count command was:
htseq-count --mode=union --stranded=no --type=exon --idattr=gene_id accepted_hits.nsorted.sam Solanum_tuberosum.3.0.17.gtf > HTSeq_counts_UNION_ENSEMBL_GTF.out

The HTSeq output is:
Total sum of gene counts 13319749
no_feature 3150201
ambiguous 346347
too_low_aQual 0
not_aligned 0
alignment_not_unique 9024735

The sum of HTSeq-count is 25841032 counts. Shouldn't the sum of counts be equal to the total number of mapped reads (35607293)? Please help me explain this discrepancy.

Thanks in advance for the answer!
Best,

Last edited by M_staats; 03-21-2013, 05:16 AM.
Tags: htseq-count
Simon Anders

Senior Member

Join Date: Feb 2010

Posts: 995
- Share
- Tweet
#2

03-21-2013, 06:16 AM

htseq-count counts read pairs, not reads. Maybe this explains it.
Comment

Previous template Next

Nine Things a Sample Prep Scientist Thinks About Before Sequencing

by SEQadmin2

I’m not a sequencing expert. I’m a purification scientist who uses NGS to evaluate workflows my group develops. With this perspective, we think about the sample first and the NGS workflow second. The sequencer is an exceptionally honest reporter, but it can only report on what you give it, so whether you get clean, interpretable data from an NGS workflow is largely determined before you begin.

Here are nine questions we think about, in roughly the order they matter, before...
- Channel: Articles
06-18-2026, 07:11 AM
From Collection to Sequencing: Why Sample Preparation and Preservation Define Sequencing Data

by SEQadmin2

Data variability is still an issue in sequencing technologies despite the advances in reproducibility and accuracy of these platforms. But the problem does not originate in the sequencing itself, but in the previous steps, before the sample reaches the sequencer.

The first step is collection, followed by preservation and sample preparation for analysis. Most scientists overlook those steps, but not being careful might just be skewing the experiment’s results.
...
- Channel: Articles
06-02-2026, 10:05 AM

Topics	Statistics	Last Post
Large-Scale Protein Screen Uncovers Hidden Regulators of Alternative Polyadenylation by SEQadmin2 Started by SEQadmin2, 06-26-2026, 11:10 AM	0 responses 16 views 0 reactions	Last Post by SEQadmin2 06-26-2026, 11:10 AM
Whole-Genome Sequencing Traces Faroe Islands Ancestry to a North Atlantic Founder Population by SEQadmin2 Started by SEQadmin2, 06-17-2026, 06:09 AM	0 responses 49 views 0 reactions	Last Post by SEQadmin2 06-17-2026, 06:09 AM
Sequencing the Two-Toed Sloth Genome Reveals Jumping Genes Tied to Its Extreme Metabolism by SEQadmin2 Started by SEQadmin2, 06-09-2026, 11:58 AM	0 responses 109 views 0 reactions	Last Post by SEQadmin2 06-09-2026, 11:58 AM
A New Method Makes Hantavirus Genome Analysis Faster and More Accessible by SEQadmin2 Started by SEQadmin2, 06-05-2026, 10:09 AM	0 responses 125 views 0 reactions	Last Post by SEQadmin2 06-05-2026, 10:09 AM

Unconfigured Ad

Discrepancy between HTSeq-count counts and total mapped reads

Comment

Latest Articles

ad_right_rmr

News