cross-species data - questions about normalization

trelek2

Junior Member

Join Date: May 2013

Posts: 2
- Share
- Tweet
#1

cross-species data - questions about normalization

05-23-2013, 06:14 AM

Hi,

I have some data form various samples (cell types) in different species.
I want to compare and analyze gene expression variability across the different species.

I've plotted the average expression (tags per million 'tpm' data) for each sample. I found that the average expression is more or less the same for all samples from a given species but the average expression values vary across species greatly (for example all human samples are about 20 and all mouse samples are about 25 but all dog samples are at about 100 tmp).

I'm guessing that this is because the data is not normalized and therefore the data is not comparable before performing a normalization.

I tried using different recently developed methods for data normalization (RLE, TMM) implemented in EdgeR as well as DeSeq and normalized the data from different species (by entrez id) separately for each cell type. This however still does not give much more similar average expressions for the different species. The averages of samples from the same species have now become much more dissimilar.

The only samples that went to a similar average expression level across the different species is Universal RNA - which is a mix of different tissues rather than a specific cell type.

I'm really confused, does the above mean the normalization is fine and I shouldn't worry about the fact that average expression values between different species are dissimilar, or does this rather imply something is wrong?

Maybe I should normalize by ignoring the gene ids and just looking at the whole expression profile (negative binomial) by sorting the genes according to their expression values and then having a different normalization factor for every expression value, whichever gene it is in every sample, so that I end up with 2 negative binomial curves that look the same but the genes on the x axis will be differently ordered depending on the species?

thanks.
Tags: edger, normalization, rle, species, tmm
mbblack

Senior Member

Join Date: Aug 2009

Posts: 245
- Share
- Tweet
#2

05-23-2013, 07:29 AM

Does each species have it's own controls? In other words, can each species be analyzed as an independent experiment?

If so, you could normalize each independently, analyze differential expression in each independently, and then simply compare homologous genes amongst the significantly differentially expressed genes in each species. I would suggest basing significance in each of the three independent analyses by simultaniously applying a statistical and a fold change cutoff to get the most robust differential gene lists for each.

That's how I've dealt with comparisons between, for example, rat and human hepatocytes exposed to dioxin. Set up the experiments separately and analyzed each species as an independent differential gene expression experiment, then compare the analyzed gene lists for significant homologous genes in each.

Otherwise, you have a very complex normalization situation where you may have to perform independent normalizations for each species followed by some form of scaling correction/normalization for the meta-analysis of all three.

Last edited by mbblack; 05-23-2013, 07:32 AM.

Michael Black, Ph.D.
ScitoVation LLC. RTP, N.C.
Comment

Previous template Next

Essential Discoveries and Tools in Epitranscriptomics

by seqadmin

The field of epigenetics has traditionally concentrated more on DNA and how changes like methylation and phosphorylation of histones impact gene expression and regulation. However, our increased understanding of RNA modifications and their importance in cellular processes has led to a rise in epitranscriptomics research. “Epitranscriptomics brings together the concepts of epigenetics and gene expression,” explained Adrien Leger, PhD, Principal Research Scientist...
- Channel: Articles
04-22-2024, 07:01 AM
Current Approaches to Protein Sequencing

by seqadmin

Proteins are often described as the workhorses of the cell, and identifying their sequences is key to understanding their role in biological processes and disease. Currently, the most common technique used to determine protein sequences is mass spectrometry. While still a valuable tool, mass spectrometry faces several limitations and requires a highly experienced scientist familiar with the equipment to operate it. Additionally, other proteomic methods, like affinity assays, are constrained...
- Channel: Articles
04-04-2024, 04:25 PM

Topics	Statistics	Last Post
Expanding the Horizons of Cellular Research with the Single Cell Atlas by seqadmin Started by seqadmin, Today, 11:49 AM	0 responses 13 views 0 likes	Last Post by seqadmin Today, 11:49 AM
Genetic Variants and Diabetes Risk in Childhood Cancer Survivors by seqadmin Started by seqadmin, Yesterday, 08:47 AM	0 responses 16 views 0 likes	Last Post by seqadmin Yesterday, 08:47 AM
Cancer Metastasis: A Deep Dive into Cellular Plasticity by seqadmin Started by seqadmin, 04-11-2024, 12:08 PM	0 responses 61 views 0 likes	Last Post by seqadmin 04-11-2024, 12:08 PM
Proteogenomic Profiles Offer New Clues in Prostate Cancer by seqadmin Started by seqadmin, 04-10-2024, 10:19 PM	0 responses 60 views 0 likes	Last Post by seqadmin 04-10-2024, 10:19 PM

Seqanswers Leaderboard Ad

Announcement

cross-species data - questions about normalization

Comment

Latest Articles

ad_right_rmr

News