What is Genome-Wide Analysis of a Gene Family?
Genome-wide analysis of a gene family is a bioinformatics study in which we identify and characterize all genes belonging to a particular gene family across the entire genome of an organism.
For example, if you study the NPR1 gene family in tomato, you search the complete tomato genome to find every NPR1-like gene, then analyze their structure, evolution, location, and possible functions.
Typical workflow
1. Genome data collection
Download the complete genome and protein sequences of the species.
2. Identification of gene-family members
Use known conserved domains, HMMER, BLAST, InterPro/Pfam, etc.
Remove false or incomplete candidates.
3. Physicochemical analysis
Protein length
Molecular weight
Isoelectric point (pI)
Subcellular localization
4. Phylogenetic analysis
Compare family members within the species and/or with other species.
Construct a phylogenetic tree.
5. Gene structure analysis
Compare exon–intron organization among family members.
6. Conserved motif analysis
Identify important conserved protein motifs using tools such as MEME.
7. Chromosomal distribution
Determine where each gene is located on the chromosomes.
8. Gene duplication analysis
Identify:
Tandem duplication
Segmental duplication
Whole-genome duplication-related copies
9. Synteny/collinearity analysis
Compare corresponding genomic regions between species or chromosomes.
10. Promoter/cis-element analysis
Examine upstream regulatory regions for elements associated with:
Stress
Hormones
Light
Pathogen response
Development
11. Expression analysis
Use RNA-seq, qRT-PCR, or public transcriptomic data to determine when and where family genes are expressed.
12. Functional prediction
Combine evolutionary, structural and expression information to propose functions for individual genes.