Genome data collection — NCBI/GenBank, RefSeq, GISAIDYes. A Q1-level, entirely in-silico paper on human virus genomes is feasible, but the key is to go beyond a routine “download genomes → make phylogenetic tree” paper. You need a well-defined biological question, large/high-quality dataset, multiple complementary analyses, and a genuinely novel conclusion.
For current taxonomy, I would use the latest ICTV release (MSL41/2025–2026) rather than relying on older virus classifications.
A strong overall workflow
1. Select the virus
For your first project, I would consider viruses with:
hundreds/thousands of publicly available genomes
good genome annotation
sufficient geographic/temporal metadata
substantial evolutionary diversity
clear relevance to human health
Good candidates include:
Virus Genome In-silico potential
SARS-CoV-2 +ssRNA ⭐⭐⭐⭐⭐
Influenza A segmented -ssRNA ⭐⭐⭐⭐⭐
Dengue virus +ssRNA ⭐⭐⭐⭐⭐
Hepatitis B virus partially dsDNA ⭐⭐⭐⭐⭐
Hepatitis C virus +ssRNA ⭐⭐⭐⭐
HIV-1 +ssRNA/RT ⭐⭐⭐⭐⭐
HPV dsDNA ⭐⭐⭐⭐⭐
Human adenoviruses dsDNA ⭐⭐⭐⭐
Enteroviruses +ssRNA ⭐⭐⭐⭐⭐
Norovirus +ssRNA ⭐⭐⭐⭐
Rotavirus segmented dsRNA ⭐⭐⭐⭐
Human herpesviruses dsDNA ⭐⭐⭐⭐⭐
For a novel Q1 paper, I would not automatically choose SARS-CoV-2 because the literature is already enormous. A less saturated virus/family can make novelty easier.
—
2. Define a strong research question
Instead of:
> “Genome-wide analysis of X virus”
use a question such as:
> “Comparative genomic, evolutionary, structural and selection analysis of X virus reveals conserved functional regions and lineage-specific adaptive signatures.”
Or:
> “Genome-wide comparative analysis of X virus isolates reveals evolutionary hotspots, conserved regions and host/geographic adaptation.”
This gives you several layers of analysis.
—
3. Genome dataset construction
Download complete genomes from:
NCBI GenBank/RefSeq
Virus databases
GISAID where appropriate and permitted
ICTV taxonomy/metadata
ICTV provides the current official taxonomy and downloadable Master Species List.
For each genome, collect:
Accession | strain/isolate | collection date | country | host | genome length | completeness | publication
Aim for approximately:
500–5,000 genomes, depending on the virus.
Avoid simply downloading everything. Establish inclusion/exclusion criteria.
—
4. Quality control
Remove:
incomplete genomes
excessive Ns
duplicate sequences
very short sequences
poorly annotated genomes
obvious contamination
sequences with problematic metadata
For phylogenetic analysis, homologous sequences should be sufficiently similar in length and alignable; mixing very short and very long sequences can produce unreliable analyses.
—
5. Genome annotation
For each genome determine:
Genome architecture
genome length
ORFs
coding regions
non-coding regions
overlapping genes
regulatory regions
structural proteins
non-structural proteins
Create a genome map.
—
6. Protein/gene-level comparative genomics
For every major gene/protein calculate:
nucleotide identity
amino-acid identity
conserved residues
variable residues
insertion/deletion regions
domain architecture
protein length variation
Then identify:
Highly conserved genes
Potentially important for:
replication
transcription
structural integrity
essential viral functions
Highly variable genes
Potentially associated with:
immune interaction
host adaptation
lineage diversification
—
7. Multiple sequence alignment
Use:
MAFFT
for nucleotide and/or protein alignment.
For large datasets, automate the process with Python/Linux.
You can also use Nextclade for viral genome alignment, mutation calling, clade assignment, quality control and phylogenetic placement.
—
8. Genome-wide variation analysis
This can become one of the major figures.
Calculate:
SNPs
substitutions
insertions
deletions
mutation frequency
nucleotide diversity
entropy
conserved regions
hypervariable regions
Generate sliding-window plots:
Genome position → nucleotide diversity
and
Genome position → mutation frequency
—
9. Phylogenetic analysis
This is essential but should not be your only analysis.
Pipeline:
MAFFT → alignment → model selection → IQ-TREE → phylogenetic tree
Nextstrain/Augur also provides workflows for viral genomic analysis, and Augur supports MAFFT alignment and IQ-TREE/other tree-building approaches.
Analyze:
major clades
lineage structure
geographic clustering
temporal clustering
ancestral relationships
—
10. Time-scaled evolutionary analysis
If collection dates are available:
BEAST / TreeTime / Nextstrain
can investigate:
evolutionary rate
time to most recent common ancestor
lineage emergence
temporal diversification
This can give substantially more biological depth than a simple phylogenetic tree.
—
11. Selection pressure analysis ⭐
This is particularly important for a strong paper.
Calculate:
dN/dS
and identify:
Positive selection
Candidate adaptive sites.
Purifying selection
Highly conserved functional regions.
Possible approaches:
HyPhy
FEL
MEME
FUBAR
SLAC
Then map significant sites onto proteins/domains.
This gives you a potentially strong result:
> “Positive-selection hotspots occur predominantly in specific functional domains, whereas replication-associated proteins show strong purifying selection.”
—
12. Protein structural analysis
For important proteins:
Sequence → protein structure → mutation mapping
Use:
AlphaFold/AlphaFold databases
PDB
PyMOL
ChimeraX
Map:
conserved residues
positively selected sites
frequent mutations
domain boundaries
Then ask:
Are evolutionary hotspots located near functional/structural regions?
That is much more interesting than simply reporting mutations.
—
13. Functional annotation
Use:
InterPro
Pfam
UniProt
Gene Ontology
KEGG where appropriate
Identify:
protein domains
catalytic regions
binding sites
conserved motifs
functional categories
—
14. Host-interaction analysis
This can make the study more biologically meaningful.
For selected viral proteins:
viral protein → predicted/known host interaction → functional pathway
Investigate:
host receptors
immune-related proteins
interferon pathways
apoptosis
inflammatory pathways
cell signaling
Use experimentally supported interaction databases where possible rather than relying only on predictions.
—
15. Mutation–structure–function integration
This is where I would try to make the paper Q1-quality.
Create an integrated analysis:
Genome variation
↓
Protein mutation
↓
Selection pressure
↓
Protein structure
↓
Functional domain
↓
Potential biological significance
For example:
> Frequent mutations → positively selected sites → located in a functional domain → structurally exposed → potentially relevant to host interaction.
Important: phrase biological consequences as predictions/hypotheses unless experimentally validated.
—
16. Geographic analysis
If metadata are available:
Country → lineage → mutations → phylogeny
You can investigate:
geographic clustering
country-associated variants
lineage distribution
mutation frequencies by region
Use appropriate statistical testing rather than simply showing maps.
—
17. Temporal analysis
Divide genomes into periods, for example:
2015–2017 → 2018–2020 → 2021–2023 → 2024–2026
Then compare:
mutation frequency
diversity
clade composition
selection
genome regions under changing evolutionary pressure
The exact periods should be determined by the virus’s sampling history.
—
18. Statistical analysis
Don’t stop at bioinformatics figures.
Include:
nucleotide diversity
Tajima’s D where appropriate
dN/dS
FST for suitable population comparisons
AMOVA where appropriate
correlation analyses
regression models
permutation/bootstrap testing
multiple-testing correction
—
19. Machine-learning component — optional but powerful
Since you are interested in Python/AI, this could be a strong addition.
For example:
Input
Genome/protein features:
mutation profile
k-mers
amino-acid substitutions
genomic regions
lineage
collection date
geographic origin
ML
Random Forest
XGBoost
SVM
clustering
PCA
UMAP
Question
Can genomic features distinguish:
lineages / geographic groups / temporal groups?
Use proper train/test separation and avoid leakage.
—
20. Build an integrated figure
A very strong conceptual figure could be:
Human virus genomes
↓
Genome dataset
↓
QC
↓
Multiple sequence alignment
↓
Genome-wide variation
↓
Phylogeny
↓
Temporal evolution
↓
Selection analysis
↓
Protein structure
↓
Functional annotation
↓
Geographic/host association
↓
Integrated evolutionary model
—
21. Suggested paper structure
Title
Comparative genomic and evolutionary analysis of [Virus] reveals conserved functional regions and lineage-specific adaptive signatures
Figure 1
Genome organization and dataset composition.
Figure 2
Genome-wide nucleotide diversity/mutation landscape.
Figure 3
Phylogenetic tree.
Figure 4
Temporal/geographic evolution.
Figure 5
Positive and purifying selection.
Figure 6
Protein structure + selected/conserved residues.
Figure 7
Integrated genotype–structure–function analysis.
Figure 8
Proposed evolutionary model.
—
22. Reproducibility is critical
For a serious computational paper, maintain:
Raw data
→ accession list
→ QC scripts
→ filtered dataset
→ alignment
→ analysis scripts
→ statistical results
→ figures
→ supplementary tables
Ideally provide:
GitHub repository
analysis workflow
exact software versions
parameters
accession numbers
supplementary datasets
Nextstrain itself emphasizes reproducible computational workflows and provides open-source tools for pathogen genome analysis.
—
23. My recommended strategy for you
Given your bioinformatics/plant-virus background, I would build a general human-virus comparative genomics pipeline that you can subsequently apply to several viruses.
Phase I — Dataset
NCBI/other public databases
↓
500–5,000 complete genomes
↓
QC
Phase II — Genome analysis
Genome architecture
↓
ORFs
↓
MAFFT
↓
SNP/indel analysis
↓
conservation/diversity
Phase III — Evolution
Phylogeny
↓
clades
↓
geographic distribution
↓
temporal evolution
↓
evolutionary rate
Phase IV — Selection
dN/dS
↓
positive selection
↓
purifying selection
↓
selection hotspots
Phase V — Protein biology
Domains
↓
structures
↓
mutation mapping
↓
functional sites
Phase VI — Advanced analysis
ML/clustering
↓
genotype–phenotype/metadata associations
↓
integrated evolutionary model
Phase VII — Publication
Figures
↓
statistics
↓
supplementary data
↓
GitHub/reproducible workflow
↓
target Q1 journal
—
The strongest approach is:
> Large dataset + rigorous QC + comparative genomics + phylogeny + temporal/geographic evolution + selection + structural mapping + functional interpretation + reproducible computational pipeline.
Enlist of all procedures
Genome quality control — completeness, duplicates, ambiguous bases
Genome annotation — ORFs and coding regions
Multiple sequence alignment — MAFFT
Genome-wide SNP/Indel analysis
Conservation and nucleotide diversity analysis
Mutation frequency analysis
Phylogenetic analysis — IQ-TREE/MEGA
Time-scaled phylogenetics — BEAST/TreeTime
Selection pressure analysis — dN/dS, HyPhy
Positive-selection site detection — MEME/FEL/FUBAR
Haplotype/lineage analysis
Geographic and temporal distribution analysis
Protein/domain analysis — InterPro/Pfam/UniProt
Protein structure prediction — AlphaFold
Mutation-to-structure mapping — PyMOL/ChimeraX
Host–virus interaction analysis
Functional enrichment/pathway analysis
Statistical analysis — correlation, clustering, PCA/UMAP
Machine-learning classification/clustering — Random Forest/XGBoost/SVM
Network analysis — mutation/protein interaction networks
Integrated evolutionary analysis
Reproducible computational pipeline — Python/R + GitHub