Human Virus genome research guidline

Genome data collection — NCBI/GenBank, RefSeq, GISAIDYes. A Q1-level, entirely in-silico paper on human virus genomes is feasible, but the key is to go beyond a routine “download genomes → make phylogenetic tree” paper. You need a well-defined biological question, large/high-quality dataset, multiple complementary analyses, and a genuinely novel conclusion.

For current taxonomy, I would use the latest ICTV release (MSL41/2025–2026) rather than relying on older virus classifications.

A strong overall workflow

1. Select the virus

For your first project, I would consider viruses with:

hundreds/thousands of publicly available genomes

good genome annotation

sufficient geographic/temporal metadata

substantial evolutionary diversity

clear relevance to human health

Good candidates include:

Virus Genome In-silico potential

SARS-CoV-2 +ssRNA ⭐⭐⭐⭐⭐
Influenza A segmented -ssRNA ⭐⭐⭐⭐⭐
Dengue virus +ssRNA ⭐⭐⭐⭐⭐
Hepatitis B virus partially dsDNA ⭐⭐⭐⭐⭐
Hepatitis C virus +ssRNA ⭐⭐⭐⭐
HIV-1 +ssRNA/RT ⭐⭐⭐⭐⭐
HPV dsDNA ⭐⭐⭐⭐⭐
Human adenoviruses dsDNA ⭐⭐⭐⭐
Enteroviruses +ssRNA ⭐⭐⭐⭐⭐
Norovirus +ssRNA ⭐⭐⭐⭐
Rotavirus segmented dsRNA ⭐⭐⭐⭐
Human herpesviruses dsDNA ⭐⭐⭐⭐⭐

For a novel Q1 paper, I would not automatically choose SARS-CoV-2 because the literature is already enormous. A less saturated virus/family can make novelty easier.

—

2. Define a strong research question

Instead of:

> “Genome-wide analysis of X virus”

use a question such as:

> “Comparative genomic, evolutionary, structural and selection analysis of X virus reveals conserved functional regions and lineage-specific adaptive signatures.”

Or:

> “Genome-wide comparative analysis of X virus isolates reveals evolutionary hotspots, conserved regions and host/geographic adaptation.”

This gives you several layers of analysis.

—

3. Genome dataset construction

Download complete genomes from:

NCBI GenBank/RefSeq

Virus databases

GISAID where appropriate and permitted

ICTV taxonomy/metadata

ICTV provides the current official taxonomy and downloadable Master Species List.

For each genome, collect:

Accession | strain/isolate | collection date | country | host | genome length | completeness | publication

Aim for approximately:

500–5,000 genomes, depending on the virus.

Avoid simply downloading everything. Establish inclusion/exclusion criteria.

—

4. Quality control

Remove:

incomplete genomes

excessive Ns

duplicate sequences

very short sequences

poorly annotated genomes

obvious contamination

sequences with problematic metadata

For phylogenetic analysis, homologous sequences should be sufficiently similar in length and alignable; mixing very short and very long sequences can produce unreliable analyses.

—

5. Genome annotation

For each genome determine:

Genome architecture

genome length

ORFs

coding regions

non-coding regions

overlapping genes

regulatory regions

structural proteins

non-structural proteins

Create a genome map.

—

6. Protein/gene-level comparative genomics

For every major gene/protein calculate:

nucleotide identity

amino-acid identity

conserved residues

variable residues

insertion/deletion regions

domain architecture

protein length variation

Then identify:

Highly conserved genes

Potentially important for:

replication

transcription

structural integrity

essential viral functions

Highly variable genes

Potentially associated with:

immune interaction

host adaptation

lineage diversification

—

7. Multiple sequence alignment

Use:

MAFFT

for nucleotide and/or protein alignment.

For large datasets, automate the process with Python/Linux.

You can also use Nextclade for viral genome alignment, mutation calling, clade assignment, quality control and phylogenetic placement.

—

8. Genome-wide variation analysis

This can become one of the major figures.

Calculate:

SNPs

substitutions

insertions

deletions

mutation frequency

nucleotide diversity

entropy

conserved regions

hypervariable regions

Generate sliding-window plots:

Genome position → nucleotide diversity

and

Genome position → mutation frequency

—

9. Phylogenetic analysis

This is essential but should not be your only analysis.

Pipeline:

MAFFT → alignment → model selection → IQ-TREE → phylogenetic tree

Nextstrain/Augur also provides workflows for viral genomic analysis, and Augur supports MAFFT alignment and IQ-TREE/other tree-building approaches.

Analyze:

major clades

lineage structure

geographic clustering

temporal clustering

ancestral relationships

—

10. Time-scaled evolutionary analysis

If collection dates are available:

BEAST / TreeTime / Nextstrain

can investigate:

evolutionary rate

time to most recent common ancestor

lineage emergence

temporal diversification

This can give substantially more biological depth than a simple phylogenetic tree.

—

11. Selection pressure analysis ⭐

This is particularly important for a strong paper.

Calculate:

dN/dS

and identify:

Positive selection

Candidate adaptive sites.

Purifying selection

Highly conserved functional regions.

Possible approaches:

HyPhy

FEL

MEME

FUBAR

SLAC

Then map significant sites onto proteins/domains.

This gives you a potentially strong result:

> “Positive-selection hotspots occur predominantly in specific functional domains, whereas replication-associated proteins show strong purifying selection.”

—

12. Protein structural analysis

For important proteins:

Sequence → protein structure → mutation mapping

Use:

AlphaFold/AlphaFold databases

PDB

PyMOL

ChimeraX

Map:

conserved residues

positively selected sites

frequent mutations

domain boundaries

Then ask:

Are evolutionary hotspots located near functional/structural regions?

That is much more interesting than simply reporting mutations.

—

13. Functional annotation

Use:

InterPro

Pfam

UniProt

Gene Ontology

KEGG where appropriate

Identify:

protein domains

catalytic regions

binding sites

conserved motifs

functional categories

—

14. Host-interaction analysis

This can make the study more biologically meaningful.

For selected viral proteins:

viral protein → predicted/known host interaction → functional pathway

Investigate:

host receptors

immune-related proteins

interferon pathways

apoptosis

inflammatory pathways

cell signaling

Use experimentally supported interaction databases where possible rather than relying only on predictions.

—

15. Mutation–structure–function integration

This is where I would try to make the paper Q1-quality.

Create an integrated analysis:

Genome variation

↓

Protein mutation

↓

Selection pressure

↓

Protein structure

↓

Functional domain

↓

Potential biological significance

For example:

> Frequent mutations → positively selected sites → located in a functional domain → structurally exposed → potentially relevant to host interaction.

Important: phrase biological consequences as predictions/hypotheses unless experimentally validated.

—

16. Geographic analysis

If metadata are available:

Country → lineage → mutations → phylogeny

You can investigate:

geographic clustering

country-associated variants

lineage distribution

mutation frequencies by region

Use appropriate statistical testing rather than simply showing maps.

—

17. Temporal analysis

Divide genomes into periods, for example:

2015–2017 → 2018–2020 → 2021–2023 → 2024–2026

Then compare:

mutation frequency

diversity

clade composition

selection

genome regions under changing evolutionary pressure

The exact periods should be determined by the virus’s sampling history.

—

18. Statistical analysis

Don’t stop at bioinformatics figures.

Include:

nucleotide diversity

Tajima’s D where appropriate

dN/dS

FST for suitable population comparisons

AMOVA where appropriate

correlation analyses

regression models

permutation/bootstrap testing

multiple-testing correction

—

19. Machine-learning component — optional but powerful

Since you are interested in Python/AI, this could be a strong addition.

For example:

Input

Genome/protein features:

mutation profile

k-mers

amino-acid substitutions

genomic regions

lineage

collection date

geographic origin

ML

Random Forest

XGBoost

SVM

clustering

PCA

UMAP

Question

Can genomic features distinguish:

lineages / geographic groups / temporal groups?

Use proper train/test separation and avoid leakage.

—

20. Build an integrated figure

A very strong conceptual figure could be:

Human virus genomes

↓

Genome dataset

↓

QC

↓

Multiple sequence alignment

↓

Genome-wide variation

↓

Phylogeny

↓

Temporal evolution

↓

Selection analysis

↓

Protein structure

↓

Functional annotation

↓

Geographic/host association

↓

Integrated evolutionary model

—

21. Suggested paper structure

Title

Comparative genomic and evolutionary analysis of [Virus] reveals conserved functional regions and lineage-specific adaptive signatures

Figure 1

Genome organization and dataset composition.

Figure 2

Genome-wide nucleotide diversity/mutation landscape.

Figure 3

Phylogenetic tree.

Figure 4

Temporal/geographic evolution.

Figure 5

Positive and purifying selection.

Figure 6

Protein structure + selected/conserved residues.

Figure 7

Integrated genotype–structure–function analysis.

Figure 8

Proposed evolutionary model.

—

22. Reproducibility is critical

For a serious computational paper, maintain:

Raw data

→ accession list

→ QC scripts

→ filtered dataset

→ alignment

→ analysis scripts

→ statistical results

→ figures

→ supplementary tables

Ideally provide:

GitHub repository

analysis workflow

exact software versions

parameters

accession numbers

supplementary datasets

Nextstrain itself emphasizes reproducible computational workflows and provides open-source tools for pathogen genome analysis.

—

23. My recommended strategy for you

Given your bioinformatics/plant-virus background, I would build a general human-virus comparative genomics pipeline that you can subsequently apply to several viruses.

Phase I — Dataset

NCBI/other public databases

↓

500–5,000 complete genomes

↓

QC

Phase II — Genome analysis

Genome architecture
↓
ORFs
↓
MAFFT
↓
SNP/indel analysis
↓
conservation/diversity

Phase III — Evolution

Phylogeny
↓
clades
↓
geographic distribution
↓
temporal evolution
↓
evolutionary rate

Phase IV — Selection

dN/dS
↓
positive selection
↓
purifying selection
↓
selection hotspots

Phase V — Protein biology

Domains
↓
structures
↓
mutation mapping
↓
functional sites

Phase VI — Advanced analysis

ML/clustering
↓
genotype–phenotype/metadata associations
↓
integrated evolutionary model

Phase VII — Publication

Figures
↓
statistics
↓
supplementary data
↓
GitHub/reproducible workflow
↓
target Q1 journal

—

 

The strongest approach is:

> Large dataset + rigorous QC + comparative genomics + phylogeny + temporal/geographic evolution + selection + structural mapping + functional interpretation + reproducible computational pipeline.

Enlist of all procedures
Genome quality control — completeness, duplicates, ambiguous bases
Genome annotation — ORFs and coding regions
Multiple sequence alignment — MAFFT
Genome-wide SNP/Indel analysis
Conservation and nucleotide diversity analysis
Mutation frequency analysis
Phylogenetic analysis — IQ-TREE/MEGA
Time-scaled phylogenetics — BEAST/TreeTime
Selection pressure analysis — dN/dS, HyPhy
Positive-selection site detection — MEME/FEL/FUBAR
Haplotype/lineage analysis
Geographic and temporal distribution analysis
Protein/domain analysis — InterPro/Pfam/UniProt
Protein structure prediction — AlphaFold
Mutation-to-structure mapping — PyMOL/ChimeraX
Host–virus interaction analysis
Functional enrichment/pathway analysis
Statistical analysis — correlation, clustering, PCA/UMAP
Machine-learning classification/clustering — Random Forest/XGBoost/SVM
Network analysis — mutation/protein interaction networks
Integrated evolutionary analysis
Reproducible computational pipeline — Python/R + GitHub

 

PPT

Leave a Comment