CHAPTER
Computational Analysis of Human Virus Genomes
A Practical Chapter for Undergraduate Students in Bioinformatics, Microbiology and Molecular Biology
Abstract
Viral
are the most abundant and the most rapidly changing sequences in public biological databases. Because they are short, numerous and richly annotated with dates and places of collection, they are also the single best training ground for learning comparative and evolutionary genomics. This chapter sets out a complete, reproducible workflow for the in-silico study of a human virus, from the choice of a study system through dataset assembly, quality control, annotation, alignment, measurement of variation, phylogenetic and time-scaled analysis, tests of natural selection, structural and functional mapping, host-virus interaction analysis, spatial and temporal comparison, statistical testing, optional machine learning, and finally the writing and archiving of the results. Each stage is presented with its underlying biological logic, the software most commonly used, the assumptions the method makes, and the errors that most often invalidate undergraduate projects. The chapter is intended to be worked through in order, and it closes with a glossary, self-assessment questions and a fully cited reference list.
Keywords: viral genomics; comparative genomics; molecular evolution; phylogenetics; molecular clock; natural selection; dN/dS; protein structure prediction; reproducible research; bioinformatics education.
| Learning Objectives
By the end of this chapter you should be able to: |
- Explain why viral genomes are especially suited to computational evolutionary study, and relate genome type to expected rate and mode of change.
- Assemble a defensible genome dataset from public repositories, with complete and auditable metadata.
- Apply quality-control criteria and justify every sequence that is removed.
- Produce and critically evaluate a multiple sequence alignment, and recognise when an alignment cannot support downstream inference.
- Quantify genome-wide variation using nucleotide diversity, Shannon entropy and sliding-window analysis.
- Build, root and interpret a phylogenetic tree, and state honestly what a tree can and cannot demonstrate.
- Estimate evolutionary rates and divergence times, and describe the assumptions of a molecular clock.
- Test for positive and purifying selection using dN/dS-based site models, and describe the principal limitations of the dN/dS statistic.
- Map conserved and positively selected residues onto a predicted protein structure and interpret the result as a hypothesis rather than a finding.
- Design appropriate statistical tests, correct for multiple testing, and avoid data leakage in any machine-learning component.
- Deposit code, data and documentation so that another student could reproduce every figure in the study.
1. Why Study Viral Genomes Computationally?
A virus is, from the point of view of a computer, an unusually convenient organism. Its genome is short – from roughly 1.7 kilobases for hepatitis delta virus to a few hundred kilobases for the largest herpesviruses – so an entire population of genomes can be aligned and analysed on an ordinary laptop. Its generation time is measured in hours, and for RNA viruses its mutation rate per site per replication is orders of magnitude higher than that of its host (Drake & Holland, 1999; Sanjuan & Domingo-Calap, 2016). The consequence is that evolution in viruses happens on a timescale we can actually observe: sequences sampled two years apart are measurably different, and the differences carry information about how the virus is adapting.
Public databases now hold millions of viral genome sequences, most of them annotated with a collection date, a country of origin and a host species (Sayers et al., 2024). This combination – large numbers, short sequences, rapid change and rich metadata – means that a well-designed undergraduate project can produce a genuine, publishable result without a laboratory bench, a freezer or a biosafety cabinet. It also means that the difference between a trivial study and a good one lies entirely in the design of the analysis.
That distinction is the theme of this chapter. It is very easy to download a few hundred sequences, align them, build a tree and describe the clusters. Thousands of papers do exactly this, and most of them add little. A study becomes valuable when several independent lines of computational evidence – the distribution of variation along the genome, the shape and timing of the phylogeny, the signature of natural selection, the position of variable residues in the three-dimensional protein, and the known function of the domain they fall in – all point towards the same biological conclusion. Building that convergence, layer by layer, is what the rest of this chapter teaches.
| Box 1. What ‘in silico’ does and does not mean
In-silico research means research conducted entirely by computation on existing data. It is a legitimate and mature branch of biology: the phylogenetic reconstruction of SARS-CoV-2 lineages, the discovery of drug-resistance mutations in HIV-1, and the structural prediction that won the 2024 Nobel Prize in Chemistry are all in-silico results. What in-silico work cannot do is demonstrate function. A computational analysis can show that a residue is conserved, that it is under positive selection, that it sits at a protein-protein interface, and that homologous residues in a related virus are known to affect receptor binding. Taken together that is a strong, testable hypothesis. It is not proof that the residue affects receptor binding in your virus. Undergraduate reports are marked down more often for overstating computational results than for any other single fault; the language of prediction (‘these results suggest’, ‘this is consistent with’, ‘we hypothesise’) is not weakness, it is accuracy. |
2. Biological Background You Need Before You Start
2.1 Genome type determines what you will see
Baltimore (1971) classified viruses by the nature of their genome and the route they take to messenger RNA. That classification is not merely taxonomic bookkeeping; it predicts how much variation you should expect to find and how you should interpret it. RNA viruses replicate using polymerases that generally lack proofreading activity, giving mutation rates of the order of 10-6 to 10-4 substitutions per nucleotide per replication cycle, whereas double-stranded DNA viruses use higher-fidelity polymerases and often the host replication machinery, and evolve several orders of magnitude more slowly (Duffy et al., 2008; Sanjuan & Domingo-Calap, 2016).
Table 1. Genome type, expected evolutionary rate and the analyses each supports.
| Genome type | Example human viruses | Typical substitution rate (subs/site/year) | What this means for your project |
| +ssRNA | SARS-CoV-2, dengue virus, enteroviruses, hepatitis C virus | 10-4 to 10-3 | Strong temporal signal over months to years; molecular clock analysis is reliable; expect abundant variation. |
| -ssRNA (segmented) | Influenza A virus | 10-3 (haemagglutinin) | Reassortment between segments is common; each segment must be analysed separately before any combined interpretation. |
| ssRNA with reverse transcription | HIV-1 | 10-3 to 10-2 (within host) | Extremely high diversity; recombination is frequent; within-host and between-host evolution must not be mixed. |
| Partially dsDNA with RT | Hepatitis B virus | 10-5 to 10-4 | Intermediate; clock signal detectable over decades; strong genotype structure. |
| dsDNA | Human papillomaviruses, human herpesviruses | 10-8 to 10-6 | Very slow; deep divergence and co-divergence with host populations; temporal signal usually absent over sampling timescales. |
The practical message of Table 1 is that the same workflow does not fit every virus. A time-scaled phylogenetic analysis of human papillomavirus using sequences collected over twenty years will fail, because twenty years is far too short for measurable change in a slowly evolving double-stranded DNA genome. The same analysis applied to dengue virus will succeed easily. Choose the analysis to fit the biology, not the other way round.
2.2 Populations, quasispecies and what a ‘genome sequence’ really is
An RNA virus does not exist inside a patient as a single sequence. It exists as a swarm of closely related variants, generated continuously by an error-prone polymerase and shaped by selection and drift – the quasispecies concept (Domingo et al., 2012). What you download from GenBank is usually a consensus sequence: the most common base at each position across that swarm. This is a summary, not an organism. It is entirely adequate for between-host comparative studies of the kind described in this chapter, but you should know that within-host diversity has been averaged away, and you should never mix consensus sequences with deep-sequencing variant data in the same alignment.
2.3 Recombination and reassortment: check before you build a tree
Phylogenetic methods assume that every site in your alignment shares a single evolutionary history. Recombination violates that assumption directly: different parts of a recombinant genome have different ancestors, and forcing them onto one tree produces a topology that is wrong in a way no amount of bootstrapping will reveal. Worse, undetected recombination systematically inflates estimates of positive selection, because sites that changed by recombination look like sites that changed by repeated substitution (Posada & Crandall, 2001).
Recombination is common in HIV-1, enteroviruses, dengue virus and coronaviruses, and segment reassortment is routine in influenza A. Screen for it before phylogenetic or selection analysis, using GARD (Kosakovsky Pond et al., 2006), RDP5 (Martin et al., 2021) or 3SEQ (Boni et al., 2007). If recombination is detected, either remove the recombinant sequences, or partition the alignment at the detected breakpoints and analyse each partition on its own tree. For segmented viruses, analyse each segment separately as a matter of course.
3. Choosing a Study System
The first decision determines the ceiling on everything that follows. A good target virus has five properties:
- Sufficient data. Several hundred to a few thousand complete genomes are enough for robust phylogenetics, diversity estimation and site-level selection analysis. Fewer than about fifty genomes leaves most statistical tests underpowered.
- Consistent annotation. Genomes deposited with coding-sequence features already defined save weeks of work and prevent frame errors.
- Rich metadata. Collection date, country and host are the raw material of the temporal, geographic and adaptation analyses. A genome lacking a date cannot enter a clock analysis; a genome lacking a country cannot enter a spatial test.
- Genuine evolutionary diversity. If every sequence is 99.9% identical to every other, there is nothing to measure.
- Human health relevance, so that the interpretation section has something to say.
Viruses that satisfy all five include SARS-CoV-2, influenza A, dengue virus, hepatitis B and C viruses, HIV-1, human papillomaviruses, the enteroviruses and the human herpesviruses. A deliberate word of caution applies to SARS-CoV-2: the data are unmatched, but so is the volume of published analysis, and a student project competing in that space must clear a very high bar for novelty. Genome availability is not the scarce resource; an unanswered question is. A less intensively studied virus – a regionally important dengue serotype, a specific enterovirus type, or a neglected hepatitis genotype – will usually yield a more original result for the same effort.
4. Framing the Research Question
A project title such as ‘Genome-wide analysis of virus X’ names a dataset but promises no finding. Reviewers and examiners cannot tell from it what would count as a result, and neither, often, can the student. Compare it with a framed question:
Comparative genomic, evolutionary, structural and selection analysis of virus X reveals conserved functional regions and lineage-specific adaptive signatures.
Genome-wide comparative analysis of virus X isolates reveals evolutionary hotspots, conserved regions and host or geographic adaptation.
Each of these states a mechanism that the analysis will test – conservation, selection, adaptation – and therefore assigns a job to every planned figure before any data are downloaded. This is not cosmetic. A framed question tells you which analyses are essential and which are decoration, and it prevents the commonest structural failure in student projects: a sequence of unrelated analyses with no argument connecting them.
Write the title and the eight planned figure legends before you download a single sequence. If you cannot write the legend for a figure, you do not yet know why you are producing it.
5. Building the Genome Dataset
5.1 Where the sequences come from
- NCBI GenBank and RefSeq. GenBank is the primary archive of submitted sequences; RefSeq provides a single curated reference genome per species or genotype and is the natural coordinate system for annotation (O’Leary et al., 2016; Sayers et al., 2024).
- Virus-specific resources. The Los Alamos HIV Sequence Database, the Bacterial and Viral Bioinformatics Resource Center (BV-BRC) and the influenza resources at NCBI provide curated, pre-aligned and better-annotated subsets than a raw GenBank query.
- Indispensable for influenza and SARS-CoV-2, but access is governed by a data access agreement that requires acknowledgement of the submitting laboratories and prohibits redistribution of the sequences themselves (Shu & McCauley, 2017; Khare et al., 2021). You may share accession lists; you may not share the FASTA file.
- The International Committee on Taxonomy of Viruses publishes the Master Species List, which defines the correct, current name and rank of your virus (Lefkowitz et al., 2018; Simmonds et al., 2017). Use it, because taxonomic names in GenBank records are often years out of date.
5.2 The metadata table is part of the data
Before any analysis, build a single spreadsheet with one row per genome and, at minimum, the following columns: accession number, strain or isolate identifier, collection date, country (and region if available), host species, genome length, completeness status, genotype or lineage if assigned, and the publication or submitting laboratory. This table is not administrative overhead. It is the input to the temporal, geographic and host-adaptation analyses, and its completeness silently determines the sample size of each. A dataset of 1,200 genomes of which only 380 carry a usable collection date is a dataset of 380 genomes for every clock-based analysis, and your methods section must say so.
| Box 2. Ethics and data citation
Viral sequences are derived from clinical samples taken from people. Although deposited sequences are de-identified, three obligations follow. First, never attempt to link sequence data to individuals or to re-identify patients. Second, cite the data: list accession numbers in a supplementary table, and where the terms of a resource require it (as GISAID does), acknowledge the originating and submitting laboratories explicitly. Third, be careful with geographic claims. Reporting that a particular country is the ‘source’ of a variant is an inference with real social consequences, and it is usually confounded by where sequencing capacity happens to exist rather than by where the virus actually circulated. |
6. Quality Control
Every downstream result inherits the errors left in the dataset at this stage, and no later analysis can detect them. Quality control is therefore not a preliminary chore; it is a methodological decision that must be documented and defended.
Table 2. Standard quality-control filters and the specific harm each prevents.
| Filter | Typical threshold | What it prevents |
| Genome completeness | Retain sequences >= 90-95% of reference length | Partial genomes create long alignment gaps that are read as deletions and distort diversity estimates. |
| Ambiguous bases (N) | Exclude sequences with > 1-5% N | Runs of N are treated as missing data and can be mistaken for conservation, because an unknown base cannot register as a difference. |
| Exact and near duplicates | Collapse identical sequences; record the count | Duplicates inflate apparent support for a clade and violate the independence assumption of most statistical tests. |
| Length outliers | Exclude sequences outside a defined length window | Homologous sequences must be of comparable length to align meaningfully. |
| Annotation quality | Require defined coding sequences in the correct frame | Frame errors invalidate every codon-based analysis, including all of dN/dS. |
| Contamination and misassembly | Screen with BLAST against a broad database | Non-viral or wrong-species sequence produces spurious long branches. |
| Metadata completeness | Flag missing date, country or host | Silently reduces the effective sample size of temporal and spatial analyses. |
Two rules make quality control defensible. First, apply the filters as a script, not by hand, so that the exact retained set can be regenerated. Second, report the attrition: a short flow statement such as ‘3,412 genomes retrieved; 402 removed for incompleteness; 118 for excessive ambiguity; 265 exact duplicates collapsed; 2,627 retained’ belongs in the methods section of every report.
7. Genome Annotation and Architecture
Annotation establishes the coordinate system in which every later result will be expressed. For each genome, or at least for the reference and one representative of each lineage, determine the genome length and overall architecture, the open reading frames, the coding and non-coding regions, any overlapping genes, the untranslated regions and regulatory elements, and the division between structural and non-structural proteins.
Overlapping reading frames deserve particular attention, because they are common in compact viral genomes – the hepatitis B virus genome encodes four overlapping frames in 3.2 kb – and they constrain evolution in a way that will distort your selection analysis if ignored. A nucleotide change that is synonymous in one frame is usually non-synonymous in the other, so a codon-based model applied to one frame alone will misattribute the resulting pattern to selection on that frame’s protein.
Produce a genome map figure early. It orients the reader, and it becomes the coordinate key for the diversity plot, the selection plot and every subsequent positional result.
8. Multiple Sequence Alignment
An alignment is a hypothesis of positional homology: the claim that the residues in a given column descend from a single ancestral residue. Every downstream method – diversity, phylogeny, selection – takes that claim as given. If the alignment is wrong, everything after it is wrong, and confidently so.
8.1 Choosing and running an aligner
MAFFT is the standard choice for viral datasets, offering a range of algorithms that trade accuracy against speed and scaling comfortably to thousands of sequences (Katoh & Standley, 2013; Katoh et al., 2019). MUSCLE remains a reasonable alternative (Edgar, 2004). For viral genomes specifically, Nextclade aligns each sequence against a curated reference, calls mutations, assigns clades and reports per-sequence quality scores in a single pass, which makes it an efficient combined alignment and quality-control step (Aksamentov et al., 2021).
- For coding regions, align at the amino-acid level and then back-translate to codons. Aligning coding nucleotides directly frequently introduces gaps that are not multiples of three, which destroys the reading frame and makes codon-based selection analysis impossible.
- Inspect the alignment visually – in AliView, Jalview or MEGA – before using it. Look at the ends, where partial sequences produce ragged edges, and at any column with an implausible concentration of gaps.
- Trim with judgement. Tools such as trimAl remove poorly aligned columns automatically (Capella-Gutierrez et al., 2009), and removing ragged termini is almost always correct. Aggressive trimming of internal regions, however, discards exactly the variable sites that carry your signal (Talavera & Castresana, 2007). Report what you trimmed and why.
9. Measuring Genome-Wide Variation
With a trustworthy alignment you can describe how variation is distributed along the genome. This is the first substantive result of the study and it generates the hypotheses that later sections test.
9.1 The core statistics
- Nucleotide diversity, pi, is the average number of nucleotide differences per site between two randomly chosen sequences (Nei & Li, 1979). It is the standard summary of standing variation in a population.
- Shannon entropy at a site measures the uncertainty of the base at that position (Shannon, 1948). Entropy is zero at a fully conserved site and maximal when all four bases are equally frequent, which makes it an intuitive per-site conservation score.
- Segregating sites, S, is simply the count of variable positions, and Watterson’s theta is derived from it. The comparison of pi with theta underlies Tajima’s D (Tajima, 1989).
- Mutation frequency per position, calculated relative to the reference, identifies the specific substitutions that recur across independent genomes – often the most immediately interpretable result in the whole study.
9.2 The sliding-window plot
Compute pi (and, separately, mutation frequency) in a window of fixed width – commonly 100 to 500 nucleotides – stepped along the genome in increments of, say, 50 nucleotides, and plot the result against genome position with the gene map aligned beneath it. This single figure usually becomes the most cited panel of a viral genomics paper, because it shows at a glance which genes are constrained and which are hypervariable.
Two cautions. First, the window width is a smoothing parameter: a wide window hides narrow peaks, a narrow window produces noise. Report the width and step, and check that your conclusions survive a change of window size. Second, a peak of diversity is not evidence of positive selection. It may equally reflect relaxed constraint, a hypermutable context, recombination, or alignment error. Distinguishing these possibilities is exactly what Section 12 is for. DnaSP (Rozas et al., 2017) computes all of these statistics directly, and the calculations are also straightforward to implement in Biopython or R.
10. Phylogenetic Analysis
10.1 Building the tree
The standard modern pipeline is: aligned sequences, then substitution-model selection, then maximum-likelihood tree inference, then branch support. ModelFinder selects the best-fitting substitution model by information criterion (Kalyaanamoorthy et al., 2017); IQ-TREE 2 performs the tree search (Nguyen et al., 2015; Minh et al., 2020); and ultrafast bootstrap approximation provides branch support with far less computation than the standard bootstrap (Felsenstein, 1985; Hoang et al., 2018). RAxML-NG (Kozlov et al., 2019) and FastTree 2 (Price et al., 2010) are alternatives, the latter for very large datasets where speed matters more than precision. MEGA 11 offers the same analyses through a graphical interface and is a reasonable starting point for students new to the field (Tamura et al., 2021).
- Interpret ultrafast bootstrap values with the correct scale: >= 95 indicates strong support, which is not the same threshold used for standard bootstrap values (>= 70). Do not mix the conventions.
- Root the tree explicitly and state how. An outgroup from a related virus species is the most defensible method; midpoint rooting assumes a roughly constant rate; and for a time-scaled tree the root is estimated from the collection dates themselves.
- A clade is a hypothesis of common ancestry, not a category. Naming clusters A, B and C is description; explaining why they differ is analysis.
10.2 What a tree can and cannot tell you
A phylogeny shows the branching pattern of inferred ancestry and the relative amount of change along each branch. It does not, by itself, explain anything. Geographic clustering in a tree may reflect genuine local transmission, or it may reflect the fact that one laboratory sequenced 400 samples from one hospital in one month. Temporal clustering may reflect a real expansion, or a change in national sequencing policy. The tree is essential, and it should never be the only analysis in the report.
11. Time-Scaled and Phylodynamic Analysis
Because viral sequences carry collection dates, the tips of a viral phylogeny are sampled at different times – the sequences are heterochronous. This allows branch lengths to be calibrated in units of real time rather than substitutions, converting a topology into a dated history (Volz et al., 2013).
- Test for temporal signal first. Use TempEst to regress root-to-tip genetic distance against collection date (Rambaut et al., 2016). A positive slope with a reasonable correlation indicates measurable evolution; a flat or noisy relationship means the dataset cannot support clock analysis, and you must say so rather than proceed anyway.
- Estimate rates and dates. TreeTime provides fast maximum-likelihood dating suitable for large datasets (Sagulenko et al., 2018). BEAST 2 (Bouckaert et al., 2019) and BEAST 1.10 (Suchard et al., 2018) implement full Bayesian inference, jointly estimating the tree, the clock model and the demographic model, at much greater computational cost.
- Report what the analysis yields: the substitution rate with its credible interval, the time to the most recent common ancestor (TMRCA) of the whole sample and of each major lineage, and where appropriate a reconstruction of effective population size through time.
Interpretation requires care on two points. The TMRCA of your sample is the age of the common ancestor of the sequences you happened to collect, which is a lower bound on the age of the virus, not its origin date. And Bayesian phylogeographic reconstruction (Lemey et al., 2009) is strongly affected by uneven sampling: countries that sequence heavily are inferred as sources far more often than they should be, a bias that structured-coalescent approaches partially address (De Maio et al., 2015).
12. Detecting Natural Selection
This is the analysis that most clearly separates a descriptive study from an explanatory one, and it is the section most often missing from undergraduate reports.
12.1 The logic of dN/dS
In a protein-coding sequence, some nucleotide changes alter the encoded amino acid (non-synonymous) and some do not (synonymous). Synonymous changes are, to a first approximation, invisible to selection acting on protein function, so their rate provides an internal estimate of the neutral substitution rate. The ratio omega = dN/dS of the non-synonymous rate per non-synonymous site to the synonymous rate per synonymous site therefore has a direct interpretation (Nei & Gojobori, 1986; Yang & Bielawski, 2000):
- omega < 1: purifying (negative) selection. Amino-acid changes are being removed. This is the normal state of most viral protein-coding sites, and it is especially strong in polymerases and other replication machinery.
- omega = 1: neutral evolution, or no detectable selection.
- omega > 1: positive (diversifying) selection. Amino-acid changes are being favoured – the signature expected at sites involved in immune escape, receptor interaction or host adaptation.
12.2 Site-level methods
Genome-wide or gene-wide averages of omega are almost always below 1, because most sites in most proteins are constrained; averaging therefore hides the handful of sites that matter. Site-level methods, implemented principally in the HyPhy package (Kosakovsky Pond et al., 2020) and available through the Datamonkey web server (Weaver et al., 2018), test each codon individually:
Table 3. Site-level selection methods in HyPhy and when to use each.
| Method | Statistical approach | Best suited to |
| SLAC | Counting-based, ancestral-state reconstruction | Large alignments; fast and conservative; a useful first pass (Kosakovsky Pond & Frost, 2005). |
| FEL | Fixed-effects maximum likelihood, site by site | Moderate datasets; detects pervasive selection acting across the whole tree (Kosakovsky Pond & Frost, 2005). |
| FUBAR | Bayesian, with information shared across sites | Large datasets; high power for pervasive selection; fast (Murrell et al., 2013). |
| MEME | Mixed-effects model allowing omega to vary across branches | Episodic selection affecting a site in only part of the tree – the realistic case for immune escape (Murrell et al., 2012). |
| aBSREL / RELAX | Branch-site models | Testing whether particular lineages experienced intensified or relaxed selection (Smith et al., 2015; Wertheim et al., 2015). |
Good practice is to run more than one method and to treat sites recovered by several methods as the robust set. PAML implements a complementary family of codon models and remains a standard reference implementation (Yang, 2007).
12.3 Limitations you must state
- dN/dS was derived for divergent sequences between species. Applied within a population, where segregating polymorphism has not yet been filtered by selection, its behaviour is different and its power is reduced (Kryazhimskiy & Plotkin, 2008). Viral datasets are population samples, so this caveat always applies.
- Undetected recombination inflates false-positive rates substantially. Screen first (Section 2.3).
- Frame errors or misalignment in coding regions generate spurious positive selection. A site under ‘positive selection’ at the ragged end of an alignment is almost always an artefact.
- Sites are tested in their thousands, so multiple-testing correction is mandatory; report the false discovery rate procedure and threshold used (Benjamini & Hochberg, 1995).
- A positively selected site is a statistical signal, not a demonstrated function.
13. Protein Structure and Mutation Mapping
Selection analysis tells you which residues are changing unusually. Structural analysis tells you where those residues are, and that spatial context is what converts a list of codon positions into a biological argument.
- Obtain a structure. Query the Protein Data Bank first for an experimentally determined structure of your protein or a close homologue (Berman et al., 2000). If none exists, use the AlphaFold Protein Structure Database (Jumper et al., 2021; Varadi et al., 2022) or generate a prediction with ColabFold (Mirdita et al., 2022).
- Check the confidence. AlphaFold reports a per-residue confidence score, pLDDT. Regions below about 70 are unreliable and regions below 50 frequently correspond to intrinsically disordered segments. Never draw structural conclusions from a low-confidence region, and always show the confidence colouring in a supplementary figure.
- Map your results onto the structure. Colour conserved residues, positively selected sites, high-frequency mutations and domain boundaries, using PyMOL (Schrodinger, 2015) or ChimeraX (Pettersen et al., 2021; Meng et al., 2023).
- Ask the structural question. Are the evolutionary hotspots clustered, or scattered? Are they surface-exposed or buried? Do they lie in or near a known functional site – a receptor-binding surface, an active site, an antibody epitope?
A cluster of positively selected, surface-exposed residues in a receptor-binding domain is a coherent and interesting result. Scattered positively selected residues buried in a hydrophobic core, by contrast, should make you suspect alignment error rather than biology – and saying so demonstrates better scientific judgement than reporting it as a finding.
14. Functional Annotation and Host-Virus Interactions
14.1 Assigning function to sequence
Functional annotation places your variable and conserved regions into the vocabulary of molecular function. InterPro integrates multiple signature databases to identify domains and families (Paysan-Lafosse et al., 2023); Pfam provides the underlying profile hidden Markov models (Mistry et al., 2021); UniProt supplies curated, literature-backed annotation of individual proteins and residues (UniProt Consortium, 2023); and Gene Ontology and KEGG provide controlled vocabularies for function and pathway membership (Ashburner et al., 2000; Kanehisa & Goto, 2000).
14.2 Host-virus interaction analysis
The interaction layer links viral proteins to host proteins and thence to host pathways: receptor usage, interferon and innate immune signalling, apoptosis, inflammatory cascades and intracellular trafficking. Use experimentally supported interaction data wherever possible – IntAct (Orchard et al., 2014), BioGRID (Oughtred et al., 2021) and the virus-specific VirHostNet (Guirimand et al., 2015) – rather than relying on prediction alone. STRING is useful for network context but mixes experimental evidence with predicted and text-mined associations, and the evidence channels must be filtered and reported (Szklarczyk et al., 2023). The SARS-CoV-2 interactome of Gordon et al. (2020) is a good model of how such data are generated and presented.
| Box 3. Interpreting an enrichment result
Pathway enrichment analysis answers a narrow question: given a list of proteins, are any annotation terms more frequent in that list than expected by chance from a defined background set? Three requirements follow. Define the background explicitly (all human proteins is rarely the right choice). Correct for multiple testing across all terms tested. And remember that annotation databases are biased towards well-studied proteins, so ‘enrichment’ partly measures where research funding has gone. |
15. Geographic and Temporal Structure
15.1 Geographic analysis
With country metadata you can ask whether lineages are geographically structured, whether particular variants are associated with particular regions, and whether mutation frequencies differ between regions. The essential requirement is that these questions be answered with statistics rather than with a coloured map. Appropriate tools include F-statistics for population differentiation (Weir & Cockerham, 1984), the analysis of molecular variance for partitioning variation among and within groups (Excoffier et al., 1992), and chi-squared or Fisher’s exact tests for variant-by-region contingency tables, with correction for multiple comparisons.
Sampling bias is the dominant confounder and must be addressed explicitly. Report the number of genomes per country and per year; consider subsampling to equalise representation; and be conservative in language, because ‘variant V is significantly more frequent among sequences from country C’ is a statement about a sequence database, not necessarily about a population.
15.2 Temporal analysis
Divide the genomes into time windows chosen to reflect the sampling history of the virus – for example 2015-2017, 2018-2020, 2021-2023 and 2024-2026 – and compare across them: mutation frequency, nucleotide diversity, clade composition, and the strength and location of selection. A finding that a genome region under strong purifying selection in early windows shows relaxed constraint in later ones is a genuinely interesting result, and it is only visible if the comparison is made.
Choose the window boundaries from the data – sampling density, known epidemiological events – and not for convenience, and report the number of genomes in each window beside every result derived from it.
16. Statistical Analysis
Computational biology generates enormous numbers of comparisons, and the discipline of the analysis rests on how those comparisons are handled.
- Descriptive population-genetic statistics: nucleotide diversity, segregating sites, haplotype diversity.
- Neutrality tests: Tajima’s D compares pi with Watterson’s theta and is sensitive to departures from neutral, constant-size evolution (Tajima, 1989). Because both selection and demographic change (a population bottleneck, an epidemic expansion) shift D in the same direction, a significant value is not by itself evidence of selection.
- Differentiation and structure: F-statistics (Weir & Cockerham, 1984) and AMOVA (Excoffier et al., 1992).
- Association and trend: correlation and regression, for example of root-to-tip distance against date, or of mutation frequency against time.
- Resampling: bootstrap and permutation tests, which make few distributional assumptions and are well suited to sequence data.
- Multiple-testing correction: the Benjamini-Hochberg false discovery rate procedure is the usual choice for site-level and gene-level testing (Benjamini & Hochberg, 1995).
One statistical issue is specific to this field and is routinely overlooked. Sequences related by a phylogeny are not independent observations: two genomes from the same recent outbreak are similar because they share an ancestor, not because they independently support a hypothesis. Standard tests that assume independent samples will therefore report p-values that are far too small. Phylogenetically aware methods, or at minimum subsampling to reduce clustering, are required (Felsenstein, 1985).
17. Machine Learning: Optional but Unforgiving
Machine learning can add real value to a viral genomics study by asking whether genomic features alone are sufficient to distinguish lineages, geographic groups or time periods, and by identifying which features carry that information. Typical input features are mutation profiles, k-mer frequencies, amino-acid substitutions, and positional or lineage encodings; common algorithms are random forests (Breiman, 2001), gradient boosting (Chen & Guestrin, 2016) and support vector machines (Cortes & Vapnik, 1995), with principal component analysis and UMAP (McInnes et al., 2018) for unsupervised visualisation. Scikit-learn provides all of these (Pedregosa et al., 2011).
The discipline required is greater than in most other applications, for one reason: sequence data leak. If two nearly identical genomes from the same transmission chain fall one into the training set and one into the test set, the model can achieve near-perfect accuracy by memorisation, and the reported performance is meaningless.
- Split the data by phylogenetic cluster, outbreak, or time period – never at random across individual sequences.
- Perform every preprocessing step (feature selection, scaling, imputation) inside the cross-validation loop, fitted on training data only.
- Report a baseline. If 80% of your sequences belong to one lineage, a classifier that always predicts that lineage is 80% accurate; your model must be compared against that.
- Report class-balanced metrics – balanced accuracy, per-class recall, or the area under the precision-recall curve – not accuracy alone.
- Treat feature-importance scores as descriptive. They indicate which features the model used, not which sites are biologically causal.
18. Integration: Turning Layers into an Argument
The individual analyses above are the evidence; the argument is what you build from them. The integration chain runs:
genome variation -> protein mutation -> selection pressure -> position in the protein structure -> functional domain -> proposed biological significance
A worked example of the reasoning: a site is identified as a mutation-frequency hotspot in the sliding-window analysis; FEL and MEME independently report it as positively selected; mapping shows it lies within an annotated receptor-binding domain; the structure shows it is surface-exposed; and the substitution is more frequent in genomes collected after 2021 than before. Each observation on its own is weak. Together they constitute a specific, coherent and falsifiable hypothesis: that this residue is under recent adaptive pressure related to host or immune interaction.
State it as exactly that – a hypothesis, with the experiment that would test it named in the discussion. Every link in the chain is an inference, and the honest reporting of that is a mark of competence, not of weak results.
19. Reproducibility
A computational study that cannot be repeated is not a result; it is an anecdote. Reproducibility in this field is achievable in full, because every input is a file and every step is a command (Sandve et al., 2013; Wilkinson et al., 2016).
- Preserve the audit trail: raw data, accession lists, quality-control scripts, the filtered dataset, alignments, analysis scripts, statistical output, figures and supplementary tables.
- Script everything. A workflow manager such as Snakemake (Molder et al., 2021) or Nextflow (Di Tommaso et al., 2017) records the dependency structure of the analysis and re-runs it end to end.
- Pin the environment. Record exact software versions and parameters, and capture them with Conda or Bioconda (Gruning et al., 2018) or a container (Merkel, 2014; Kurtzer et al., 2017). ‘MAFFT’ is not a method; ‘MAFFT v7.520, –auto’ is.
- Publish the code in a public Git repository, and archive a release with a persistent identifier (for example through Zenodo) so that the version used in the paper remains citable.
- Provide accession numbers, not descriptions. ‘Sequences were downloaded from GenBank’ is not reproducible; a supplementary table of 2,627 accessions is.
- Document enough that a stranger can regenerate every figure. A README specifying the dependencies, the run order and the expected outputs is the minimum.
Note that reproducibility pays a dividend beyond good practice: a pipeline built this way can be pointed at a second virus in days rather than months, which turns one project into a research programme.
20. Writing It Up
Plan the figures before running the analyses. A conventional and effective structure for a study of this kind uses eight main figures:
- Genome organisation and dataset composition – the coordinate system and the sample.
- Genome-wide nucleotide diversity and mutation frequency (the sliding-window plot).
- Phylogenetic tree with support values, showing major clades.
- Time-scaled or geographically annotated phylogeny.
- Selection landscape: dN/dS by site, with positively selected sites marked.
- Protein structure with conserved and selected residues mapped.
- Integrated genotype-structure-function panel.
- Proposed evolutionary model – the synthesis the paper argues for.
Figures 1 to 6 present evidence, Figure 7 integrates it and Figure 8 interprets it. If you cannot draw Figure 8, the study has not yet reached a conclusion, and the remedy is more thinking rather than more analysis.
In the methods section, give versions, parameters and thresholds for every tool. In the results, separate what was observed from what it might mean. In the discussion, state the limitations explicitly – sampling bias, the assumptions of the clock, the caveats of dN/dS, the absence of experimental validation – because a reviewer who finds a limitation you did not mention will doubt everything else you wrote.
21. Ten Common Pitfalls
- Choosing a saturated study system. Competing with thousands of SARS-CoV-2 papers is a high-risk strategy for a first project.
- Presenting a phylogenetic tree as a conclusion. A topology describes structure; it does not explain it.
- Skipping quality control, or performing it by hand so that it cannot be reproduced.
- Building a codon-based analysis on an alignment that was never checked for frame integrity.
- Omitting the recombination screen before phylogenetic and selection analysis.
- Omitting selection analysis altogether, leaving the study descriptive.
- Ignoring sampling bias in geographic and temporal comparisons.
- Reporting thousands of site-level tests without multiple-testing correction.
- Allowing data leakage in a machine-learning component, then reporting the inflated accuracy.
- Stating computational predictions as demonstrated biological facts.
22. A Suggested Practical Exercise
The following exercise can be completed in roughly two weeks of part-time work on a standard laptop, and it exercises every stage of the workflow at a manageable scale.
- Choose a single well-sampled virus and one of its protein-coding genes – for example the dengue virus envelope (E) gene, or the VP1 gene of a specific enterovirus type.
- Retrieve 200 to 500 complete coding sequences from GenBank with collection date and country, and build the metadata table.
- Apply and document quality-control filters; report the attrition at each step.
- Align at the amino-acid level with MAFFT and back-translate to codons; inspect and trim the alignment.
- Screen for recombination with GARD or RDP5 and act on the result.
- Compute nucleotide diversity and Shannon entropy in sliding windows; plot against gene position.
- Select a substitution model with ModelFinder and infer a maximum-likelihood tree with IQ-TREE 2 and ultrafast bootstrap.
- Test for temporal signal with TempEst; if present, date the tree with TreeTime.
- Run FEL, FUBAR and MEME on Datamonkey; take the intersection as the robust set of positively selected sites.
- Retrieve or predict the protein structure, map the conserved and selected sites, and produce the integrated figure.
- Write a 2,500-word report with the eight-figure structure, a full methods section including versions and parameters, and a public repository containing the scripts and the accession list.
Glossary
Terms are listed alphabetically. Each entry gives the working definition used in this chapter; fuller treatments are available in the cited literature.
| Alignment (multiple sequence) | A hypothesis that residues placed in the same column of a set of sequences descend from a common ancestral residue. |
| AMOVA | Analysis of molecular variance; partitions genetic variation among and within predefined groups such as countries or time periods. |
| Clade | A group on a phylogeny consisting of an ancestor and all of its descendants. |
| Consensus sequence | The most frequent base at each position across a viral population within a host; what is usually deposited in public databases. |
| dN/dS (omega) | The ratio of the non-synonymous substitution rate per non-synonymous site to the synonymous substitution rate per synonymous site; values above 1 indicate positive selection, below 1 purifying selection. |
| Heterochronous sampling | Sequences collected at different, recorded points in time, which allows branch lengths to be calibrated in real time units. |
| Nucleotide diversity (pi) | The mean number of nucleotide differences per site between two randomly chosen sequences from a sample. |
| Open reading frame (ORF) | A stretch of sequence between a start and a stop codon capable of encoding a protein. |
| pLDDT | The per-residue confidence score reported by AlphaFold; values below about 70 indicate unreliable local structure. |
| Positive (diversifying) selection | Selection favouring amino-acid change, typically at sites involved in immune escape or host adaptation. |
| Purifying (negative) selection | Selection removing amino-acid change; the usual state of functionally constrained sites. |
| Quasispecies | The cloud of closely related genetic variants that constitutes an RNA virus population within a single host. |
| Reassortment | Exchange of whole genome segments between co-infecting segmented viruses, as in influenza A. |
| Recombination | Exchange of genetic material producing a genome whose regions have different evolutionary histories; violates the single-tree assumption of phylogenetics. |
| Shannon entropy | An information-theoretic measure of variability at an alignment position; zero when the position is invariant. |
| Tajima’s D | A neutrality test comparing two estimators of genetic diversity; affected by both selection and demographic change. |
| TMRCA | Time to the most recent common ancestor of a set of sampled sequences. |
| Temporal signal | A detectable positive relationship between a sequence’s collection date and its genetic distance from the root; a prerequisite for molecular clock analysis. |
Self-Assessment Questions
These questions test understanding rather than recall. Several have no single correct answer; the reasoning matters more than the conclusion.
- Explain why a study of human papillomavirus using sequences collected between 2010 and 2025 is unlikely to support a time-scaled phylogenetic analysis, while a comparable dengue virus dataset would.
- A classmate reports that 12% of sites in a viral gene show dN/dS greater than 1. List four possible technical explanations that must be excluded before this is interpreted as positive selection.
- Why must coding sequences be aligned as amino acids and back-translated, rather than aligned directly as nucleotides?
- A sliding-window plot shows a sharp diversity peak in one gene. Give three biological and two technical explanations for such a peak, and describe an analysis that would help distinguish them.
- A machine-learning classifier distinguishes sequences from Country A and Country B with 97% accuracy. Describe two ways this result could be an artefact, and the design changes that would guard against each.
- Explain, in terms a first-year student would understand, why two genomes from the same outbreak are not two independent pieces of evidence.
- Your dataset contains 1,500 genomes but only 620 have collection dates. State precisely which analyses in this chapter are affected and how you would report the discrepancy.
- Distinguish between a positively selected site, a frequently mutated site and a functionally important site, and explain how a study can have evidence for the first two without having evidence for the third.
- Why is Tajima’s D not, on its own, evidence of natural selection?
- Describe the minimum set of materials you would deposit publicly so that another student could reproduce Figure 2 of your report exactly
- .PPT
https://docs.google.com/presentation/d/1-jx_Sik74XL39dY_GIK3csr6MQ5pqicp/edit?usp=drive_link&ouid=102247424304535240260&rtpof=true&sd=true
References
References follow APA (7th edition) style. Students should consult the original publication before citing any of these works in their own writing, and should check for a more recent database update paper, since resources such as GenBank, UniProt, InterPro and STRING publish a new reference article most years.
Aksamentov, I., Roemer, C., Hodcroft, E. B., & Neher, R. A. (2021). Nextclade: clade assignment, mutation calling and quality control for viral genomes. Journal of Open Source Software, 6(67), 3773.
Ashburner, M., Ball, C. A., Blake, J. A., Botstein, D., Butler, H., Cherry, J. M., … Sherlock, G. (2000). Gene Ontology: tool for the unification of biology. Nature Genetics, 25(1), 25-29.
Baltimore, D. (1971). Expression of animal virus genomes. Bacteriological Reviews, 35(3), 235-241.
Benjamini, Y., & Hochberg, Y. (1995). Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society, Series B, 57(1), 289-300.
Berman, H. M., Westbrook, J., Feng, Z., Gilliland, G., Bhat, T. N., Weissig, H., … Bourne, P. E. (2000). The Protein Data Bank. Nucleic Acids Research, 28(1), 235-242.
Boni, M. F., Posada, D., & Feldman, M. W. (2007). An exact nonparametric method for inferring mosaic structure in sequence triplets. Genetics, 176(2), 1035-1047.
Bouckaert, R., Vaughan, T. G., Barido-Sottani, J., Duchene, S., Fourment, M., Gavryushkina, A., … Drummond, A. J. (2019). BEAST 2.5: An advanced software platform for Bayesian evolutionary analysis. PLOS Computational Biology, 15(4), e1006650.
Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5-32.
Capella-Gutierrez, S., Silla-Martinez, J. M., & Gabaldon, T. (2009). trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses. Bioinformatics, 25(15), 1972-1973.
Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785-794).
Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273-297.
De Maio, N., Wu, C.-H., O’Reilly, K. M., & Wilson, D. (2015). New routes to phylogeography: a Bayesian structured coalescent approximation. PLOS Genetics, 11(8), e1005421.
Di Tommaso, P., Chatzou, M., Floden, E. W., Barja, P. P., Palumbo, E., & Notredame, C. (2017). Nextflow enables reproducible computational workflows. Nature Biotechnology, 35(4), 316-319.
Domingo, E., Sheldon, J., & Perales, C. (2012). Viral quasispecies evolution. Microbiology and Molecular Biology Reviews, 76(2), 159-216.
Drake, J. W., & Holland, J. J. (1999). Mutation rates among RNA viruses. Proceedings of the National Academy of Sciences, 96(24), 13910-13913.
Duffy, S., Shackelton, L. A., & Holmes, E. C. (2008). Rates of evolutionary change in viruses: patterns and determinants. Nature Reviews Genetics, 9(4), 267-276.
Edgar, R. C. (2004). MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Research, 32(5), 1792-1797.
Excoffier, L., Smouse, P. E., & Quattro, J. M. (1992). Analysis of molecular variance inferred from metric distances among DNA haplotypes. Genetics, 131(2), 479-491.
Felsenstein, J. (1985). Confidence limits on phylogenies: an approach using the bootstrap. Evolution, 39(4), 783-791.
Gordon, D. E., Jang, G. M., Bouhaddou, M., Xu, J., Obernier, K., White, K. M., … Krogan, N. J. (2020). A SARS-CoV-2 protein interaction map reveals targets for drug repurposing. Nature, 583(7816), 459-468.
Gruning, B., Dale, R., Sjodin, A., Chapman, B. A., Rowe, J., Tomkins-Tinch, C. H., … Koster, J. (2018). Bioconda: sustainable and comprehensive software distribution for the life sciences. Nature Methods, 15(7), 475-476.
Guirimand, T., Delmotte, S., & Navratil, V. (2015). VirHostNet 2.0: surfing on the web of virus/host molecular interactions data. Nucleic Acids Research, 43(D1), D583-D587.
Hadfield, J., Megill, C., Bell, S. M., Huddleston, J., Potter, B., Callender, C., … Neher, R. A. (2018). Nextstrain: real-time tracking of pathogen evolution. Bioinformatics, 34(23), 4121-4123.
Hoang, D. T., Chernomor, O., von Haeseler, A., Minh, B. Q., & Vinh, L. S. (2018). UFBoot2: improving the ultrafast bootstrap approximation. Molecular Biology and Evolution, 35(2), 518-522.
Holmes, E. C. (2009). The evolution and emergence of RNA viruses. Oxford University Press.
Huddleston, J., Hadfield, J., Sibley, T. R., Lee, J., Fay, K., Ilcisin, M., … Bedford, T. (2021). Augur: a bioinformatics toolkit for phylogenetic analyses of human pathogens. Journal of Open Source Software, 6(57), 2906.
Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., … Hassabis, D. (2021). Highly accurate protein structure prediction with AlphaFold. Nature, 596(7873), 583-589.
Kalyaanamoorthy, S., Minh, B. Q., Wong, T. K. F., von Haeseler, A., & Jermiin, L. S. (2017). ModelFinder: fast model selection for accurate phylogenetic estimates. Nature Methods, 14(6), 587-589.
Kanehisa, M., & Goto, S. (2000). KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Research, 28(1), 27-30.
Katoh, K., Rozewicki, J., & Yamada, K. D. (2019). MAFFT online service: multiple sequence alignment, interactive sequence choice and visualization. Briefings in Bioinformatics, 20(4), 1160-1166.
Katoh, K., & Standley, D. M. (2013). MAFFT multiple sequence alignment software version 7: improvements in performance and usability. Molecular Biology and Evolution, 30(4), 772-780.
Khare, S., Gurry, C., Freitas, L., Schultz, M. B., Bach, G., Diallo, A., … Maurer-Stroh, S. (2021). GISAID’s role in pandemic response. China CDC Weekly, 3(49), 1049-1051.
Kosakovsky Pond, S. L., & Frost, S. D. W. (2005). Not so different after all: a comparison of methods for detecting amino acid sites under selection. Molecular Biology and Evolution, 22(5), 1208-1222.
Kosakovsky Pond, S. L., Posada, D., Gravenor, M. B., Woelk, C. H., & Frost, S. D. W. (2006). GARD: a genetic algorithm for recombination detection. Bioinformatics, 22(24), 3096-3098.
Kosakovsky Pond, S. L., Poon, A. F. Y., Velazquez, R., Weaver, S., Hepler, N. L., Murrell, B., … Muse, S. V. (2020). HyPhy 2.5 – a customizable platform for evolutionary hypothesis testing using phylogenies. Molecular Biology and Evolution, 37(1), 295-299.
Kozlov, A. M., Darriba, D., Flouri, T., Morel, B., & Stamatakis, A. (2019). RAxML-NG: a fast, scalable and user-friendly tool for maximum likelihood phylogenetic inference. Bioinformatics, 35(21), 4453-4455.
Kryazhimskiy, S., & Plotkin, J. B. (2008). The population genetics of dN/dS. PLOS Genetics, 4(12), e1000304.
Kurtzer, G. M., Sochat, V., & Bauer, M. W. (2017). Singularity: Scientific containers for mobility of compute. PLOS ONE, 12(5), e0177459.
Lefkowitz, E. J., Dempsey, D. M., Hendrickson, R. C., Orton, R. J., Siddell, S. G., & Smith, D. B. (2018). Virus taxonomy: the database of the International Committee on Taxonomy of Viruses (ICTV). Nucleic Acids Research, 46(D1), D708-D717.
Lemey, P., Rambaut, A., Drummond, A. J., & Suchard, M. A. (2009). Bayesian phylogeography finds its roots. PLOS Computational Biology, 5(9), e1000520.
Martin, D. P., Varsani, A., Roumagnac, P., Botha, G., Maslamoney, S., Schwab, T., … Murrell, B. (2021). RDP5: a computer program for analysing recombination in, and removing signals of recombination from, nucleotide sequence datasets. Virus Evolution, 7(1), veaa087.
McInnes, L., Healy, J., & Melville, J. (2018). UMAP: Uniform Manifold Approximation and Projection for dimension reduction. arXiv preprint arXiv:1802.03426.
Meng, E. C., Goddard, T. D., Pettersen, E. F., Couch, G. S., Pearson, Z. J., Morris, J. H., & Ferrin, T. E. (2023). UCSF ChimeraX: Tools for structure building and analysis. Protein Science, 32(11), e4792.
Merkel, D. (2014). Docker: lightweight Linux containers for consistent development and deployment. Linux Journal, 2014(239), 2.
Minh, B. Q., Schmidt, H. A., Chernomor, O., Schrempf, D., Woodhams, M. D., von Haeseler, A., & Lanfear, R. (2020). IQ-TREE 2: New models and efficient methods for phylogenetic inference in the genomic era. Molecular Biology and Evolution, 37(5), 1530-1534.
Mirdita, M., Schutze, K., Moriwaki, Y., Heo, L., Ovchinnikov, S., & Steinegger, M. (2022). ColabFold: making protein folding accessible to all. Nature Methods, 19(6), 679-682.
Mistry, J., Chuguransky, S., Williams, L., Qureshi, M., Salazar, G. A., Sonnhammer, E. L. L., … Bateman, A. (2021). Pfam: The protein families database in 2021. Nucleic Acids Research, 49(D1), D412-D419.
Molder, F., Jablonski, K. P., Letcher, B., Hall, M. B., Tomkins-Tinch, C. H., Sochat, V., … Koster, J. (2021). Sustainable data analysis with Snakemake. F1000Research, 10, 33.
Murrell, B., Moola, S., Mabona, A., Weighill, T., Sheward, D., Kosakovsky Pond, S. L., & Scheffler, K. (2013). FUBAR: a fast, unconstrained Bayesian approximation for inferring selection. Molecular Biology and Evolution, 30(5), 1196-1205.
Murrell, B., Wertheim, J. O., Moola, S., Weighill, T., Scheffler, K., & Kosakovsky Pond, S. L. (2012). Detecting individual sites subject to episodic diversifying selection. PLOS Genetics, 8(7), e1002764.
Nei, M., & Gojobori, T. (1986). Simple methods for estimating the numbers of synonymous and nonsynonymous nucleotide substitutions. Molecular Biology and Evolution, 3(5), 418-426.
Nei, M., & Li, W.-H. (1979). Mathematical model for studying genetic variation in terms of restriction endonucleases. Proceedings of the National Academy of Sciences, 76(10), 5269-5273.
Nguyen, L.-T., Schmidt, H. A., von Haeseler, A., & Minh, B. Q. (2015). IQ-TREE: A fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies. Molecular Biology and Evolution, 32(1), 268-274.
O’Leary, N. A., Wright, M. W., Brister, J. R., Ciufo, S., Haddad, D., McVeigh, R., … Pruitt, K. D. (2016). Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Research, 44(D1), D733-D745.
Orchard, S., Ammari, M., Aranda, B., Breuza, L., Briganti, L., Broackes-Carter, F., … Hermjakob, H. (2014). The MIntAct project – IntAct as a common curation platform for 11 molecular interaction databases. Nucleic Acids Research, 42(D1), D358-D363.
Oughtred, R., Rust, J., Chang, C., Breitkreutz, B.-J., Stark, C., Willems, A., … Tyers, M. (2021). The BioGRID database: A comprehensive biomedical resource of curated protein, genetic, and chemical interactions. Protein Science, 30(1), 187-200.
Paysan-Lafosse, T., Blum, M., Chuguransky, S., Grego, T., Pinto, B. L., Salazar, G. A., … Bateman, A. (2023). InterPro in 2022. Nucleic Acids Research, 51(D1), D418-D427.
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., … Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825-2830.
Pettersen, E. F., Goddard, T. D., Huang, C. C., Meng, E. C., Couch, G. S., Croll, T. I., … Ferrin, T. E. (2021). UCSF ChimeraX: Structure visualization for researchers, educators, and developers. Protein Science, 30(1), 70-82.
Posada, D., & Crandall, K. A. (2001). Evaluation of methods for detecting recombination from DNA sequences: computer simulations. Proceedings of the National Academy of Sciences, 98(24), 13757-13762.
Price, M. N., Dehal, P. S., & Arkin, A. P. (2010). FastTree 2 – approximately maximum-likelihood trees for large alignments. PLOS ONE, 5(3), e9490.
Rambaut, A., Lam, T. T., Max Carvalho, L., & Pybus, O. G. (2016). Exploring the temporal structure of heterochronous sequences using TempEst (formerly Path-O-Gen). Virus Evolution, 2(1), vew007.
Rozas, J., Ferrer-Mata, A., Sanchez-DelBarrio, J. C., Guirao-Rico, S., Librado, P., Ramos-Onsins, S. E., & Sanchez-Gracia, A. (2017). DnaSP 6: DNA sequence polymorphism analysis of large data sets. Molecular Biology and Evolution, 34(12), 3299-3302.
Sagulenko, P., Puller, V., & Neher, R. A. (2018). TreeTime: Maximum-likelihood phylodynamic analysis. Virus Evolution, 4(1), vex042.
Sandve, G. K., Nekrutenko, A., Taylor, J., & Hovig, E. (2013). Ten simple rules for reproducible computational research. PLOS Computational Biology, 9(10), e1003285.
Sanjuan, R., & Domingo-Calap, P. (2016). Mechanisms of viral mutation. Cellular and Molecular Life Sciences, 73(23), 4433-4448.
Sayers, E. W., Cavanaugh, M., Clark, K., Pruitt, K. D., Sherry, S. T., Yankie, L., & Karsch-Mizrachi, I. (2024). GenBank 2024 update. Nucleic Acids Research, 52(D1), D134-D137.
Schrodinger, LLC. (2015). The PyMOL molecular graphics system, version 2.0.
Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27(3), 379-423.
Shu, Y., & McCauley, J. (2017). GISAID: Global initiative on sharing all influenza data – from vision to reality. Eurosurveillance, 22(13), 30494.
Simmonds, P., Adams, M. J., Benko, M., Breitbart, M., Brister, J. R., Carstens, E. B., … Zerbini, F. M. (2017). Virus taxonomy in the age of metagenomics. Nature Reviews Microbiology, 15(3), 161-168.
Smith, M. D., Wertheim, J. O., Weaver, S., Murrell, B., Scheffler, K., & Kosakovsky Pond, S. L. (2015). Less is more: an adaptive branch-site random effects model for efficient detection of episodic diversifying selection. Molecular Biology and Evolution, 32(5), 1342-1353.
Suchard, M. A., Lemey, P., Baele, G., Ayres, D. L., Drummond, A. J., & Rambaut, A. (2018). Bayesian phylogenetic and phylodynamic data integration using BEAST 1.10. Virus Evolution, 4(1), vey016.
Szklarczyk, D., Kirsch, R., Koutrouli, M., Nastou, K., Mehryary, F., Hachilif, R., … von Mering, C. (2023). The STRING database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Research, 51(D1), D638-D646.
Tajima, F. (1989). Statistical method for testing the neutral mutation hypothesis by DNA polymorphism. Genetics, 123(3), 585-595.
Talavera, G., & Castresana, J. (2007). Improvement of phylogenies after removing divergent and ambiguously aligned blocks from protein sequence alignments. Systematic Biology, 56(4), 564-577.
Tamura, K., Stecher, G., & Kumar, S. (2021). MEGA11: Molecular Evolutionary Genetics Analysis version 11. Molecular Biology and Evolution, 38(7), 3022-3027.
UniProt Consortium. (2023). UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Research, 51(D1), D523-D531.
Varadi, M., Anyango, S., Deshpande, M., Nair, S., Natassia, C., Yordanova, G., … Velankar, S. (2022). AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Research, 50(D1), D439-D444.
Volz, E. M., Koelle, K., & Bedford, T. (2013). Viral phylodynamics. PLOS Computational Biology, 9(3), e1002947.
Weaver, S., Shank, S. D., Spielman, S. J., Li, M., Muse, S. V., & Kosakovsky Pond, S. L. (2018). Datamonkey 2.0: A modern web application for characterizing selective and other evolutionary processes. Molecular Biology and Evolution, 35(3), 773-777.
Weir, B. S., & Cockerham, C. C. (1984). Estimating F-statistics for the analysis of population structure. Evolution, 38(6), 1358-1370.
Wertheim, J. O., Murrell, B., Smith, M. D., Kosakovsky Pond, S. L., & Scheffler, K. (2015). RELAX: detecting relaxed selection in a phylogenetic framework. Molecular Biology and Evolution, 32(3), 820-832.
Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., Appleton, G., Axton, M., Baak, A., … Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018.
Yang, Z. (2007). PAML 4: Phylogenetic analysis by maximum likelihood. Molecular Biology and Evolution, 24(8), 1586-1591.
Yang, Z., & Bielawski, J. P. (2000). Statistical methods for detecting molecular adaptation. Trends in Ecology & Evolution, 15(12), 496-503.