Skip to content

Allele Clustering

Quick Facts

Workflow Type Applicable Kingdom Last Known Changes Command-line Compatibility Workflow Level Dockstore
Phylogenetic Construction Bacteria vX.X.X Yes Set-level Allele_Clustering_PHB

Allele_Clustering_PHB

The Allele_Clustering_PHB workflow provides tools for generating distance matrices and inferring distance-based trees from core- or whole-genome allele profiles produced by the TheiaProk workflows.

This workflow supports:

  • Calculation of absolute or normalized allele differences between samples in a set;
  • Tree inference using UPGMA, single-linkage, complete-linkage, neighbor-joining, and minimum-spanning tree algorithms;
  • Generation of trees directly from TheiaProk's Allele Calling hashes, which represent hierarchical groupings based on the provided core- or whole-genome MLST results.

Allele Clustering Workflow Overview

Allele Clustering workflow taking cg- or wgMLST results and generating a dendrogram in Newick format and an allele difference distance matrix.

Inputs

tree_building_algorithm options

The options for the tree_building_algorithm input are as follows:

  • "upgma" - Iteratively merges the two closest clusters, updating inter-cluster distances with the size-weighted average of the merged cluster's distances, producing a clock-like tree. It is fast and simple but assumes a constant evolutionary rate, so it can distort topology when that assumption is violated.
  • "single_linkage" - Merges the two clusters whose closest cross-cluster sample pair distance is the smallest across all cluster pairs, using the minimum of the two merged clusters' distances. This can cause "chaining," where a sample slightly closer to one cluster than another pulls entire groups together, producing long, straggly trees
  • "absolute_linkage" - Merges the two clusters whose farthest cross-cluster sample pair distance is the smallest across all other cluster pairs, using the maximum of the two merged clusters' distances. It tends to produce compact, evenly sized clusters, but can over-split groups that are internally diverse.
  • "neighbor_joining" - A distance-based method that corrects raw distances by the average divergence of each sample before selecting the pair to join, producing an unrooted tree with branch lengths. It handles rate variation well, but it requires all pairwise distances up front and can produce negative branch lengths in noisy data.
  • "minimum_spanning" - An implementation intended to reproduce the behavior of the BioNumerics algorithm, this method builds a minimum spanning network by greedily connecting each unjoined sample to its nearest already-connected neighbor with a single edge. It produces a network rather than a bifurcating tree, making it ideal for visualizing microevolutionary relationships, but the resulting graph is not a phylogenetic tree and cannot be interpreted as one.

distance_algorithm options

Distance metrics are calculated by specifying one of the following options:

  • "absolute_allele_differences" - Counts the raw number of loci where two samples have different alleles, ignoring any loci where either sample has a missing value. This is the simplest and most interpretable metric, but it is sensitive to the number of loci typed; samples with more loci called will naturally accumulate more differences.
  • "normalized_allele_differences" - Divides the count of differing alleles (including missing data) by the number of loci where both samples have a call, expressing distance as a percentage (0-100). This corrects for variable typing completeness across samples, making comparisons fairer, but the percentage scale can obscure the absolute magnitude of differences.
Terra Task Name Variable Type Description Default Value Terra Status
allele_clustering allele_jsons Array[File] The array of input of either core-genome or whole-genome MLST allele calls in JSON format from the Allele Calling module of TheiaProk Required
allele_clustering distance_algorithm String The distance metric algorithm for calculating distances between samples; see above for options Required
allele_clustering tree_building_algorithm String The tree building algorithm for constructing the tree; see above for options Required
allele_clustering tree_name String The name for the output NWK file Required
allele_clustering_task cpu Int Number of CPUs to allocate to the task 2 Optional
allele_clustering_task disk_size Int Amount of storage (in GB) to allocate to the task 100 Optional
allele_clustering_task docker String The Docker container to use for the task us-docker.pkg.dev/general-theiagen/theiagen/allele-clustering:1.0.0 Optional
allele_clustering_task memory Int Amount of memory/RAM (in GB) to allocate to the task 4 Optional
version_capture docker String The Docker container to use for the task us-docker.pkg.dev/general-theiagen/theiagen/alpine-plus-bash:3.20.0 Optional
version_capture timezone String Set the time zone to get an accurate date of analysis (uses UTC by default) Optional

Workflow Tasks

versioning: Version Capture

The versioning task captures the workflow version from the GitHub (code repository) version.

Version Capture Technical details

Links
Task task_versioning.wdl
allele_clustering: PulseNet 2.0 Hash-Based Allele Clustering

The Allele Clustering module is used by PulseNet 2.0 to generate NWK trees for visualization, using the results from the allele_clustering task available in the TheiaProk workflows.

To run this task, a tree building algorithm and distance algorithm must be specified; these options are available in the inputs section of the workflow documentation.

This module adapts the PulseNet 2.0 code for the calculation of distance matrices from allele profiles or sequence alignment and distance-based tree inference for implementation on Terra.bio, and was developed in collaboration with the Association of Public Health Laboratories (APHL) and Centers for Disease Control and Prevention (CDC)'s Enteric Diseases Laboratory Branch (EDLB). We gratefully acknowledge the developers of the allele clustering algorithm and the CDC EDLB team.

Allele Clustering Technical Details

Links
Task task_allele_clustering.wdl
Software Source Code PulseNet 2.0 Trees
Software Documentation PulseNet 2.0 Trees on GitHub

Outputs

Variable Type Description
allele_clustering_distance_matrix File The distance matrix of the resulting tree
allele_clustering_tree File The NWK file containing the resulting tree
allele_clustering_wf_analysis_date String Date of analysis
allele_clustering_wf_version String Version of PHB used for analysis
concatenated_allele_jsons File The file containing all of the input JSONs concatenated, primarily useful for debugging

References

pulsenet2.0-trees