Allele Clustering¶
Quick Facts¶
| Workflow Type | Applicable Kingdom | Last Known Changes | Command-line Compatibility | Workflow Level | Dockstore |
|---|---|---|---|---|---|
| Phylogenetic Construction | Bacteria | vX.X.X | Yes | Set-level | Allele_Clustering_PHB |
Allele_Clustering_PHB¶
The Allele_Clustering_PHB workflow provides tools for generating distance matrices and inferring distance-based trees from core- or whole-genome allele profiles produced by the TheiaProk workflows.
This workflow supports:
- Calculation of absolute or normalized allele differences between samples in a set;
- Tree inference using UPGMA, single-linkage, complete-linkage, neighbor-joining, and minimum-spanning tree algorithms;
- Generation of trees directly from TheiaProk's Allele Calling hashes, which represent hierarchical groupings based on the provided core- or whole-genome MLST results.
Inputs¶
tree_building_algorithm options
The options for the tree_building_algorithm input are as follows:
- "upgma" - Iteratively merges the two closest clusters, updating inter-cluster distances with the size-weighted average of the merged cluster's distances, producing a clock-like tree. It is fast and simple but assumes a constant evolutionary rate, so it can distort topology when that assumption is violated.
- "single_linkage" - Merges the two clusters whose closest cross-cluster sample pair distance is the smallest across all cluster pairs, using the minimum of the two merged clusters' distances. This can cause "chaining," where a sample slightly closer to one cluster than another pulls entire groups together, producing long, straggly trees
- "absolute_linkage" - Merges the two clusters whose farthest cross-cluster sample pair distance is the smallest across all other cluster pairs, using the maximum of the two merged clusters' distances. It tends to produce compact, evenly sized clusters, but can over-split groups that are internally diverse.
- "neighbor_joining" - A distance-based method that corrects raw distances by the average divergence of each sample before selecting the pair to join, producing an unrooted tree with branch lengths. It handles rate variation well, but it requires all pairwise distances up front and can produce negative branch lengths in noisy data.
- "minimum_spanning" - An implementation intended to reproduce the behavior of the BioNumerics algorithm, this method builds a minimum spanning network by greedily connecting each unjoined sample to its nearest already-connected neighbor with a single edge. It produces a network rather than a bifurcating tree, making it ideal for visualizing microevolutionary relationships, but the resulting graph is not a phylogenetic tree and cannot be interpreted as one.
distance_algorithm options
Distance metrics are calculated by specifying one of the following options:
- "absolute_allele_differences" - Counts the raw number of loci where two samples have different alleles, ignoring any loci where either sample has a missing value. This is the simplest and most interpretable metric, but it is sensitive to the number of loci typed; samples with more loci called will naturally accumulate more differences.
- "normalized_allele_differences" - Divides the count of differing alleles (including missing data) by the number of loci where both samples have a call, expressing distance as a percentage (0-100). This corrects for variable typing completeness across samples, making comparisons fairer, but the percentage scale can obscure the absolute magnitude of differences.
| Terra Task Name | Variable | Type | Description | Default Value | Terra Status |
|---|---|---|---|---|---|
| allele_clustering | allele_jsons | Array[File] | The array of input of either core-genome or whole-genome MLST allele calls in JSON format from the Allele Calling module of TheiaProk | Required | |
| allele_clustering | distance_algorithm | String | The distance metric algorithm for calculating distances between samples; see above for options | Required | |
| allele_clustering | tree_building_algorithm | String | The tree building algorithm for constructing the tree; see above for options | Required | |
| allele_clustering | tree_name | String | The name for the output NWK file | Required | |
| allele_clustering_task | cpu | Int | Number of CPUs to allocate to the task | 2 | Optional |
| allele_clustering_task | disk_size | Int | Amount of storage (in GB) to allocate to the task | 100 | Optional |
| allele_clustering_task | docker | String | The Docker container to use for the task | us-docker.pkg.dev/general-theiagen/theiagen/allele-clustering:1.0.0 | Optional |
| allele_clustering_task | memory | Int | Amount of memory/RAM (in GB) to allocate to the task | 4 | Optional |
| version_capture | docker | String | The Docker container to use for the task | us-docker.pkg.dev/general-theiagen/theiagen/alpine-plus-bash:3.20.0 | Optional |
| version_capture | timezone | String | Set the time zone to get an accurate date of analysis (uses UTC by default) | Optional |
Workflow Tasks¶
versioning: Version Capture
The versioning task captures the workflow version from the GitHub (code repository) version.
Version Capture Technical details
| Links | |
|---|---|
| Task | task_versioning.wdl |
allele_clustering: PulseNet 2.0 Hash-Based Allele Clustering
The Allele Clustering module is used by PulseNet 2.0 to generate NWK trees for visualization, using the results from the allele_clustering task available in the TheiaProk workflows.
To run this task, a tree building algorithm and distance algorithm must be specified; these options are available in the inputs section of the workflow documentation.
This module adapts the PulseNet 2.0 code for the calculation of distance matrices from allele profiles or sequence alignment and distance-based tree inference for implementation on Terra.bio, and was developed in collaboration with the Association of Public Health Laboratories (APHL) and Centers for Disease Control and Prevention (CDC)'s Enteric Diseases Laboratory Branch (EDLB). We gratefully acknowledge the developers of the allele clustering algorithm and the CDC EDLB team.
Allele Clustering Technical Details
| Links | |
|---|---|
| Task | task_allele_clustering.wdl |
| Software Source Code | PulseNet 2.0 Trees |
| Software Documentation | PulseNet 2.0 Trees on GitHub |
Outputs¶
| Variable | Type | Description |
|---|---|---|
| allele_clustering_distance_matrix | File | The distance matrix of the resulting tree |
| allele_clustering_tree | File | The NWK file containing the resulting tree |
| allele_clustering_wf_analysis_date | String | Date of analysis |
| allele_clustering_wf_version | String | Version of PHB used for analysis |
| concatenated_allele_jsons | File | The file containing all of the input JSONs concatenated, primarily useful for debugging |
