7. Challenges and limitations
Deep learning-based microbial species differentiation remains highly dependent on the quality, diversity, and
representativeness of training data and reference databases. Training datasets may be biased toward well-
studied organisms, geographic regions, environments, sequencing platforms, and publicly available genomes,
causing models to perform poorly on underrepresented or novel taxa. Likewise, incomplete or biased reference
databases can skew taxonomic classification and prevent identification of organisms with no close
references.
[52,53]
Another challenge is differentiating closely related species, especially when their genomes are
highly similar. HGT exacerbates the problem as genes can move across taxonomic groups, which may lead to
models that rely heavily on gene content or sequence features to assign incorrect taxa. These issues are
especially relevant for metagenomic datasets that contain novel, rare, or highly diverse species.
[2,4]
The second
main issue concerns model reliability and interpretability. Deep learning models can be black boxes, and we
often do not know which genomic features lead the model to assign a particular species, or whether those
features reflect a biologically relevant signal. Domain shift also poses a challenge: models trained on data from
one laboratory, geographic area, sequencing platform, or community may not perform as well on another.
Finally, data leakage can occur when closely related genomes or samples appear in both training and testing,
inflating performance estimates. Therefore, external validation on independent datasets that differ
taxonomically and/or geographically is important yet infrequently done. A recent review showed that many
studies applying deep learning to metagenomics use relatively small datasets, simulations, or closely related
validation datasets, which limits our understanding of generalizability.
[2,54]
Lastly, widespread adoption requires
addressing computational cost, reproducibility, and standardization. Large deep learning models can require
significant computational power for training and inference, while differences in sequencing protocols,
preprocessing steps, databases, model architectures, hyperparameters, and performance metrics can limit
reproducibility and comparability between studies. To ensure that models perform well and can be compared
across platforms and different biological contexts, we need standardized datasets, benchmarking methods,
reporting guidelines, and independent validation. Studies should provide full details of preprocessing and
model specifications, avoid data leakage, and assess generalizability on truly independent datasets, with
additional attention to explainability and uncertainty estimation. Ultimately, the continued advancement of
microbial species differentiation requires not only the development of more accurate models but also the
creation of standardized, diverse, reproducible, and biologically representative benchmarking frameworks.
[2,55,53]
Overall, for species-level taxonomic profiling, Kraken2 and MetaPhlAn remain highly competitive, while
DL-based genome binners SemiBin and COMEBin outperform commonly used methods such as MetaBAT2.
Therefore, current results suggest that DL-based methods are not always superior to standard methods for all
metagenomic applications. Another major challenge is strain-level taxonomic profiling. Differentiating strains
remains an unsolved challenge not only for DL-based approaches but also for standard methods, owing to the
large number of shared sequences and similarities between strains. CAMI results also reflect inconsistent
improvements from applying DL at the strain level. To understand if DL offers any improvement for strain-level
profiling rather than just species-level profiling, extensive comparisons with state-of-the-art tools using
identical data with similar evaluation measures such as precision, recall, F1-score, and abundance estimation,
which unfortunately are lacking in the current studies, are necessary.
[57]
Conclusion
Deep learning has moved from a niche method to a mainstream tool for metagenomics and has entered
alignment-free architectures for taxonomic classification, joint representation learning approaches for genome
binning, sequence-based approaches for predicting genes and functions, and increasingly powerful
representations for associating microbiome data with host phenotypes. We expect these applications to
continue to benefit from improved data representations (e.g., from one-hot encoding to k-mer embeddings and
pretrained genomic language models), as well as improvements to model architectures. As the field moves
forward, we see increasing interest in combining multiple model architectures (e.g., graph-aware transformer,
taxonomy-informed autoencoder) and multiple types of features (e.g., composition, abundance, assembly graph,
and prior biological knowledge) within a single model, while at the same time placing greater emphasis on
interpretability, computational cost, and reference-free generalization to the vast amounts of microbial diversity
that have yet to be characterized. Despite this progress, notable challenges remain. Although deep learning
models show impressive results on benchmark datasets, it is unclear how well these translate to biological
generalization because of data leakage, taxonomic overlap, benchmark contamination, reliance on references,
sequencing platform differences, and geographic and environmental biases. Additionally, functional and
phenotype predictions are not yet mature because the underlying processes that drive these traits are governed
by complex biological and environmental factors that cannot always be inferred accurately from sequence alone.
No consensus exists on the gold standard for species-level classification, and most studies lack external
validation, making direct comparisons across models challenging. Together, these limitations make it difficult
to apply these tools in clinical or environmental scenarios, where models must be robust to variation in