Background: Bioinformatics pipelines spanning genomics, transcriptomics, proteomics, and metagenomics face a pervasive reproducibility crisis driven by software dependency drift, resource exhaustion, and non-deterministic tool behaviour. Despite substantial investment in workflow management systems such as Nextflow, Snakemake, and Galaxy, pipeline failure responses remain predominantly manual. Advances in large language models (LLMs), retrieval-augmented generation (RAG), and multi-agent orchestration offer technically credible pathways to automated pipeline failure detection and autonomous remediation, yet their systematic application to bioinformatics-specific infrastructure has not been reviewed. Results: This PRISMA 2020-compliant systematic review (PROSPERO: CRD420261361756) synthesised 26 studies from six databases covering January 2019 to April 2026, addressing five research questions spanning LLM-based log parsing, multi-agent workflow orchestration, human-in-the-loop governance, fault-tolerant infrastructure patterns, and benchmarking gaps. LLM-RAG frameworks achieved up to 80% workflow step recall and reduced manual curation time by over 90%. Multi-agent systems including BioMaster and MARWA demonstrated superior error recovery across 18 omics modalities and 102 bioinformatics tools, consistently outperforming single-agent baselines. Infrastructure foundations are technically mature, but no included study demonstrated an end-to-end integrated self-healing pipeline combining monitoring, anomaly detection, remediation, and governance-compliant audit trail generation. Evidence quality was moderate (mean 6.2/10; range 5–9). Conclusions: Governance frameworks, bioinformatics-specific benchmarks, and regulatory alignment for clinical contexts remain critically absent and represent the field’s primary bottleneck. We propose a four-layer conceptual model (Infrastructure, Observability, Orchestration, Governance) to organise current evidence and identify research priorities. Three directions warrant priority investment: standardised failure-injection benchmarks, end-to-end integrated pipeline validation in production bioinformatics environments, and LLM fine-tuning on bioinformatics-specific log formats.
Loading....