ALL Metrics
-
Views
-
Downloads
Get PDF
Get XML
Cite
Export
Track
Systematic Review

Comparative Genomic Insights into Pangenome Diversity and Functional Adaptation of Akkermansia and Faecalibacterium as Gut-Associated Next-Generation Probiotics: A Systematic Review

[version 1; peer review: awaiting peer review]
PUBLISHED 25 Jul 2026
Author details Author details
OPEN PEER REVIEW
REVIEWER STATUS AWAITING PEER REVIEW

Abstract

Background

Gut-associated next-generation probiotics (NGPs), particularly Akkermansia and Faecalibacterium, are promising candidates for restoring gut homeostasis and managing dysbiosis-related diseases. Comparative genomics can reveal their pangenome diversity, lineage-specific traits, and functional adaptations. This review aimed to synthesize comparative genomic evidence on their probiotic-associated genomic features and identify current research gaps.

Methods

A systematic review was conducted in accordance with the PRISMA 2020 guidelines. Comprehensive literature searches were performed in PubMed, Scopus, and The Lens to identify original comparative genomic studies published between January 2020 and December 2025. From 167 records initially identified, 17 studies met the predefined eligibility criteria and were included in the qualitative synthesis. Data were systematically extracted and narratively synthesized to compare comparative genomic approaches, pangenome characteristics, and functional adaptations of Akkermansia and Faecalibacterium.

Results

Comparative genomics showed that Akkermansia and Faecalibacterium have high genomic diversity, open pangenomes, and large accessory or strain-specific gene repertoires, indicating lineage- and strain-dependent probiotic potential. Akkermansia was mainly characterised by functional traits associated with the utilization of mucin, human milk oligosaccharides, and host-glycan. These traits were accompanied by mucin-degrading enzymes, epithelial adherence-related features, pili, autotransporters, oxygen-tolerance mechanisms, and genes associated with vitamin B12 biosynthesis. In contrast, Faecalibacterium showed broader carbohydrate utilization, glycan-degrading enzymes, trehalose metabolism, extracellular polysaccharide (EPS) or capsule genes, acetate/butyrate-associated metabolism, and strain-specific anti-inflammatory traits.

Conclusions

Comparative genomics provides a useful framework for strain-level prioritization of Akkermansia and Faecalibacterium as gut-associated next-generation probiotic candidates. The evidence indicated broad pangenome diversity and lineage-dependent functional potential in both genera. However, safety-related genomic features, including antimicrobial resistance genes, virulence or pathogenicity markers, horizontal gene transfer signals, and risk indices, were assessed unevenly across studies. Future development of these taxa as next-generation probiotics should therefore combine comparative genomics with standardized genome-based safety screening and experimental functional validation.

Systematic Review Registration

This systematic review was registered in the International Prospective Register of Systematic Reviews (PROSPERO; Registration No. CRD420261437920) on July 2, 2026.

Keywords

Akkermansia muciniphila, comparative pangenomics, Faecalibacterium prausnitzii, functional genomics, gut microbiota, microbiome therapeutics, whole-genome sequencing

Introduction

The human gastrointestinal tract harbors a complex and highly diverse microbial ecosystem known as the gut microbiota, which plays an essential role in maintaining host physiology and health.1 The human gut microbiota plays a fundamental role in regulating nutrient metabolism, maintaining intestinal barrier integrity, immune homeostasis, and protection against invading pathogens.2,3 Through these activities, gut microbiota produce a broad range of bioactive metabolites, including short-chain fatty acids (SCFAs), vitamins, and other bioactive compounds that regulate host metabolic and immunological functions.4 Conversely, disturbances in the composition and function of the gut microbiota, commonly referred to as gut dysbiosis, have been implicated in the pathogenesis of numerous chronic disorders, including asthma, autism spectrum disorder (ASD), colorectal cancer, Clostridium difficile infection, diabetes mellitus, hypertension, inflammatory bowel disease (IBD), obesity, neurodegenerative diseases, and other immune-mediated conditions.5 These findings highlight the importance of strategies aimed at restoring and maintaining gut microbiota homeostasis.

Among microbiota-targeted interventions, probiotics have been widely investigated for their ability to promote gut health. According to the joint Food and Agriculture Organization of the United Nations (FAO) and World Health Organization (WHO), probiotics refer to live microorganisms that, when administered in adequate amounts, confer a health benefit on the host. Conventional probiotics are predominantly lactic acid bacteria, particularly Lactobacillus and Bifidobacterium, which promote host health by inhibiting pathogens, producing beneficial metabolites, and modulating immune responses.6 Recent advances in gut microbiota research have identified the human gut as a rich reservoir of previously unexplored microorganisms with promising probiotic potential. Consequently, increasing research efforts have focused on identifying and characterising novel gut-derived bacterial taxa as candidates for next-generation probiotics.7

Next-generation probiotics are a group of gut-derived commensal microorganisms that meet the conventional definition of probiotics but have not previously been used as health-promoting probiotic agents, making them promising candidates for targeted microbiome-based therapeutics.8 To date, numerous gut-associated bacterial taxa have been proposed as next-generation probiotic candidates, including Akkermansia, Bacteroides, Christensenella, Faecalibacterium, and many others.9 Among these, Akkermansia muciniphila and Faecalibacterium prausnitzii have been recognized as leading candidates because of their potential roles in the prevention and treatment of dysbiosis-associated diseases.10 Therefore, these two genera provide an ideal genomic framework for exploring lineage diversification, functional adaptation, and probiotic-associated traits.

The rapid expansion of whole genome sequencing (WGS) has substantially advanced comparative genomic investigations of gut-associated NGPs, allowing researchers to characterize genomic diversity beyond taxonomic classification.11 Comparative genomic and pangenome analyses have increasingly been applied to Akkermansia and Faecalibacterium to investigate core and accessory genome composition, biosynthetic gene clusters, stress adaptation, and other genomic determinants associated with probiotic functionality and ecological adaptations.1215 However, the available evidence remains dispersed across individual comparative genomic studies that differ considerably in genome datasets, analytical pipelines, and functional interpretation, making it difficult to obtain an integrated understanding of the conserved and lineage-specific genomic features underlying probiotic potential.

Although several studies have summarized the biological functions, therapeutic potential, and clinical application of next-generation probiotics, a systematic synthesis integrating comparative genomic evidence on pangenome diversity and functional adaptation of Akkermansia and Faecalibacterium remains lacking.16 Therefore, this systematic review focused on original comparative genomic studies investigating two representative gut-associated next-generation probiotic genera. The objective was to synthesize current evidence on how comparative genomics has been used to characterize pangenomic diversity, functional adaptation, host-associated traits, and safety-related genomic features, while identifying current research trends, knowledge gaps, and priorities for genome-informed strain selection.

Materials and methods

Study design

This systematic literature review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines. The review protocol was developed to ensure methodological rigor, reproducibility, and transparency. The protocol was registered in the International Prospective Register of Systematic Reviews (PROSPERO) under the registration number CRD420261437920. To minimize selection bias and improve reproducibility, the review objectives, eligibility criteria, literature search strategy, study selection procedures, data extraction framework, and qualitative synthesis approach were specified in advance.

Literature search strategy

A comprehensive literature search was conducted in three electronic databases, namely Scopus, PubMed, and The Lens, to identify relevant studies published between January 2020 and December 2025. The search was performed in June 2026 using keyword combinations related to gut-associated next-generation probiotics and genome-scale comparative analyses. The search strategy combined bacterial taxa (Akkermansia and Faecalibacterium) with comparative genomic terms, including comparative genomics, comparative genome analysis, phylogenomics, pangenome, and pan-genome. Database-specific search fields were used in accordance with each database’s indexing system, including Title, Abstract, and Keywords (Scopus); Title and Abstract (PubMed); and Title, Abstract, Keywords, or Field of Study (The Lens). Database-specific filters were subsequently applied to retain only original research articles published in English between 2020 and 2025.

Eligibility criteria

Studies were considered eligible if they met all of the following criteria: (i) original research articles published in peer-reviewed journals between 1 January 2020 and 31 December 2025; (ii) studies investigating one or more of the gut-associated next-generation probiotic genera included within the scope of this review, namely Akkermansia and Faecalibacterium; (iii) studies employing whole-genome sequencing (WGS) data to perform comparative genomics analyses, including whole-genome comparison, pangenomics, genome-informed phylogenomics, or other genome-scale comparative approaches; (iv) studies were required to report comparative genomic characteristics associated with functional adaptation, genomic diversity, metabolism, host interaction and adaptation, or probiotic-associated traits.

Studies were excluded if they were review articles, conference papers, editorials, book chapters, genome announcements, taxonomic descriptions, methodological reports without comparative genomic analyses, or genome-based species reclassification lacking functional comparative genomic investigation, metagenomic association studies without comparative genomic analysis of target taxa, or non-English publications.

Study selection

All records identified through PubMed, Scopus, and The Lens were exported into EndNote 21 (Clarivate Analytics) for reference management. Database-specific filters were applied before export, and duplicate records were subsequently identified and removed. The remaining records were independently screened by two reviewers in a two-stage process using the predefined eligibility criteria. In the first stage, titles, abstracts, and keywords were screened to exclude clearly irrelevant records. Articles considered potentially eligible proceeded to full-text assessment. During the full-text assessment, studies were retained only when comparative genomic analysis represented one of the principal analytical components of the study. Any disagreements regarding study eligibility were resolved through discussion. When necessary, a third reviewer was consulted to reach a final decision. The reasons for excluding full-text articles were documented throughout the selection process. The study selection process was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines.

Data extraction and analysis

Data were extracted using a predefined extraction framework covering study characteristics, genomic datasets and their sources, comparative genomic approaches, pangenome features, phylogenomic findings, functional genomic traits, host-adaptation features, and safety-related genomic information. Extracted data were then organized into three synthesis tables: characteristics of included comparative genomic studies ( Table 2), comparative genomic, pangenomic, and phylogenomic features in included Akkermansia and Faecalibacterium studies ( Table 3), and comparative functional genomic traits associated with gut adaptation and probiotic-associated potential in Akkermansia and Faecalibacterium ( Table 4).

Quality assessment

Given the nature of the included studies, which consisted of comparative genomic and pangenomic analyses rather than clinical intervention or observational studies in living subjects, conventional risk-of-bias tools such as the Cochrane Risk of Bias tool (RoB 2), ROBINS-I, or QUADAS-2 were considered methodologically inappropriate. Therefore, the tools were not applied. Instead, a custom instrument termed the Genome and Pangenome Quality Assessment Checklist (GPQAC) was developed specifically for this review to evaluate the methodological rigour and reporting quality of the included genomic studies. The complete checklist, including all assessment domains, scoring criteria, and score interpretation, is publicly available as extended data in the Open Science Framework (OSF) repository associated with this review (see Extended Data File Quality Assessment Checklist Table S1 and S2; https://doi.org/10.17605/OSF.IO/C3N6M).17

The GPQAC comprises two sections: Section A assesses input genome quality across five domains, including genome completeness and contamination, assembly quality, taxonomic identity validity, sequencing platform and coverage, and data availability and reproducibility (maximum score: 10 points); Section B evaluates the methodological quality of pangenome construction and functional annotation across seven domains, including genome representativeness, tool and parameter reporting, gene category definitions, functional annotation methods, validation of probiotic functional claims, supporting phylogenetic analysis, and pipeline reproducibility (maximum score: 14 points). Each domain is scored 0 (not reported or poor), 1 (partially reported), or 2 (fully reported and meets the defined threshold), for a maximum score of 24 points. Studies were classified as low (0–8), moderate,916 or high quality1724 ( Table 1).

Table 1. Genome and pangenome quality assessment checklist (GPQAC).

Author (year) Section A (Input Genome Quality) Section B (Pangenome and Functional Quality Analysis) (Section A + B) Total ScoreQuality
Bai et al. (2023)3941418High quality
Becken et al. (2021)1271320High quality
Bukhari et al. (2022)4051217High quality
De Filippis et al. (2020)1371219High quality
Fatima et al. (2021)224812Moderate
Gao et al. (2022)316814Moderate
Geerlings et al. (2021)236915Moderate
González et al. (2023)1461420High quality
Kim et al. (2022)2081321High quality
Kirmiz et al. (2020)2181220High quality
Li et al. (2022)4441317High quality
Li et al. (2024)1571119High quality
Lu et al. (2024)416915Moderate
Luna et al. (2022)305914Moderate
Ouwerkerk et al. (2022)247815Moderate
Seo et al. (2025)325914Moderate
Ueda et al. (2021)3371118High quality

The checklist was adapted from the Minimum Information about a Metagenome-Assembled Genome (MIMAG) and Single Amplified Genome (MISAG) standards established by the Genomic Standards Consortium18 and from established pangenome analysis methodology.19 Before full-scale application, the GPQAC was piloted on a subset of five included studies to assess item clarity and consistency. Inter-rater reliability between two independent reviewers was calculated using Cohen’s kappa, with discrepancies resolved through discussion or adjudication by a third reviewer. Quality assessment scores were reported descriptively and used to contextualise the interpretation of findings. However, no studies were excluded based on their quality assessment scores to preserve a comprehensive evidence base for synthesis.

Results

Literature search and study selection

The literature search identified a total of 167 records across three electronic databases (Scopus, PubMed, and The Lens). After applying database-specific filters, including publication year (2020–2025), document type (original research articles), and language (English), 102 records were exported for screening. Following the removal of 54 duplicate records, 48 unique records remained for title, abstract, and keyword screening. Of these, 22 records were excluded, and 26 reports were sought for full-text retrieval. All reports were successfully retrieved and assessed for eligibility. After full-text assessment, nine studies were excluded for not meeting the predefined eligibility criteria, resulting in 17 studies being included in the qualitative synthesis. The study selection process followed the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines. A detailed overview of the identification, screening, eligibility assessment, and inclusion process is presented in the PRISMA 2020 flow diagram ( Figure 1).

f1df7fd7-c3f6-4d51-be1a-c7991e881cbf_figure1.gif

Figure 1. PRISMA 2020 flow diagram of the literature selection process.

Quality assessment of included studies

The methodological quality of all 17 included studies was evaluated using the Genome and Pangenome Quality Assessment Checklist (GPQAC), with total scores ranging from 12 to 21 out of 24 points ( Table 1). Overall, the majority of included studies demonstrated high methodological quality, with 10 studies (59%) classified as high quality (total score 17–21) and 7 studies (41%) classified as moderate quality (total score 12–15). No studies were classified as low quality, indicating an acceptable level of methodological rigour across the evidence base. Among the high-quality studies, the highest scores were achieved by Kim et al. (2022)20 (score: 21), followed by Becken et al. (2021),12 González et al. (2023),14 and Kirmiz et al. (2020)21 (score: 20 each), and De Filippis et al. (2020)13 (score: 19). Among the moderate-quality studies, scores ranged from 12 (Fatima et al. 2021)22 to 15 (Geerlings et al. 202123; Ouwerkerk et al. 202224).

Regarding section-specific performance, Section B scores (Pangenome and Functional Analysis Methodological Quality, maximum 14 points) were generally higher relative to their respective maximum than Section A scores (Input Genome Quality, maximum 10 points), suggesting that while most studies applied rigorous pangenomic and functional annotation approaches, reporting of input genome quality metrics, such as completeness, contamination estimates, and assembly statistics was comparatively less consistent across studies. This pattern highlights a recurring methodological gap in the transparent reporting of genome-level quality control in comparative genomic studies of gut-associated bacteria.

Characteristic and genomic scope of the included studies

Table 2 summarizes the characteristics of the included comparative genomic studies on Akkermansia and Faecalibacterium. The studies varied in taxonomic scope, ranging from strain-level analyses to broader phylogroup, clade, or genus-level comparisons. Genome sources included newly sequenced isolates, public reference genomes, complete genomes, and metagenome-assembled genomes. The approaches used across studies included comparative genomics, pangenomics, phylogenomics, ANI-based analysis, functional annotation, pathway reconstruction, CAZyme profiling, CRISPR/Cas screening, phenotypic validation, and genome-based safety assessment.

Table 2. Characteristics of included comparative genomic studies.

Author (year)Genus/SpeciesGenome type and sourceComparative genomic approaches
Bai et al. (2023)39Faecalibacterium prausnitzii 84 complete genomes from NCBIComparative genomic, pangenomic, and functional analyses
Becken et al. (2021)12Akkermansia muciniphila 43 newly sequenced genomes from the POMMS cohort and 6 public NCBI/JGI IMG reference genomesComparative genomic and pangenomic study with phenotypic and in vivo validation
Bukhari et al. (2022)40Akkermansia muciniphila 19 complete genomes from NCBIComparative genomic and pangenomic study with functional analyses
De Filippis et al. (2020)13Faecalibacterium 2,859 MAGs, including 64 genomes from NCBI, 2,784 human gut MAGs, and 11 NHP gut MAGsComparative genomic study with functional analyses
Fatima et al. (2021)22Akkermansia muciniphila 19 complete genomes from NCBIComparative genomic and pangenomic study with functional analyses
Gao et al. (2022)31Akkermansia muciniphila 7 complete genomes, including a new strain from mouse feces and 6 reference genomes from NCBIComparative genomic study with functional analyses
Geerlings et al. (2021)23Akkermansia muciniphila 10 complete genomes, including 9 sequenced genomes from several animal feces and a reference strain from NCBIComparative genomic and pangenomic study with functional analyses
González et al. (2023)14Akkermansia 367 MAGs from NCBIComparative genomic and pangenomic study with functional analyses
Kim et al. (2022)20Akkermansia muciniphila 27 genomes, including 5 newly sequenced human feces-derived strains and public NCBI genomesComparative genomic and pangenomic study with sulfatase activity validation
Kirmiz et al. (2020)21Akkermansia 35 reconstructed Akkermansia MAGs from human fecesComparative genomics, pangenomics, and vitamin B12 biosynthesis validation
Li et al. (2022)44Akkermansia muciniphila 112 genomes from NCBIComparative genomic, pangenomic, and recombination analysis
Li et al. (2024)15Faecalibacterium 136 genomes, including 29 newly sequenced human gut-derived strains and 107 public genomes from NCBIComparative genomic and pangenomic study with functional analyses
Lu et al. (2024)41Akkermansia 20 genomes, including 4 newly sequenced human feces-derived isolates and 16 reference genomes from NCBIComparative genomic study with functional analyses
Luna et al. (2022)30Akkermansia 85 genomes, including 11 newly sequenced isolates from human feces and 74 public genomes from NCBIComparative genomic study with functional analyses
Ouwerkerk et al. (2022)24Akkermansia muciniphila 8 complete genomes, including 6 newly isolated human-intestinal strains from healthy donors and two type-strain genomes from NCBIComparative genomic, physiological, and proteomic characterization
Seo et al. (2025)32Faecalibacterium prausnitzii 52 complete genomes, including 2 newly sequenced isolates from humans and 50 genomes from NCBIComparative genomic and pangenomic study with functional analyses
Ueda et al. (2021)33Faecalibacterium prausnitzii 12 complete genomes from fecal samples of MCI and Alzheimer’s disease subjectsComparative genomic and pangenomic study with functional analyses

Comparative genomic, pangenomic, and phylogenomic features

Overall, Faecalibacterium tended to show broader genome-size variation and deeper strain-, clade-, or species-level diversification, whereas Akkermansia was more consistently structured into AmI–AmIV phylogroups, Amuc-related clusters, or genomic species clusters. Pangenome analyses in both genera commonly indicated open or highly variable pangenomes, with differences in core, accessory, cloud, unique, and strain-specific gene fractions across studies. Several articles also reported phylogenomic and evolutionary patterns, including ANI-based clustering, lineage divergence, gene gain/loss, recombination, HGT, and effective-strain-specific ortholog profiles. Summaries of the comparative genomic, pangenomic, and phylogenomic features reported in Akkermansia and Faecalibacterium studies are provided in Table 3.

Table 3. Comparative genomic, pangenomic, and phylogenomic features in included Akkermansia and Faecalibacterium studies.

Author (year)Comparative genome featuresPangenome composition and architecturePhylogenomic and evolutionary insights
Bai et al. (2023)39Genome sizes ranged from 2.66–3.42 Mb, with 52.9–58.1% GC content and 1,991–3,264 predicted proteins across 84 strains8,240 homologous gene clusters were identified, including 375 core genes, 207 unique genes, and extensive accessory genes. The pangenome was open, with 4.5% core genesPhylogenomic analyses based on 303 single-copy core genes, 16S rRNA, and pan-genome composition resolved 84 strains into six phylogenetic groups
Becken et al. (2021)12Genome sizes ranged from 2.6 to 3.3 Mb, with AmII and AmIV genomes generally larger than AmI genomes4,982 gene clusters were identified, including 1,647 core genes and 506 singleton gene clusters; pangenome status was not reportedComparative phylogenomics resolved four phylogroups (AmI–AmIV), with further subdivision into AmIa and AmIb at 96% ANI
Bukhari et al. (2022)40Gene numbers ranged from 2,140 to 2,738 per genome, with an average of 2,335 genesAn open pangenome comprised 7,625 genes, including 1,906 persistent, 626 shell, 2,134 cloud, and 2,945 unique genesComparative phylogenomics and ANI resolved four phylogroups (AmI–AmIV)
De Filippis et al. (2020)13Genome statistics were not reported; genome-wide comparison indicated inter-clade differentiation based on MASH and gene-content distancesPangenome analysis was not performedPhylogenomics resolved 22 Faecalibacterium-like clades, with host- and lifestyle-associated stratification
Fatima et al. (2021)22Analysis of 19 complete genomes showed genome sizes of 2.66–2.82 Mb, 55.25–55.82% GC content, 2,130–2,315 protein-coding genes, 52–53 tRNAs, and 9 rRNAsThe species exhibited an open pangenome comprising 4,933 gene families, including 1,035 core genes, with increasing unique genes as additional genomes were incorporatedCore-genome phylogenetic analysis resolved evolutionary relationships among the 19 strains based on shared orthologous genes
Gao et al. (2022)31Genome size was 2.81 Mb with 55.32% GC content and 2,905 CDS; the genome was 142 kb larger than the reference strainPangenome analysis was not performedPhylogenomic analysis assigned MucX to the AmI phylogroup. ANI values were > 98% within AmI but 85.11–86.71% with other phylogroups
Geerlings et al. (2021)23Genome sizes ranged from 2.7–2.9 Mb with 55.2–55.9% GC content and 2,233–2,629 predicted genesComparative pangenome analysis was performed, but quantitative metrics and pangenome status were not reportedPhylogenomic analyses identified highly similar isolates (>99.7% ANI) and divergent lineages (93.9–97.4% ANI)
González et al. (2023)14Comparative analysis focused on 367 high-quality genomes selected from 600 candidates using ≥90% completeness and < 5% contamination thresholdsThe four largest genomic species clusters exhibited open pangenomes, with cluster-specific variation in total, core, and unique genesPhylogenomic and ANI analyses resolved two major clades and 25 genomic species clusters; evolutionary analyses indicated gene gain/loss and horizontal gene transfer
Kim et al. (2022)20Genome sizes ranged from 2.66–2.86 Mb with 55.23–55.76% GC content and 2,149–2,363 CDSsThe open pangenome comprised 3,811 gene families, including 1,749 core genes, 1,255 accessory genes, and 807 unique genesPhylogenomics resolved two major lineages associated with geographic origin
Kirmiz et al. (2020)21Genome sizes ranged from 2.27–3.20 Mb with 55.2–58.7% GC content; genome completeness ranged from 68.5–100%The pangenome comprised 6,557 gene clusters, including 1,021 core genes and 1,240 unique gene clustersComparative phylogenomics and ANI resolved four phylogroups (AmI–AmIV)
Li et al. (2022)44Analysis of 112 genomes revealed differences in genome size, GC content, and CDS number among genetic lineagesThe species possessed an open pangenome of 9,419 genes, including core, soft-core, shell, and cloud gene fractionsCore-genome phylogeny and ANI resolved three genetic lineages (A–C); lineage-specific recombination patterns were identified
Li et al. (2024)15Comparative analysis of 136 genomes revealed significant inter-cluster variation in genome size, GC content, and gene numberThe genus exhibited an open pangenome comprising 15,261 gene families, with high accessory and unique gene fractionsCore-genome phylogeny resolved 11 species-level clusters; high 16S rRNA similarity contrasted with inter-cluster ANI <95%
Lu et al. (2024)41Four isolates possessed larger genomes and more CDSs than the type strain, with similar tRNA counts and lower rRNA copy numbersThe isolates possessed 2,184 shared homologous genes and 22–105 strain-specific genes; the composition and status of the pangenome were not assessedPhylogenomic analysis assigned all four isolates to the AmII phylogroup, distinct from the AmI type strain DSM22959
Luna et al. (2022)3085 genomes showed genome-size variation across Akkermansia strainsPangenome analysis was not performedPhylogenomics resolved four phylogroups (AmI–AmIV)
Ouwerkerk et al. (2022)24Genome sizes ranged from 2.5–2.8 Mb with 55.04–56.11% GC content and 2,055–2,350 genesPangenome analysis was not performedANI and comparative phylogenomics resolved two phylogenetic clusters (Amuc1 and AmucU)
Seo et al. (2025)32KBL1026, KBL1027, and DSM17677 had similar genome sizes and GC contents; ANI showed clear differentiation between KBL1027 and DSM17677.Pan-genome analysis of 50 isolates showed a very small core gene fraction and high strain-specific gene contentWhole-genome phylogenomics placed KBL1026 and KBL1027 in Faecalibacterium Clade B; strain-specific cell wall/capsule genes distinguished KBL1027
Ueda et al. (2021)33Whole-genome comparison was performed on 12 complete genomes, but standard genome statistics were not reportedClassical pangenome metrics were not reported; ortholog-based comparison identified effective-strain-specific orthologsStrains were distinguished by specific ortholog profiles, with selected orthologs enriched in healthy subjects compared with MCI subjects

Functional genomic traits linked to gut adaptation and probiotic potential

Table 4 summarizes functional genomic traits linked to gut adaptation and probiotic potential in Akkermansia and Faecalibacterium. Overall, Akkermansia was primarily associated with the utilization of mucin, human milk oligosaccharides (HMOs), and host-glycan, as well as mucosal adaptation traits, oxygen tolerance, and selected vitamin B12-related functions. In contrast, Faecalibacterium exhibited broader carbohydrate-related functions, including plant-polysaccharide degradation, glycan-degrading enzymes, trehalose metabolism, EPS/capsule biosynthesis, and acetate- and butyrate-associated metabolism. Additional probiotic-relevant traits included BGCs, CRISPR/Cas systems, bacteriocin-related genes, anti-inflammatory activity, and IL-10 induction. Safety-related features were unevenly assessed, particularly for ARGs, VFs, PGs, phage remnants, HGT-linked genes, AMR genes, and risk indices.

Table 4. Comparative functional genomic traits associated with gut adaptation and probiotic-associated potential in Akkermansia and Faecalibacterium.

Author (year)Carbohydrate utilizationSCFA traitsHost interaction and colonizationOther probiotic-relevant traitsSafety-related genomic features
Bai et al. (2023)39CAZyme analysis identified GT, CE, and GH familiesNot reportedNot reportedBGCs and CRISPR/Cas diversity were detectedARGs, VFs, PGs, and PPRI/PPRS-based probiotic risk classification were assessed
Becken et al. (2021)12Phylogroup-specific GH profiles were identifiedAcetate and propionate production was reportedEpithelial adherence and aggregation genes were reportedVitamin B12 genes, oxygen tolerance, iron acquisition, and sulfur metabolism were reportedCRISPR/Cas and possible phage-related regions were reported
Bukhari et al. (2022)40Carbohydrate metabolism genes were reportedPropionate-related pathways were identifiedNot reportedThiamine and methane metabolism genes were reportedARGs and autotransporter toxin genes were reported
De Filippis et al. (2020)13Clade-specific plant-polysaccharide and complex carbohydrate degradation genes were reportedNot reportedClade diversity was associated with host, diet, geography, age, lifestyle, obesity, and inflammatory statusDistinct functional potentials were observed among cladesARG-related genes were more prevalent in Western-associated profiles
Fatima et al. (2021)22Carbohydrate metabolism genes were reportedNot reportedNot reportedBacteriocin-related genes were detectedadeF/ermB resistance genes were detected
Gao et al. (2022)31HMO-degrading genes were reportedNot reportedNot reportedAmuc_1100/Amuc_1434/P9 homologs were reportedNot reported
Geerlings et al. (2021)23Conserved mucin-degradation genes were reportedIsolates produced acetate, propionate, and 1,2-propanediolNot reportedNot reportedPotential phage remnants were detected
González et al. (2023)14Lineage-specific mucin-degradation CAZymes were predictedGenome-inferred acetate production was reportedType IV pili, autotransporters, and adhesin-related genes were reportedOxidative-stress genes, mucin-associated peptidases, and conserved sulfatases were reportedHGT-linked GH families and AcrAB efflux pump-related genes were detected
Kim et al. (2022)20GH genes from multiple GH families were identifiedNot reportedNot reportedSulfur-metabolism genes were reportedNot reported
Kirmiz et al. (2020)21Not reportedSome strains produced acetate and propionateCore Type IV pilus genes were presentVitamin B12 synthesis genes and microaerobic metabolism were reportedNot reported
Li et al. (2022)44Lineage-specific CAZyme profiles were identifiedNot reportedNot reportedNot reportedNot reported
Li et al. (2024)15Multiple CAZyme families and glycan-degrading enzymes were identifiedComplete acetate and butyrate biosynthesis pathways were presentGH33-mediated mucin utilization and BSH genes were reportedBGCs, RiPPs, and thiamine metabolism genes were reported20 ARG types were detected in a few genomes
Lu et al. (2024)41GH and GT were the dominant CAZyme familiesNot reportedNot reportedComplete CRISPR-Cas systems were detectedAMR and virulence-related genes were detected
Luna et al. (2022)30Phylogroup-specific GH repertoires supported HMO degradationAcetate, succinate, and propionate production were reportedHMO utilization was linked to early gut colonization potentialVitamin B12 biosynthesis genes were reportedNot reported
Ouwerkerk et al. (2022)24Strain-specific GH variation was reportedAcetate and propionate production was reportedMucosal adaptation genes were reportedCRISPR diversity across strains was reportedNot reported
Seo et al. (2025)32Specific EPS/capsule-related glycan biosynthesis genes were identifiedButyrate production was reportedCapsule genes and EPS production were reportedStrain-specific anti-inflammatory activity and IL-10 induction were reportedNot reported
Ueda et al. (2021)33Trehalose metabolism was identifiedNot reportedPilus assembly genes were identifiedNot reportedNot reported

Cross-genus comparative synthesis of functional adaptation and probiotic potential

Cross-genus synthesis showed that the included studies differed in taxonomic scope, genome source, and comparative genomic depth, but collectively indicated distinct genomic patterns between Akkermansia and Faecalibacterium. Akkermansia studies were mostly centered on A. muciniphila or broader phylogroup-level diversity, with recurrent classification into AmI–AmIV, Amuc-related clusters, or genomic species clusters. In contrast, studies on Faecalibacterium examined strain-level comparisons of F. prausnitzii, as well as larger clade and genus-level datasets. Pangenome analyses in both genera commonly reported open or highly variable pangenomes, although several studies did not perform full pangenome reconstruction. Faecalibacterium generally showed broader strain-, clade-, and species-level diversification, whereas Akkermansia showed more consistent phylogroup-based structuring.

Functional genomic comparison further showed that the two genera differed in the main traits reported across studies. Akkermansia was mainly associated with the utilization of mucin, HMOs, and host-glycan, as well as with mucosal adaptation traits such as epithelial adherence, pili, autotransporters, sulfatases, oxygen tolerance, and selected vitamin B12-related functions. Faecalibacterium showed broader carbohydrate-related traits, including plant-polysaccharide degradation, glycan-degrading enzymes, trehalose metabolism, EPS/capsule-related biosynthesis, and acetate/butyrate-associated metabolism. The reporting of additional probiotic-relevant traits was uneven across studies and included BGCs, CRISPR/Cas systems, bacteriocin-related genes, sulfur metabolism, anti-inflammatory activity, and IL-10 induction. Safety-related evidence was also inconsistent across studies. Some studies assessed safety-related features, including ARGs, VFs, PGs, AMR genes, phage remnants, HGT-linked genes, AcrAB-related genes, or risk indices. In addition, the other studies did not include a systematic genome-based safety assessment.

Discussion

Next-generation probiotics have gained increasing attention as candidates for biotherapeutic, functional food, and nutraceutical applications; however, their development remains constrained by strain-specific nutritional requirements, oxygen sensitivity, and unresolved safety considerations.25 Comparative genomics is therefore important because it can resolve evolutionary divergence and strain-level functional variation that may be obscured by genus- or 16S rRNA-based classification, particularly in gut bacteria with high intra-taxon genomic diversity.26 In this review, Akkermansia was repeatedly structured into phylogroups or genomic species clusters. In addition, Faecalibacterium exhibited broader diversification at the clade and species levels, indicating that evolutionary divergence provides a framework for interpreting strain-level functional differences. Such clustering is biologically meaningful because phylogroup or clade separation can correspond to differences in oxygen tolerance, epithelial adherence, genome plasticity, and other traits relevant to gut colonization and next-generation probiotic candidate selection.12,13

The open pangenome patterns observed in both genera provide a genomic explanation for why probiotic-associated traits cannot be assumed to be uniformly conserved across all strains. In bacterial populations, the core genome generally represents conserved functions. In contrast, the accessory genome is more dynamic and may include genes acquired through mutation, recombination, horizontal gene transfer, plasmids, phages, transposons, or insertion sequences.27,28 In this review, both Akkermansia and Faecalibacterium showed large accessory, cloud, unique, or strain-specific gene fractions, indicating that many traits relevant to substrate utilization, host adaptation, metabolic specialization, and safety are likely distributed unevenly among lineages or strains. This supports a pan-ecological interpretation, in which pangenome variation reflects how bacterial populations adapt to different ecological niches, host interactions, and selective pressures rather than simply representing random gene-content differences.29 Therefore, for next-generation probiotic development, pangenome analyses should be used not only to describe genomic diversity but also to prioritize strains with the most appropriate combination of functional traits and low-risk genomic profiles.

The functional traits identified in this review suggest that Akkermansia and Faecalibacterium may support gut health through distinct but interconnected ecological strategies. Akkermansia was mainly associated with mucin degradation, human milk oligosaccharide utilization, and host-glycan utilization, reflecting its adaptation to the mucus layer and host-derived nutrient niches.14,30,31 Faecalibacterium, by comparison, exhibited broader carbohydrate-related functions, including dietary polysaccharide utilization, glycan-degrading enzymes, trehalose-related metabolism, and extracellular polysaccharide or capsule-associated genes.13,15,32,33 These differences are biologically relevant because carbohydrate utilization shapes bacterial growth, gut niche preference, and downstream metabolite production, which may influence host physiology.34 Akkermansia muciniphila has been experimentally shown to use mucin-derived monosaccharides, including fucose, galactose, N-acetylglucosamine, and N-acetylgalactosamine, supporting its adaptation to the mucus layer.35 In contrast, cultured representatives of major Faecalibacterium prausnitzii phylogroups have been shown to use pectin, uronic acids, and host-derived substrates, supporting their role as anaerobic fermenters in the intestinal ecosystem.36 This metabolic difference was also reflected in SCFA traits, with Akkermansia associated with acetate, propionate, succinate, and 1,2-propanediol production, and Faecalibacterium more strongly linked to acetate and butyrate-associated metabolism.15,23,30,32 Butyrate is a key energy source for intestinal epithelial cells and contributes to barrier integrity, T-cell regulation, pathogen resistance, inflammatory response, and metabolic homeostasis.36,37

These metabolic traits are further connected to host interaction, immune regulation, and ecological persistence. In this review, Akkermansia was characterized by several host-interaction and persistence-related features, including epithelial adherence, aggregation, pili, autotransporters, sulfatases, mucin-associated enzymes, oxygen tolerance, Amuc-related proteins, and genes associated with B12/corrin-related biosynthesis.12,14,21,24,31 In contrast, Faecalibacterium displayed extracellular polysaccharides or capsule genes, pilus-related orthologs, bile salt hydrolase genes, IL-10-associated responses, and strain-specific anti-inflammatory activity.15,32,33,37 In Akkermansia, adhesion to the mucus layer is relevant because access to mucins supports its ecological role.24,35 The outer membrane protein Amuc_1100 has been reported to interact with TLR2 and improve gut barrier function, partly recapitulating the beneficial effect of Akkermansia muciniphila.38 In Faecalibacterium, butyrate and secreted anti-inflammatory metabolites have been associated with inhibition of NF-κB activation and IL-8 production.37 Together, these findings indicate that Akkermansia and Faecalibacterium are complementary rather than interchangeable next-generation probiotic candidates.

Although the included studies reported several functional genomic traits supporting the potential of Akkermansia and Faecalibacterium as next-generation probiotic candidates, the safety evidence remains incomplete and uneven across studies. Based on this review, safety-related features mainly included antimicrobial resistance, virulence or pathogenicity markers, phage-related regions, efflux-related genes, CRISPR/Cas systems, and probiotic risk classification, but these assessments were not applied consistently across all datasets.13,15,22,23,3941 This uneven safety assessment weakens the current evidence because probiotic safety cannot be inferred from beneficial functional traits alone and must be supported by strain-level safety evaluation, including assessment of transferable antimicrobial resistance, virulence-associated genes, mobile genetic elements, and toxicological evidence.42 This issue is particularly relevant because ARGs carried by probiotic candidates may be transferred to intestinal bacteria through horizontal gene transfer, potentially reducing the effectiveness of antibiotic therapy when acquired by pathogenic bacteria.43

Overall, comparative genomic evidence shows that Akkermansia and Faecalibacterium are promising but strain-dependent next-generation probiotic candidates, with potential shaped by pangenome diversity, functional adaptation, and safety profiles. These findings emphasize the need for strain-specific genomic, functional, and safety validation before their translation into probiotic, functional food, or biotherapeutic applications.

Implications, limitations, and research gaps

Implications

The findings of this review suggest that comparative genomics can serve as an early-stage decision-making framework for selecting and prioritizing Akkermansia and Faecalibacterium strains before experimental validation. Rather than treating probiotic potential as a genus-level property, future development should adopt a strain-centered strategy that integrates genomic function, ecological adaptation, and evidence of safety. This approach can help reduce the number of unsuitable candidates at an early stage, guide the design of functional assays, and support more targeted preclinical or clinical evaluation. Therefore, comparative genomics should be viewed not only as a descriptive tool but as a practical screening step for developing safer, more evidence-based next-generation probiotics.

Limitations

This review is limited by the heterogeneity of the included studies in terms of dataset size, genome source, sequencing quality, taxonomic resolution, and analytical depth. Some studies used complete genomes, whereas others relied on draft genomes, public genome datasets, or genome-resolved assemblies, which may affect the consistency of pangenome and functional comparisons. In addition, not all studies performed full pangenome analysis, pathway reconstruction, or systematic safety screening, making direct comparison across articles uneven. Many reported probiotic-associated traits were also genome-inferred, meaning that the presence of a trait did not necessarily confirm gene expression, metabolite production, colonization ability, immune modulation, or safety under gut-like conditions. Therefore, the synthesis should be interpreted as a comparative genomic overview of potential traits, not as direct evidence of probiotic efficacy.

Research gaps

Future research should move beyond descriptive genome annotation toward standardized, experimentally validated strain-level assessment. High-quality complete genomes are still needed to refine species boundaries, phylogroups, pangenome architecture, and lineage-specific functional markers in both genera. Functional validation should combine comparative genomics with transcriptomics, proteomics, and metabolomics, together with in vivo or clinical studies. Safety assessment also needs to be standardized across studies, including consistent screening for ARGs, virulence/pathogenicity-related genes, toxin-related annotations, phage elements, HGT signals, mobile genetic elements, and efflux-related genes. These gaps are important because promising next-generation probiotic candidates must be selected not only for beneficial genomic traits but also for demonstrated functionality, stability, and safety.

Conclusion

This systematic review demonstrates that comparative genomics is valuable for characterizing pangenome diversity, functional adaptation, and probiotic-associated traits in Akkermansia and Faecalibacterium. Both genera showed high genomic diversity, open pangenomes, and extensive strain- or lineage-specific gene repertoires, indicating that their next-generation probiotic potential should not be generalized at the genus level. Akkermansia was mainly associated with mucin, HMO, and host-glycan utilization, mucosal adaptation, oxygen tolerance, and genes associated with vitamin biosynthesis. In addition, Faecalibacterium was characterized by broader carbohydrate utilization, acetate- and butyrate-related metabolism, EPS/capsule genes, and strain-specific anti-inflammatory traits. However, many findings remain genome-inferred, and safety assessments were inconsistent across studies. Future research should combine comparative genomics with standardized safety screening, multi-omics validation, host-interaction assays, and controlled in vivo or clinical studies to support strain-level selection of next-generation probiotic candidates.

Ethics and consent

This systematic review synthesized evidence from previously published studies and did not directly involve human participants, animal experimentation, or primary biological sampling. Therefore, ethical approval and consent were not required.

Use of AI tools

ChatGPT (OpenAI, GPT-5.5 Thinking) was used for language refinement, editorial polishing, and formatting support during manuscript preparation. The authors reviewed and verified all content and take full responsibility for the final manuscript.

Comments on this article Comments (0)

Version 1
VERSION 1 PUBLISHED 25 Jul 2026
Comment
Author details Author details
Competing interests
Grant information
Copyright
Download
 
Export To
metrics
Views Downloads
F1000Research - -
PubMed Central
Data from PMC are received and updated monthly.
- -
Citations
CITE
how to cite this article
Arlan L, Lestari RD, Israyusnita F et al. Comparative Genomic Insights into Pangenome Diversity and Functional Adaptation of Akkermansia and Faecalibacterium as Gut-Associated Next-Generation Probiotics: A Systematic Review [version 1; peer review: awaiting peer review]. F1000Research 2026, 15:1222 (https://doi.org/10.12688/f1000research.186241.1)
NOTE: If applicable, it is important to ensure the information in square brackets after the title is included in all citations of this article.
track
receive updates on this article
Track an article to receive email alerts on any updates to this article.

Open Peer Review

Current Reviewer Status:
AWAITING PEER REVIEW
AWAITING PEER REVIEW
?
Key to Reviewer Statuses VIEW
ApprovedThe paper is scientifically sound in its current form and only minor, if any, improvements are suggested
Approved with reservations A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit.
Not approvedFundamental flaws in the paper seriously undermine the findings and conclusions

Comments on this article Comments (0)

Version 1
VERSION 1 PUBLISHED 25 Jul 2026
Comment
Alongside their report, reviewers assign a status to the article:
Approved - the paper is scientifically sound in its current form and only minor, if any, improvements are suggested
Approved with reservations - A number of small changes, sometimes more significant revisions are required to address specific details and improve the papers academic merit.
Not approved - fundamental flaws in the paper seriously undermine the findings and conclusions
Sign In
If you've forgotten your password, please enter your email address below and we'll send you instructions on how to reset your password.

The email address should be the one you originally registered with F1000.

Email address not valid, please try again

You registered with F1000 via Google, so we cannot reset your password.

To sign in, please click here.

If you still need help with your Google account password, please click here.

You registered with F1000 via Facebook, so we cannot reset your password.

To sign in, please click here.

If you still need help with your Facebook account password, please click here.

Code not correct, please try again
Email us for further assistance.
Server error, please try again.