Pub Date : 2026-08-13eCollection Date: 2026-01-01DOI: 10.3389/fbinf.2026.1926188
Xin Li, Yaoyu Liu, Ming Xu
Introduction: MicroRNAs (miRNAs) regulate gene expression and are closely linked to the onset and progression of immune‑related diseases. Experimental discovery of disease‑associated miRNAs remains costly and time‑consuming, motivating computational prioritization. However, existing matrix‑completion and deep graph‑learning methods either underuse local biological neighborhood evidence or require complex multi‑view neural architectures.
Methods: In this study, we present RLNSF‑MDA, a reliability‑guided graph‑regularized logistic matrix factorization framework for predicting potential miRNA-disease associations, with an emphasis on immune disease analysis. The method integrates disease semantic similarity; miRNA functional, semantic, and sequence similarities; and training‑only Gaussian interaction profile similarities through data‑driven reliability weights. It then combines multi‑scale neighborhood evidence, diffusion scores, low‑rank reconstruction features, and contrastive graph‑regularized latent factors.
Results: On the HMDD v3.2 benchmark containing 788 miRNAs and 374 diseases, RLNSF‑MDA achieved an average accuracy of 0.8525, an AUC of 0.9194, and an AUPR of 0.9055 in five‑fold cross‑validation, and an average accuracy of 0.8569, an AUC of 0.9266, and an AUPR of 0.9197 in ten‑fold cross‑validation. Ablation experiments showed that the full model outperformed reduced feature combinations and training strategies, supporting the contribution of reliability‑guided fusion, side‑score construction, graph regularization, and contrastive ranking. Evidence from immune‑related case studies in dbDEMC V2.0 and miR2Disease further covered lupus nephritis, lymphoma, and leukemia, with 45, 47, and 47 confirmed miRNAs, respectively.
Discussion: These results suggest that RLNSF‑MDA provides an effective and interpretable framework for prioritizing candidate miRNAs associated with immune‑related diseases.
{"title":"RLNSF-MDA: reliability-guided graph-regularized matrix factorization for immune-related miRNA-disease association prediction.","authors":"Xin Li, Yaoyu Liu, Ming Xu","doi":"10.3389/fbinf.2026.1926188","DOIUrl":"10.3389/fbinf.2026.1926188","url":null,"abstract":"<p><strong>Introduction: </strong>MicroRNAs (miRNAs) regulate gene expression and are closely linked to the onset and progression of immune‑related diseases. Experimental discovery of disease‑associated miRNAs remains costly and time‑consuming, motivating computational prioritization. However, existing matrix‑completion and deep graph‑learning methods either underuse local biological neighborhood evidence or require complex multi‑view neural architectures.</p><p><strong>Methods: </strong>In this study, we present RLNSF‑MDA, a reliability‑guided graph‑regularized logistic matrix factorization framework for predicting potential miRNA-disease associations, with an emphasis on immune disease analysis. The method integrates disease semantic similarity; miRNA functional, semantic, and sequence similarities; and training‑only Gaussian interaction profile similarities through data‑driven reliability weights. It then combines multi‑scale neighborhood evidence, diffusion scores, low‑rank reconstruction features, and contrastive graph‑regularized latent factors.</p><p><strong>Results: </strong>On the HMDD v3.2 benchmark containing 788 miRNAs and 374 diseases, RLNSF‑MDA achieved an average accuracy of 0.8525, an AUC of 0.9194, and an AUPR of 0.9055 in five‑fold cross‑validation, and an average accuracy of 0.8569, an AUC of 0.9266, and an AUPR of 0.9197 in ten‑fold cross‑validation. Ablation experiments showed that the full model outperformed reduced feature combinations and training strategies, supporting the contribution of reliability‑guided fusion, side‑score construction, graph regularization, and contrastive ranking. Evidence from immune‑related case studies in dbDEMC V2.0 and miR2Disease further covered lupus nephritis, lymphoma, and leukemia, with 45, 47, and 47 confirmed miRNAs, respectively.</p><p><strong>Discussion: </strong>These results suggest that RLNSF‑MDA provides an effective and interpretable framework for prioritizing candidate miRNAs associated with immune‑related diseases.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1926188"},"PeriodicalIF":3.6,"publicationDate":"2026-08-13","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13518331/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148842243","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-12eCollection Date: 2026-01-01DOI: 10.3389/fbinf.2026.1856011
Katie Farrell, Kevin Tang, Melchizedek Mashiku, Dawit Abay, Joseph Longo, Edward Ramos, Gabriel Leventhal-Douglas, Margaret Rohrbaugh, Paul Chenoweth, Cara C Burns, Kun Zhao
One approach to find possible poliovirus sources that have not been detected by the Global Polio Eradication Initiative's (GPEI) surveillance systems is to scan reads in the Sequence Read Archive (SRA) for potential poliovirus sequences. In the post-eradication era, the identification of poliovirus sequences in the SRA database could signal a potential biosafety risk which may set back the enormous achievements of the GPEI. To advance the use of the SRA database for detecting poliovirus, we hypothesized that bioinformatics alignment tools like Bowtie2, BLASTn, Magic-BLAST, MegaBLAST, STAT, and ElasticBLAST could distinguish between non-poliovirus and poliovirus reads from samples represented in the SRA database. Short poliovirus sequencing reads were simulated using poliovirus Sabin strain genomes. Simulation was also done for sequences other than poliovirus (referred here as "non-poliovirus reads"). Simulated reads were aligned to reference poliovirus genomes using different alignment tools to benchmark the accuracy and computing time of each tool. Parameters were also established to identify previously unknown poliovirus reads using percent identity and alignment length from BLASTn results. Bowtie2 was the most accurate and efficient tool, correctly identifying all simulated poliovirus reads. The STAT tool detected 99.6% of known control poliovirus accessions using the enterovirus query but only 77.6% using the poliovirus query, demonstrating strong but incomplete detection capability. This study demonstrates the feasibility of screening the SRA for poliovirus sequences as a tool to strengthen poliovirus containment and mitigate post-eradication risks, which can serve as an additional safety net for the GPEI.
{"title":"Detecting poliovirus sequences in the sequence read archive database using bioinformatics tools.","authors":"Katie Farrell, Kevin Tang, Melchizedek Mashiku, Dawit Abay, Joseph Longo, Edward Ramos, Gabriel Leventhal-Douglas, Margaret Rohrbaugh, Paul Chenoweth, Cara C Burns, Kun Zhao","doi":"10.3389/fbinf.2026.1856011","DOIUrl":"10.3389/fbinf.2026.1856011","url":null,"abstract":"<p><p>One approach to find possible poliovirus sources that have not been detected by the Global Polio Eradication Initiative's (GPEI) surveillance systems is to scan reads in the Sequence Read Archive (SRA) for potential poliovirus sequences. In the post-eradication era, the identification of poliovirus sequences in the SRA database could signal a potential biosafety risk which may set back the enormous achievements of the GPEI. To advance the use of the SRA database for detecting poliovirus, we hypothesized that bioinformatics alignment tools like Bowtie2, BLASTn, Magic-BLAST, MegaBLAST, STAT, and ElasticBLAST could distinguish between non-poliovirus and poliovirus reads from samples represented in the SRA database. Short poliovirus sequencing reads were simulated using poliovirus Sabin strain genomes. Simulation was also done for sequences other than poliovirus (referred here as \"non-poliovirus reads\"). Simulated reads were aligned to reference poliovirus genomes using different alignment tools to benchmark the accuracy and computing time of each tool. Parameters were also established to identify previously unknown poliovirus reads using percent identity and alignment length from BLASTn results. Bowtie2 was the most accurate and efficient tool, correctly identifying all simulated poliovirus reads. The STAT tool detected 99.6% of known control poliovirus accessions using the enterovirus query but only 77.6% using the poliovirus query, demonstrating strong but incomplete detection capability. This study demonstrates the feasibility of screening the SRA for poliovirus sequences as a tool to strengthen poliovirus containment and mitigate post-eradication risks, which can serve as an additional safety net for the GPEI.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1856011"},"PeriodicalIF":3.6,"publicationDate":"2026-08-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13506724/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148835412","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-12eCollection Date: 2026-01-01DOI: 10.3389/fbinf.2026.1903746
Fahim Sufi
Synthetic microbial genomic data are becoming increasingly important for benchmarking microbial genome analysis pipelines, simulating rare taxa, evaluating metagenomic workflows, and supporting reproducible computational biology. Recent genomic foundation models demonstrate that biological sequences can be modelled at unprecedented scale, with emerging capacity for genome-level interpretation, generation, and design. However, the scientific value of synthetic microbial genomic data depends not only on whether sequences can be generated, but whether they are biologically plausible, computationally useful, reproducible, and responsibly governed. This Perspective argues that agentic AI can provide the missing orchestration layer for trustworthy synthetic microbial genomics. Rather than treating synthetic data generation as a single model output, agentic workflows can coordinate specialised roles for sequence generation, biological plausibility assessment, taxonomic validation, functional annotation, contamination detection, downstream benchmarking, provenance logging, and governance review. I propose a validation-first agentic framework in which synthetic microbial genomes, plasmids, phages, and metagenomic profiles are iteratively generated, evaluated, revised, and documented before release or downstream use. Such a framework can help transform synthetic microbial genomic data from computational artefacts into auditable scientific infrastructure with explicit validation gates, escalation criteria, and machine-readable provenance.
{"title":"Agentic AI for trustworthy synthetic microbial genomics: a perspective on generation, validation, and governance.","authors":"Fahim Sufi","doi":"10.3389/fbinf.2026.1903746","DOIUrl":"10.3389/fbinf.2026.1903746","url":null,"abstract":"<p><p>Synthetic microbial genomic data are becoming increasingly important for benchmarking microbial genome analysis pipelines, simulating rare taxa, evaluating metagenomic workflows, and supporting reproducible computational biology. Recent genomic foundation models demonstrate that biological sequences can be modelled at unprecedented scale, with emerging capacity for genome-level interpretation, generation, and design. However, the scientific value of synthetic microbial genomic data depends not only on whether sequences can be generated, but whether they are biologically plausible, computationally useful, reproducible, and responsibly governed. This Perspective argues that agentic AI can provide the missing orchestration layer for trustworthy synthetic microbial genomics. Rather than treating synthetic data generation as a single model output, agentic workflows can coordinate specialised roles for sequence generation, biological plausibility assessment, taxonomic validation, functional annotation, contamination detection, downstream benchmarking, provenance logging, and governance review. I propose a validation-first agentic framework in which synthetic microbial genomes, plasmids, phages, and metagenomic profiles are iteratively generated, evaluated, revised, and documented before release or downstream use. Such a framework can help transform synthetic microbial genomic data from computational artefacts into auditable scientific infrastructure with explicit validation gates, escalation criteria, and machine-readable provenance.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1903746"},"PeriodicalIF":3.6,"publicationDate":"2026-08-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13506882/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148835394","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-12eCollection Date: 2026-01-01DOI: 10.3389/fbinf.2026.1896572
Gerald Mboowa, Ivan Sserwadda, Stephen Kanyerezi, Benson R Kidenya, Jonani Bwambale, Benson Musinguzi
<p><strong>Background: </strong><i>Candida</i> spp. infections are an increasing public health concern, particularly in settings where laboratory mycology, genomic surveillance infrastructure, and antifungal susceptibility testing remain limited. Accurate species identification, reproducible assembly assessment, conservative genomic screening for antifungal-resistance markers, and interpretable phylogenomic outputs are essential for surveillance and outbreak preparedness. However, fungal whole-genome sequencing workflows remain fragmented, difficult to reproduce across computing environments, and insufficiently adapted for implementation in low-resource public health genomics settings.</p><p><strong>Methods: </strong>We developed rMAP-Candida, a modular, Dockerized WDL/Cromwell workflow for paired-end <i>Candida</i> spp. whole-genome sequencing analysis. The workflow performs read quality control and trimming with fastp, Candida-focused species typing using Kraken2/Bracken, <i>de novo</i> assembly with MEGAHIT, assembly-contiguity assessment with QUAST, optional genome-completeness assessment using Compleasm or BUSCO, antifungal-resistance marker screening using ChroQueTas/FungAMR-derived outputs, species-aware core-SNP phylogenomics, pairwise SNP-distance summarization, closest-neighbor analysis, and integrated HTML surveillance reporting. To improve independent reproducibility, the repository includes a quick-start local Cromwell test, a corrected two-sample input JSON, documented checks for public container and database access, and a two-sample reproducibility run.</p><p><strong>Results: </strong>rMAP-Candida generated reproducible species assignments, assembly-contiguity metrics, optional completeness summaries, antifungal-resistance marker outputs, species-aware phylogenomic summaries, pairwise SNP-distance tables, closest-neighbor summaries, and integrated HTML reports. In the Ugandan validation dataset, the workflow identified six principal species groups, dominated by <i>Candida albicans</i>, followed by <i>Candida tropicalis</i>, <i>Pichia kudriavzevii</i>, <i>Nakaseomyces glabratus</i>, <i>Clavispora lusitaniae</i>, and <i>Candida parapsilosis</i>. The integrated report summarized 24 antifungal-resistance marker hits and a median assembly N50 of 35,946 bp. Species-aware phylogenomics was performed for eligible species groups, while ineligible or skipped groups were explicitly reported with reasons. The report also distinguished "no curated genomic antifungal-resistance marker detected" from phenotypic susceptibility, supporting conservative interpretation of resistance-screening outputs.</p><p><strong>Conclusion: </strong>rMAP-Candida provides a portable, reproducible, modular, and surveillance-oriented WDL/Cromwell workflow for <i>Candida</i> spp. genomic analysis. By integrating species identification, assembly-contiguity assessment, optional completeness evaluation, antifungal-resistance marker screening, species-aware phylogenomics,
{"title":"rMAP-Candida: a modular Dockerized WDL/Cromwell workflow for reproducible <i>Candida species</i> typing, assembly-contiguity assessment, antifungal-resistance marker screening, and phylogenomic surveillance.","authors":"Gerald Mboowa, Ivan Sserwadda, Stephen Kanyerezi, Benson R Kidenya, Jonani Bwambale, Benson Musinguzi","doi":"10.3389/fbinf.2026.1896572","DOIUrl":"10.3389/fbinf.2026.1896572","url":null,"abstract":"<p><strong>Background: </strong><i>Candida</i> spp. infections are an increasing public health concern, particularly in settings where laboratory mycology, genomic surveillance infrastructure, and antifungal susceptibility testing remain limited. Accurate species identification, reproducible assembly assessment, conservative genomic screening for antifungal-resistance markers, and interpretable phylogenomic outputs are essential for surveillance and outbreak preparedness. However, fungal whole-genome sequencing workflows remain fragmented, difficult to reproduce across computing environments, and insufficiently adapted for implementation in low-resource public health genomics settings.</p><p><strong>Methods: </strong>We developed rMAP-Candida, a modular, Dockerized WDL/Cromwell workflow for paired-end <i>Candida</i> spp. whole-genome sequencing analysis. The workflow performs read quality control and trimming with fastp, Candida-focused species typing using Kraken2/Bracken, <i>de novo</i> assembly with MEGAHIT, assembly-contiguity assessment with QUAST, optional genome-completeness assessment using Compleasm or BUSCO, antifungal-resistance marker screening using ChroQueTas/FungAMR-derived outputs, species-aware core-SNP phylogenomics, pairwise SNP-distance summarization, closest-neighbor analysis, and integrated HTML surveillance reporting. To improve independent reproducibility, the repository includes a quick-start local Cromwell test, a corrected two-sample input JSON, documented checks for public container and database access, and a two-sample reproducibility run.</p><p><strong>Results: </strong>rMAP-Candida generated reproducible species assignments, assembly-contiguity metrics, optional completeness summaries, antifungal-resistance marker outputs, species-aware phylogenomic summaries, pairwise SNP-distance tables, closest-neighbor summaries, and integrated HTML reports. In the Ugandan validation dataset, the workflow identified six principal species groups, dominated by <i>Candida albicans</i>, followed by <i>Candida tropicalis</i>, <i>Pichia kudriavzevii</i>, <i>Nakaseomyces glabratus</i>, <i>Clavispora lusitaniae</i>, and <i>Candida parapsilosis</i>. The integrated report summarized 24 antifungal-resistance marker hits and a median assembly N50 of 35,946 bp. Species-aware phylogenomics was performed for eligible species groups, while ineligible or skipped groups were explicitly reported with reasons. The report also distinguished \"no curated genomic antifungal-resistance marker detected\" from phenotypic susceptibility, supporting conservative interpretation of resistance-screening outputs.</p><p><strong>Conclusion: </strong>rMAP-Candida provides a portable, reproducible, modular, and surveillance-oriented WDL/Cromwell workflow for <i>Candida</i> spp. genomic analysis. By integrating species identification, assembly-contiguity assessment, optional completeness evaluation, antifungal-resistance marker screening, species-aware phylogenomics,","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1896572"},"PeriodicalIF":3.6,"publicationDate":"2026-08-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13506686/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148835417","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-12eCollection Date: 2026-01-01DOI: 10.3389/fbinf.2026.1871436
Jing Li, Juncai Li, Qijia Chen, Zhong Wang, Xianzhi Liu, Mingmin Liang, Junzhuang Wang, Hongyuan Ding, Bin Zeng, Lei Wang
Background: In recent years, with the diversification and expansion of drug research in the medical field, the widespread use of drugs, particularly antibiotics, has led to increased microbial resistance. Consequently, exploring potential associations between drugs and microbes has become critically important. However, traditional biological experiments are extremely expensive and time-consuming. Therefore, developing more effective computational models for predicting potential associations between microbes and drugs is both essential and challenging.
Results: We proposed GATROF, a hybrid heterogeneous graph-based framework for microbe-drug association prediction. In GATROF, by integrating multiple microbe-drug-disease similarity measures, we first constructed two distinct microbe-drug networks. In addition, based on different features of microbes and drugs, we further constructed two novel microbe-drug feature matrices. On this basis, the microbe-drug networks and the constructed feature matrices were further used in a Graph Attention Network to learn complementary topology-aware representations of microbes and drugs. These GAT-derived representations were then integrated with the constructed drug-side and microbe-side feature matrices and input into a Rotation Forest classifier for final association prediction. Experimental results and case studies demonstrated that GATROF predicts microbe-drug associations more accurately than existing state-of-the-art methods.
Conclusion: GATROF provides a new integrated predictive framework for predicting potential microbe-drug associations. By combining heterogeneous biological information, GAT-based topological representation learning, and Rotation Forest classification, GATROF may help prioritize candidate drug-microbe associations for further biological validation.
{"title":"A microbe-drug association prediction model based on graph attention networks and rotation forest.","authors":"Jing Li, Juncai Li, Qijia Chen, Zhong Wang, Xianzhi Liu, Mingmin Liang, Junzhuang Wang, Hongyuan Ding, Bin Zeng, Lei Wang","doi":"10.3389/fbinf.2026.1871436","DOIUrl":"10.3389/fbinf.2026.1871436","url":null,"abstract":"<p><strong>Background: </strong>In recent years, with the diversification and expansion of drug research in the medical field, the widespread use of drugs, particularly antibiotics, has led to increased microbial resistance. Consequently, exploring potential associations between drugs and microbes has become critically important. However, traditional biological experiments are extremely expensive and time-consuming. Therefore, developing more effective computational models for predicting potential associations between microbes and drugs is both essential and challenging.</p><p><strong>Results: </strong>We proposed GATROF, a hybrid heterogeneous graph-based framework for microbe-drug association prediction. In GATROF, by integrating multiple microbe-drug-disease similarity measures, we first constructed two distinct microbe-drug networks. In addition, based on different features of microbes and drugs, we further constructed two novel microbe-drug feature matrices. On this basis, the microbe-drug networks and the constructed feature matrices were further used in a Graph Attention Network to learn complementary topology-aware representations of microbes and drugs. These GAT-derived representations were then integrated with the constructed drug-side and microbe-side feature matrices and input into a Rotation Forest classifier for final association prediction. Experimental results and case studies demonstrated that GATROF predicts microbe-drug associations more accurately than existing state-of-the-art methods.</p><p><strong>Conclusion: </strong>GATROF provides a new integrated predictive framework for predicting potential microbe-drug associations. By combining heterogeneous biological information, GAT-based topological representation learning, and Rotation Forest classification, GATROF may help prioritize candidate drug-microbe associations for further biological validation.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1871436"},"PeriodicalIF":3.6,"publicationDate":"2026-08-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13506770/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148835425","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-12eCollection Date: 2026-01-01DOI: 10.3389/fbinf.2026.1833416
Sajid Khan, Sabahat Jamil, Muhammad Hamza, Zarlish Attique, Suping Zhang
Metastatic breast cancer remains difficult to cure, and the way B and T lymphocytes adapt across metastatic niches especially under therapy remains insufficiently defined. Clarifying compartment specific immune remodeling may help explain resistance to PD-1/PD-L1 blockade and identify actionable targets. We performed an integrated meta-analysis of single cell RNA-seq datasets from normal breast tissue, primary tumors, tumor-draining lymph nodes (TLNs), and peripheral blood mononuclear cells (PBMCs), focusing on B and Tcell states. Immune composition differed notably by compartment. Tumors were enriched for effector CD8 states (CD8 cytotoxic 20.1%; CD8 activated 13.5%), whereas TLNs preserved larger naïve and memory reservoirs (CD4 naïve 40.7%; B naïve 11.4%; B memory 12.0%) and contained a higher B cell fraction than tumors (39.6% vs. 19.5%). Post therapy, PBMCs and TLNs showed increased BTLA-HVEM (TNFRSF14) checkpoint signaling and enhanced MIF-CD74 interactions with a shift from CD44 toward CXCR4, consistent with CXCR4 driven migratory and survival programs. In TLNs, TNFRSF14 signaling was unidirectional (B→T), absent in the reverse direction, and not detected in tumors. Clinically, higher tumor CXCR4 combined with lower TNFRSF14 was associated with shorter progression free survival in TCGA-BRCA, most evident in node positive, early stage disease. To target the BTLA-HVEM checkpoint axis, we performed structure guided de novo peptide design using the native HVEM (23-39) peptide as an active structural template, followed by docking and molecular dynamics simulations. The optimized De novo-P2 peptide showed stable and favorable interactions at the BTLA interface, supporting its potential as a competitive modulator of BTLA-HVEM signaling. These data define niche specific lymphocyte remodeling and implicate BTLA-HVEM and CXCL12-CXCR4 as candidate biomarkers and therapeutic targets linked to PD-1/PD-L1 resistance.
转移性乳腺癌仍然难以治愈,B淋巴细胞和T淋巴细胞适应转移性壁龛的方式,特别是在治疗下,仍然没有充分的定义。阐明室特异性免疫重构可能有助于解释PD-1/PD-L1阻断的耐药性,并确定可行的靶点。我们对来自正常乳腺组织、原发肿瘤、肿瘤引流淋巴结(tln)和外周血单核细胞(PBMCs)的单细胞RNA-seq数据集进行了综合荟萃分析,重点关注B细胞和t细胞状态。不同细胞间的免疫组成差异显著。肿瘤富集了CD8效应态(CD8细胞毒性20.1%,CD8活化13.5%),而tln保留了更大的naïve和记忆库(CD4 naïve 40.7%; B naïve 11.4%; B记忆12.0%),并且比肿瘤含有更高的B细胞比例(39.6%比19.5%)。治疗后,pbmc和tln显示BTLA-HVEM (TNFRSF14)检查点信号增加,MIF-CD74相互作用增强,从CD44向CXCR4转移,与CXCR4驱动的迁移和生存计划一致。在tln中,TNFRSF14信号是单向的(B→T),相反方向不存在,在肿瘤中未检测到。在临床上,TCGA-BRCA中,较高的肿瘤CXCR4合并较低的TNFRSF14与较短的无进展生存期相关,在淋巴结阳性的早期疾病中最为明显。为了靶向BTLA-HVEM检查点轴,我们以天然HVEM(23-39)肽作为活性结构模板,进行了结构引导的从头肽设计,然后进行对接和分子动力学模拟。优化后的De novo-P2肽在BTLA界面上表现出稳定和良好的相互作用,支持其作为BTLA- hvem信号传导的竞争性调节剂的潜力。这些数据定义了小生境特异性淋巴细胞重塑,并暗示BTLA-HVEM和CXCL12-CXCR4是与PD-1/PD-L1耐药相关的候选生物标志物和治疗靶点。
{"title":"Single cell mapping of B and T cell dynamics in breast cancer lymph node metastasis.","authors":"Sajid Khan, Sabahat Jamil, Muhammad Hamza, Zarlish Attique, Suping Zhang","doi":"10.3389/fbinf.2026.1833416","DOIUrl":"10.3389/fbinf.2026.1833416","url":null,"abstract":"<p><p>Metastatic breast cancer remains difficult to cure, and the way B and T lymphocytes adapt across metastatic niches especially under therapy remains insufficiently defined. Clarifying compartment specific immune remodeling may help explain resistance to PD-1/PD-L1 blockade and identify actionable targets. We performed an integrated meta-analysis of single cell RNA-seq datasets from normal breast tissue, primary tumors, tumor-draining lymph nodes (TLNs), and peripheral blood mononuclear cells (PBMCs), focusing on B and Tcell states. Immune composition differed notably by compartment. Tumors were enriched for effector CD8 states (CD8 cytotoxic 20.1%; CD8 activated 13.5%), whereas TLNs preserved larger naïve and memory reservoirs (CD4 naïve 40.7%; B naïve 11.4%; B memory 12.0%) and contained a higher B cell fraction than tumors (39.6% vs. 19.5%). Post therapy, PBMCs and TLNs showed increased BTLA-HVEM (TNFRSF14) checkpoint signaling and enhanced MIF-CD74 interactions with a shift from CD44 toward CXCR4, consistent with CXCR4 driven migratory and survival programs. In TLNs, TNFRSF14 signaling was unidirectional (B→T), absent in the reverse direction, and not detected in tumors. Clinically, higher tumor CXCR4 combined with lower TNFRSF14 was associated with shorter progression free survival in TCGA-BRCA, most evident in node positive, early stage disease. To target the BTLA-HVEM checkpoint axis, we performed structure guided <i>de novo</i> peptide design using the native HVEM (23-39) peptide as an active structural template, followed by docking and molecular dynamics simulations. The optimized De novo-P2 peptide showed stable and favorable interactions at the BTLA interface, supporting its potential as a competitive modulator of BTLA-HVEM signaling. These data define niche specific lymphocyte remodeling and implicate BTLA-HVEM and CXCL12-CXCR4 as candidate biomarkers and therapeutic targets linked to PD-1/PD-L1 resistance.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1833416"},"PeriodicalIF":3.6,"publicationDate":"2026-08-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13506802/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148835433","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-11eCollection Date: 2026-01-01DOI: 10.3389/fbinf.2026.1953979
Andrea Angarita-Rodríguez, Viviana Vargas-López, Andrés Pinzón, Adrián Sandoval-Hernandez, Kai Kang, Leping Li, Jason Papin, Pedro Puentes-Rozo, Andrés Felipe Aristizábal, Janneth González
[This corrects the article DOI: 10.3389/fbinf.2026.1816121.].
[这更正了文章DOI: 10.3389/fbinf.2026.1816121.]。
{"title":"Correction: Transcriptome-informed metabolic modeling reveals astrocyte-specific vulnerabilities in mild cognitive impairment and Alzheimer's disease progression.","authors":"Andrea Angarita-Rodríguez, Viviana Vargas-López, Andrés Pinzón, Adrián Sandoval-Hernandez, Kai Kang, Leping Li, Jason Papin, Pedro Puentes-Rozo, Andrés Felipe Aristizábal, Janneth González","doi":"10.3389/fbinf.2026.1953979","DOIUrl":"https://doi.org/10.3389/fbinf.2026.1953979","url":null,"abstract":"<p><p>[This corrects the article DOI: 10.3389/fbinf.2026.1816121.].</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1953979"},"PeriodicalIF":3.6,"publicationDate":"2026-08-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13504194/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148820441","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-11eCollection Date: 2026-01-01DOI: 10.3389/fbinf.2026.1873262
Abhirami Pv, Gupta Dheeraj Rajesh, N V L Sirisha Mulukuri, Ranjitha A, Niyas Rehman, Shiv Basant Kumar, Dileep Kumar, Pankaj Kumar
Introduction: Diabetes mellitus is a chronic metabolic disorder characterised by persistent hyperglycemia resulting from impaired insulin secretion, action, or both. Prolonged hyperglycemia disrupts metabolic homeostasis and increases the risk of complications, including cardiovascular disease, neuropathy, nephropathy, and retinopathy. Current therapies target multiple pathways but are limited by reduced efficacy and adverse effects. Peroxisome proliferator-activated receptor gamma (PPARγ) is a key regulator of glucose and lipid metabolism and an important therapeutic target. Thiazole and isoxazole scaffolds possess antidiabetic potential, and their hybridisation may improve efficacy and safety.
Objective: To design and evaluate a novel series of thiazole-linked isoxazole derivatives targeting the alternate ligand-binding domain of PPARγ for potential antidiabetic activity.
Result: The designed compounds (C1-C210) were docked into PPARγ (Protein Data Bank: 5 GTN), yielding nine candidates with binding affinities of -11.11 to -9.97 kcal/mol. Reference ligands included pioglitazone, Q-35, and S-35, with C139 showing the highest affinity. Molecular dynamics simulations (200 ns) confirmed stability. Advanced Trajectory analyses indicate favourable conformational sampling, correlated residue motions, and a relatively confined conformational space. In silico drug-likeness profiling further supported their activity profile.
Conclusion: In this computational study, the thiazole-linked isoxazole compound C139 demonstrated probable antidiabetic activity against PPARγ. However, these findings remain predictive and require further validation through comprehensive in vitro and in vivo studies.
{"title":"Design and <i>in silico</i> evaluation of Thiazole-Isoxazole hybrids as PPARγ-Targeted antidiabetic agents.","authors":"Abhirami Pv, Gupta Dheeraj Rajesh, N V L Sirisha Mulukuri, Ranjitha A, Niyas Rehman, Shiv Basant Kumar, Dileep Kumar, Pankaj Kumar","doi":"10.3389/fbinf.2026.1873262","DOIUrl":"10.3389/fbinf.2026.1873262","url":null,"abstract":"<p><strong>Introduction: </strong>Diabetes mellitus is a chronic metabolic disorder characterised by persistent hyperglycemia resulting from impaired insulin secretion, action, or both. Prolonged hyperglycemia disrupts metabolic homeostasis and increases the risk of complications, including cardiovascular disease, neuropathy, nephropathy, and retinopathy. Current therapies target multiple pathways but are limited by reduced efficacy and adverse effects. Peroxisome proliferator-activated receptor gamma (PPARγ) is a key regulator of glucose and lipid metabolism and an important therapeutic target. Thiazole and isoxazole scaffolds possess antidiabetic potential, and their hybridisation may improve efficacy and safety.</p><p><strong>Objective: </strong>To design and evaluate a novel series of thiazole-linked isoxazole derivatives targeting the alternate ligand-binding domain of PPARγ for potential antidiabetic activity.</p><p><strong>Result: </strong>The designed compounds (C1-C210) were docked into PPARγ (Protein Data Bank: 5 GTN), yielding nine candidates with binding affinities of -11.11 to -9.97 kcal/mol. Reference ligands included pioglitazone, Q-35, and S-35, with C139 showing the highest affinity. Molecular dynamics simulations (200 ns) confirmed stability. Advanced Trajectory analyses indicate favourable conformational sampling, correlated residue motions, and a relatively confined conformational space. <i>In silico</i> drug-likeness profiling further supported their activity profile.</p><p><strong>Conclusion: </strong>In this computational study, the thiazole-linked isoxazole compound C139 demonstrated probable antidiabetic activity against PPARγ. However, these findings remain predictive and require further validation through comprehensive <i>in vitro</i> and <i>in vivo</i> studies.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1873262"},"PeriodicalIF":3.6,"publicationDate":"2026-08-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13503372/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148820449","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Background: Environmental microbiome research has advanced through three complementary sequencing modalities, targeted 16S rRNA amplicon sequencing, whole-genome shotgun (WGS) metagenomics, and long-read full-length 16S rRNA profiling, each supported by distinct toolsets with heterogeneous outputs, variable configurations, and different levels of reproducibility documentation. Existing pipelines are typically modality-specific, require substantial configuration expertise, or produce outputs that need further custom scripting before standard ecological analyses can begin. This analytical fragmentation introduces avoidable technical variability and complicates cross-study reproducibility and comparability. PUDU addresses this by integrating all three modalities into a single reproducible workflow with simplified configuration, harmonized outputs across classifiers, and direct compatibility with downstream ecological analysis frameworks.
Results: We present PUDU (Pipeline for Universal Diversity Unveiling), a modular Snakemake workflow that supports amplicon (short-read 16S), shotgun metagenomics (WGS), and long-read 16S analyses from raw reads to standardized outputs for downstream microbial ecology. PUDU performs technology-aware preprocessing and centralized quality control, and integrates established taxonomic approaches, including DADA2 for amplicons, Emu for full-length 16S long reads, and Kraken2/Bracken and Centrifuger for WGS. Across methods, PUDU produces harmonized count and relative-abundance tables at user-defined taxonomic ranks, Krona files, and a standardized Phyloseq-compatible R object to streamline diversity analyses and statistical workflows. PUDU also provides an integrated Shiny interface for metadata-aware alpha/beta diversity, ordination, community composition, and shared-taxa exploration with exportable figures and taxa tables. We demonstrate PUDU on two publicly available environmental datasets spanning rhizosphere WGS and long-read marine sediment 16S, yielding broadly consistent community-level patterns across classifiers (Spearman ρ = 0.936 at phylum level; PERMANOVA R2 = 0.87-0.95) with peak memory below 45 GB on a standard Linux workstation.
Conclusion: PUDU is an end-to-end, reproducible, and extensible framework that enables standardized taxonomic profiling and ecology-oriented analysis across sequencing modalities. By combining harmonized outputs, Phyloseq interoperability, and an integrated visualization layer, PUDU facilitates reproducible, standardized, and comparable environmental microbiome analysis from raw reads to interpretable ecological insights.
{"title":"PUDU (pipeline for universal diversity unveiling): an accessible end-to-end workflow for taxonomic profiling and ecological visualization of environmental microbiomes across amplicon, shotgun, and long-read sequencing.","authors":"Alejandro Medaglia-Mata, Pablo Rojas-Rodríguez, Vojtěch Bystrý, Rossy Guillén-Watson, Olman Gómez-Espinoza, Kattia Núñez-Montero","doi":"10.3389/fbinf.2026.1909327","DOIUrl":"10.3389/fbinf.2026.1909327","url":null,"abstract":"<p><strong>Background: </strong>Environmental microbiome research has advanced through three complementary sequencing modalities, targeted 16S rRNA amplicon sequencing, whole-genome shotgun (WGS) metagenomics, and long-read full-length 16S rRNA profiling, each supported by distinct toolsets with heterogeneous outputs, variable configurations, and different levels of reproducibility documentation. Existing pipelines are typically modality-specific, require substantial configuration expertise, or produce outputs that need further custom scripting before standard ecological analyses can begin. This analytical fragmentation introduces avoidable technical variability and complicates cross-study reproducibility and comparability. PUDU addresses this by integrating all three modalities into a single reproducible workflow with simplified configuration, harmonized outputs across classifiers, and direct compatibility with downstream ecological analysis frameworks.</p><p><strong>Results: </strong>We present PUDU (Pipeline for Universal Diversity Unveiling), a modular Snakemake workflow that supports amplicon (short-read 16S), shotgun metagenomics (WGS), and long-read 16S analyses from raw reads to standardized outputs for downstream microbial ecology. PUDU performs technology-aware preprocessing and centralized quality control, and integrates established taxonomic approaches, including DADA2 for amplicons, Emu for full-length 16S long reads, and Kraken2/Bracken and Centrifuger for WGS. Across methods, PUDU produces harmonized count and relative-abundance tables at user-defined taxonomic ranks, Krona files, and a standardized Phyloseq-compatible R object to streamline diversity analyses and statistical workflows. PUDU also provides an integrated Shiny interface for metadata-aware alpha/beta diversity, ordination, community composition, and shared-taxa exploration with exportable figures and taxa tables. We demonstrate PUDU on two publicly available environmental datasets spanning rhizosphere WGS and long-read marine sediment 16S, yielding broadly consistent community-level patterns across classifiers (Spearman ρ = 0.936 at phylum level; PERMANOVA R<sup>2</sup> = 0.87-0.95) with peak memory below 45 GB on a standard Linux workstation.</p><p><strong>Conclusion: </strong>PUDU is an end-to-end, reproducible, and extensible framework that enables standardized taxonomic profiling and ecology-oriented analysis across sequencing modalities. By combining harmonized outputs, Phyloseq interoperability, and an integrated visualization layer, PUDU facilitates reproducible, standardized, and comparable environmental microbiome analysis from raw reads to interpretable ecological insights.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1909327"},"PeriodicalIF":3.6,"publicationDate":"2026-08-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13503585/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148820468","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-10eCollection Date: 2026-01-01DOI: 10.3389/fbinf.2026.1906036
Nan Li, Jing Hu, Wanning Tong, Chengdong Liu, Yun Ding, Ning Li, Zhigang Cai
Introduction: Bronchoalveolar lavage fluid metagenomic next-generation sequencing captures microbial profiles and host-derived molecular measurements from the same respiratory specimen, but downstream analysis requires coordinated handling of low-biomass microbial signals, negative-control information and multiple feature tables.
Methods: We developed LungMicroHostR, an R package for downstream host-microbiome analysis of bronchoalveolar lavage fluid metagenomic sequencing data. The package brings processed microbial profiles, host-derived molecular measurements, sample metadata and negative-control information into a unified R workflow for feature filtering, comparative model evaluation, visualization and reproducible reporting.
Results: Using the public GSE252118 resource comprising 402 samples from patients with lung cancer or pulmonary infections, LungMicroHostR assembled matched microbial, host and clinical feature tables, estimated prevalence in negative controls and compared host transcriptomic, microbial-profile and combined host-microbial models. In the test set, the 10-feature host transcriptome nearest-centroid model achieved an AUC of 0.772 (95% confidence interval, 0.680-0.860), the five-feature RNA microbial logistic model achieved an AUC of 0.745 (0.655-0.832), and the combined host transcriptome-RNA microbial logistic model achieved an AUC of 0.765 (0.655-0.866) with balanced accuracy of 0.720. An external PRJNA714488 BALF shotgun metagenomic dataset was additionally analysed at the mOTU level; LungMicroHostR matched the resulting feature table with phenotype metadata and generated a 388-feature by 26-sample microbial abundance matrix.
Discussion: LungMicroHostR provides documented functions for respiratory metagenomic analyses that require joint evaluation of microbial profiles, host-derived measurements, negative-control information and external microbial feature tables.
{"title":"LungMicroHostR: an R package for integrated host-microbiome analysis of bronchoalveolar lavage fluid metagenomic sequencing data.","authors":"Nan Li, Jing Hu, Wanning Tong, Chengdong Liu, Yun Ding, Ning Li, Zhigang Cai","doi":"10.3389/fbinf.2026.1906036","DOIUrl":"10.3389/fbinf.2026.1906036","url":null,"abstract":"<p><strong>Introduction: </strong>Bronchoalveolar lavage fluid metagenomic next-generation sequencing captures microbial profiles and host-derived molecular measurements from the same respiratory specimen, but downstream analysis requires coordinated handling of low-biomass microbial signals, negative-control information and multiple feature tables.</p><p><strong>Methods: </strong>We developed LungMicroHostR, an R package for downstream host-microbiome analysis of bronchoalveolar lavage fluid metagenomic sequencing data. The package brings processed microbial profiles, host-derived molecular measurements, sample metadata and negative-control information into a unified R workflow for feature filtering, comparative model evaluation, visualization and reproducible reporting.</p><p><strong>Results: </strong>Using the public GSE252118 resource comprising 402 samples from patients with lung cancer or pulmonary infections, LungMicroHostR assembled matched microbial, host and clinical feature tables, estimated prevalence in negative controls and compared host transcriptomic, microbial-profile and combined host-microbial models. In the test set, the 10-feature host transcriptome nearest-centroid model achieved an AUC of 0.772 (95% confidence interval, 0.680-0.860), the five-feature RNA microbial logistic model achieved an AUC of 0.745 (0.655-0.832), and the combined host transcriptome-RNA microbial logistic model achieved an AUC of 0.765 (0.655-0.866) with balanced accuracy of 0.720. An external PRJNA714488 BALF shotgun metagenomic dataset was additionally analysed at the mOTU level; LungMicroHostR matched the resulting feature table with phenotype metadata and generated a 388-feature by 26-sample microbial abundance matrix.</p><p><strong>Discussion: </strong>LungMicroHostR provides documented functions for respiratory metagenomic analyses that require joint evaluation of microbial profiles, host-derived measurements, negative-control information and external microbial feature tables.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1906036"},"PeriodicalIF":3.6,"publicationDate":"2026-08-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13500579/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148814970","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}