首页 > 最新文献

Frontiers in bioinformatics最新文献

英文 中文
RLNSF-MDA: reliability-guided graph-regularized matrix factorization for immune-related miRNA-disease association prediction. RLNSF-MDA:用于免疫相关mirna疾病关联预测的可靠性指导图正则化矩阵分解。
IF 3.6 Q2 MATHEMATICAL & COMPUTATIONAL BIOLOGY Pub Date : 2026-08-13 eCollection Date: 2026-01-01 DOI: 10.3389/fbinf.2026.1926188
Xin Li, Yaoyu Liu, Ming Xu

Introduction: MicroRNAs (miRNAs) regulate gene expression and are closely linked to the onset and progression of immune‑related diseases. Experimental discovery of disease‑associated miRNAs remains costly and time‑consuming, motivating computational prioritization. However, existing matrix‑completion and deep graph‑learning methods either underuse local biological neighborhood evidence or require complex multi‑view neural architectures.

Methods: In this study, we present RLNSF‑MDA, a reliability‑guided graph‑regularized logistic matrix factorization framework for predicting potential miRNA-disease associations, with an emphasis on immune disease analysis. The method integrates disease semantic similarity; miRNA functional, semantic, and sequence similarities; and training‑only Gaussian interaction profile similarities through data‑driven reliability weights. It then combines multi‑scale neighborhood evidence, diffusion scores, low‑rank reconstruction features, and contrastive graph‑regularized latent factors.

Results: On the HMDD v3.2 benchmark containing 788 miRNAs and 374 diseases, RLNSF‑MDA achieved an average accuracy of 0.8525, an AUC of 0.9194, and an AUPR of 0.9055 in five‑fold cross‑validation, and an average accuracy of 0.8569, an AUC of 0.9266, and an AUPR of 0.9197 in ten‑fold cross‑validation. Ablation experiments showed that the full model outperformed reduced feature combinations and training strategies, supporting the contribution of reliability‑guided fusion, side‑score construction, graph regularization, and contrastive ranking. Evidence from immune‑related case studies in dbDEMC V2.0 and miR2Disease further covered lupus nephritis, lymphoma, and leukemia, with 45, 47, and 47 confirmed miRNAs, respectively.

Discussion: These results suggest that RLNSF‑MDA provides an effective and interpretable framework for prioritizing candidate miRNAs associated with immune‑related diseases.

MicroRNAs (miRNAs)调节基因表达,与免疫相关疾病的发生和进展密切相关。疾病相关mirna的实验发现仍然是昂贵和耗时的,激励计算优先级。然而,现有的矩阵补全和深度图学习方法要么未充分利用局部生物邻域证据,要么需要复杂的多视图神经架构。方法:在本研究中,我们提出了RLNSF - MDA,这是一个可靠性指导的图正则化逻辑矩阵分解框架,用于预测潜在的mirna -疾病关联,重点是免疫疾病分析。该方法集成了疾病语义相似度;miRNA功能、语义和序列相似性;以及通过数据驱动的可靠性权重来训练仅高斯交互剖面的相似性。然后,它结合了多尺度邻域证据、扩散分数、低秩重建特征和对比图正则化潜在因素。结果:在包含788个mirna和374种疾病的HMDD v3.2基准上,RLNSF - MDA在5倍交叉验证中平均准确率为0.8525,AUC为0.9194,AUPR为0.9055;在10倍交叉验证中平均准确率为0.8569,AUC为0.9266,AUPR为0.9197。消融实验表明,完整模型优于简化特征组合和训练策略,支持可靠性引导融合、侧分数构建、图正则化和对比排序的贡献。来自dbDEMC V2.0和miR2Disease中免疫相关病例研究的证据进一步涵盖了狼疮肾炎、淋巴瘤和白血病,分别有45个、47个和47个确认的mirna。讨论:这些结果表明,RLNSF - MDA为优先考虑与免疫相关疾病相关的候选mirna提供了一个有效且可解释的框架。
{"title":"RLNSF-MDA: reliability-guided graph-regularized matrix factorization for immune-related miRNA-disease association prediction.","authors":"Xin Li, Yaoyu Liu, Ming Xu","doi":"10.3389/fbinf.2026.1926188","DOIUrl":"10.3389/fbinf.2026.1926188","url":null,"abstract":"<p><strong>Introduction: </strong>MicroRNAs (miRNAs) regulate gene expression and are closely linked to the onset and progression of immune‑related diseases. Experimental discovery of disease‑associated miRNAs remains costly and time‑consuming, motivating computational prioritization. However, existing matrix‑completion and deep graph‑learning methods either underuse local biological neighborhood evidence or require complex multi‑view neural architectures.</p><p><strong>Methods: </strong>In this study, we present RLNSF‑MDA, a reliability‑guided graph‑regularized logistic matrix factorization framework for predicting potential miRNA-disease associations, with an emphasis on immune disease analysis. The method integrates disease semantic similarity; miRNA functional, semantic, and sequence similarities; and training‑only Gaussian interaction profile similarities through data‑driven reliability weights. It then combines multi‑scale neighborhood evidence, diffusion scores, low‑rank reconstruction features, and contrastive graph‑regularized latent factors.</p><p><strong>Results: </strong>On the HMDD v3.2 benchmark containing 788 miRNAs and 374 diseases, RLNSF‑MDA achieved an average accuracy of 0.8525, an AUC of 0.9194, and an AUPR of 0.9055 in five‑fold cross‑validation, and an average accuracy of 0.8569, an AUC of 0.9266, and an AUPR of 0.9197 in ten‑fold cross‑validation. Ablation experiments showed that the full model outperformed reduced feature combinations and training strategies, supporting the contribution of reliability‑guided fusion, side‑score construction, graph regularization, and contrastive ranking. Evidence from immune‑related case studies in dbDEMC V2.0 and miR2Disease further covered lupus nephritis, lymphoma, and leukemia, with 45, 47, and 47 confirmed miRNAs, respectively.</p><p><strong>Discussion: </strong>These results suggest that RLNSF‑MDA provides an effective and interpretable framework for prioritizing candidate miRNAs associated with immune‑related diseases.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1926188"},"PeriodicalIF":3.6,"publicationDate":"2026-08-13","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13518331/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148842243","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Detecting poliovirus sequences in the sequence read archive database using bioinformatics tools. 利用生物信息学工具检测序列读取档案数据库中的脊髓灰质炎病毒序列。
IF 3.6 Q2 MATHEMATICAL & COMPUTATIONAL BIOLOGY Pub Date : 2026-08-12 eCollection Date: 2026-01-01 DOI: 10.3389/fbinf.2026.1856011
Katie Farrell, Kevin Tang, Melchizedek Mashiku, Dawit Abay, Joseph Longo, Edward Ramos, Gabriel Leventhal-Douglas, Margaret Rohrbaugh, Paul Chenoweth, Cara C Burns, Kun Zhao

One approach to find possible poliovirus sources that have not been detected by the Global Polio Eradication Initiative's (GPEI) surveillance systems is to scan reads in the Sequence Read Archive (SRA) for potential poliovirus sequences. In the post-eradication era, the identification of poliovirus sequences in the SRA database could signal a potential biosafety risk which may set back the enormous achievements of the GPEI. To advance the use of the SRA database for detecting poliovirus, we hypothesized that bioinformatics alignment tools like Bowtie2, BLASTn, Magic-BLAST, MegaBLAST, STAT, and ElasticBLAST could distinguish between non-poliovirus and poliovirus reads from samples represented in the SRA database. Short poliovirus sequencing reads were simulated using poliovirus Sabin strain genomes. Simulation was also done for sequences other than poliovirus (referred here as "non-poliovirus reads"). Simulated reads were aligned to reference poliovirus genomes using different alignment tools to benchmark the accuracy and computing time of each tool. Parameters were also established to identify previously unknown poliovirus reads using percent identity and alignment length from BLASTn results. Bowtie2 was the most accurate and efficient tool, correctly identifying all simulated poliovirus reads. The STAT tool detected 99.6% of known control poliovirus accessions using the enterovirus query but only 77.6% using the poliovirus query, demonstrating strong but incomplete detection capability. This study demonstrates the feasibility of screening the SRA for poliovirus sequences as a tool to strengthen poliovirus containment and mitigate post-eradication risks, which can serve as an additional safety net for the GPEI.

寻找全球根除脊髓灰质炎行动(GPEI)监测系统未发现的可能脊髓灰质炎病毒来源的一种方法是扫描序列读取档案(SRA)中的读取序列,寻找潜在的脊髓灰质炎病毒序列。在根除脊髓灰质炎后时代,在SRA数据库中识别脊髓灰质炎病毒序列可能预示着潜在的生物安全风险,这可能使GPEI取得的巨大成就倒退。为了进一步利用SRA数据库检测脊髓灰质炎病毒,我们假设生物信息学比对工具如Bowtie2、BLASTn、Magic-BLAST、MegaBLAST、STAT和ElasticBLAST可以区分SRA数据库中样本的非脊髓灰质炎病毒和脊髓灰质炎病毒。利用脊髓灰质炎病毒Sabin株基因组模拟脊髓灰质炎病毒短序列读取。对脊髓灰质炎病毒以外的序列(此处称为“非脊髓灰质炎病毒序列”)也进行了模拟。使用不同的比对工具将模拟读数与参考脊髓灰质炎病毒基因组比对,以基准每种工具的准确性和计算时间。还建立了参数,利用BLASTn结果的百分比识别和比对长度来鉴定以前未知的脊髓灰质炎病毒reads。Bowtie2是最准确和有效的工具,正确识别所有模拟的脊髓灰质炎病毒。STAT工具使用肠道病毒查询检测到99.6%的已知对照脊髓灰质炎病毒,但使用脊髓灰质炎病毒查询仅检测到77.6%,显示出强大但不完整的检测能力。这项研究表明,筛选SRA脊髓灰质炎病毒序列作为加强脊髓灰质炎病毒遏制和减轻根除后风险的工具是可行的,这可以作为GPEI的额外安全网。
{"title":"Detecting poliovirus sequences in the sequence read archive database using bioinformatics tools.","authors":"Katie Farrell, Kevin Tang, Melchizedek Mashiku, Dawit Abay, Joseph Longo, Edward Ramos, Gabriel Leventhal-Douglas, Margaret Rohrbaugh, Paul Chenoweth, Cara C Burns, Kun Zhao","doi":"10.3389/fbinf.2026.1856011","DOIUrl":"10.3389/fbinf.2026.1856011","url":null,"abstract":"<p><p>One approach to find possible poliovirus sources that have not been detected by the Global Polio Eradication Initiative's (GPEI) surveillance systems is to scan reads in the Sequence Read Archive (SRA) for potential poliovirus sequences. In the post-eradication era, the identification of poliovirus sequences in the SRA database could signal a potential biosafety risk which may set back the enormous achievements of the GPEI. To advance the use of the SRA database for detecting poliovirus, we hypothesized that bioinformatics alignment tools like Bowtie2, BLASTn, Magic-BLAST, MegaBLAST, STAT, and ElasticBLAST could distinguish between non-poliovirus and poliovirus reads from samples represented in the SRA database. Short poliovirus sequencing reads were simulated using poliovirus Sabin strain genomes. Simulation was also done for sequences other than poliovirus (referred here as \"non-poliovirus reads\"). Simulated reads were aligned to reference poliovirus genomes using different alignment tools to benchmark the accuracy and computing time of each tool. Parameters were also established to identify previously unknown poliovirus reads using percent identity and alignment length from BLASTn results. Bowtie2 was the most accurate and efficient tool, correctly identifying all simulated poliovirus reads. The STAT tool detected 99.6% of known control poliovirus accessions using the enterovirus query but only 77.6% using the poliovirus query, demonstrating strong but incomplete detection capability. This study demonstrates the feasibility of screening the SRA for poliovirus sequences as a tool to strengthen poliovirus containment and mitigate post-eradication risks, which can serve as an additional safety net for the GPEI.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1856011"},"PeriodicalIF":3.6,"publicationDate":"2026-08-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13506724/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148835412","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Agentic AI for trustworthy synthetic microbial genomics: a perspective on generation, validation, and governance. 可信合成微生物基因组学的人工智能:生成、验证和治理的视角。
IF 3.6 Q2 MATHEMATICAL & COMPUTATIONAL BIOLOGY Pub Date : 2026-08-12 eCollection Date: 2026-01-01 DOI: 10.3389/fbinf.2026.1903746
Fahim Sufi

Synthetic microbial genomic data are becoming increasingly important for benchmarking microbial genome analysis pipelines, simulating rare taxa, evaluating metagenomic workflows, and supporting reproducible computational biology. Recent genomic foundation models demonstrate that biological sequences can be modelled at unprecedented scale, with emerging capacity for genome-level interpretation, generation, and design. However, the scientific value of synthetic microbial genomic data depends not only on whether sequences can be generated, but whether they are biologically plausible, computationally useful, reproducible, and responsibly governed. This Perspective argues that agentic AI can provide the missing orchestration layer for trustworthy synthetic microbial genomics. Rather than treating synthetic data generation as a single model output, agentic workflows can coordinate specialised roles for sequence generation, biological plausibility assessment, taxonomic validation, functional annotation, contamination detection, downstream benchmarking, provenance logging, and governance review. I propose a validation-first agentic framework in which synthetic microbial genomes, plasmids, phages, and metagenomic profiles are iteratively generated, evaluated, revised, and documented before release or downstream use. Such a framework can help transform synthetic microbial genomic data from computational artefacts into auditable scientific infrastructure with explicit validation gates, escalation criteria, and machine-readable provenance.

合成微生物基因组数据对于微生物基因组分析管道的基准测试、模拟稀有分类群、评估宏基因组工作流程以及支持可重复的计算生物学变得越来越重要。最近的基因组基础模型表明,生物序列可以以前所未有的规模建模,具有基因组水平解释、生成和设计的新兴能力。然而,合成微生物基因组数据的科学价值不仅取决于是否可以生成序列,还取决于它们在生物学上是否合理、计算上是否有用、可重复性和负责任的管理。本观点认为,人工智能可以为值得信赖的合成微生物基因组学提供缺失的编排层。与将合成数据生成视为单一模型输出相比,代理工作流可以协调序列生成、生物合理性评估、分类验证、功能注释、污染检测、下游基准测试、来源记录和治理审查等专门角色。我提出了一个验证优先的代理框架,在该框架中,合成微生物基因组、质粒、噬菌体和宏基因组图谱在发布或下游使用之前迭代生成、评估、修订和记录。这样的框架可以帮助将合成的微生物基因组数据从计算工件转换为具有显式验证门、升级标准和机器可读来源的可审计的科学基础设施。
{"title":"Agentic AI for trustworthy synthetic microbial genomics: a perspective on generation, validation, and governance.","authors":"Fahim Sufi","doi":"10.3389/fbinf.2026.1903746","DOIUrl":"10.3389/fbinf.2026.1903746","url":null,"abstract":"<p><p>Synthetic microbial genomic data are becoming increasingly important for benchmarking microbial genome analysis pipelines, simulating rare taxa, evaluating metagenomic workflows, and supporting reproducible computational biology. Recent genomic foundation models demonstrate that biological sequences can be modelled at unprecedented scale, with emerging capacity for genome-level interpretation, generation, and design. However, the scientific value of synthetic microbial genomic data depends not only on whether sequences can be generated, but whether they are biologically plausible, computationally useful, reproducible, and responsibly governed. This Perspective argues that agentic AI can provide the missing orchestration layer for trustworthy synthetic microbial genomics. Rather than treating synthetic data generation as a single model output, agentic workflows can coordinate specialised roles for sequence generation, biological plausibility assessment, taxonomic validation, functional annotation, contamination detection, downstream benchmarking, provenance logging, and governance review. I propose a validation-first agentic framework in which synthetic microbial genomes, plasmids, phages, and metagenomic profiles are iteratively generated, evaluated, revised, and documented before release or downstream use. Such a framework can help transform synthetic microbial genomic data from computational artefacts into auditable scientific infrastructure with explicit validation gates, escalation criteria, and machine-readable provenance.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1903746"},"PeriodicalIF":3.6,"publicationDate":"2026-08-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13506882/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148835394","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
rMAP-Candida: a modular Dockerized WDL/Cromwell workflow for reproducible Candida species typing, assembly-contiguity assessment, antifungal-resistance marker screening, and phylogenomic surveillance. rmap -念珠菌:模块化Dockerized WDL/Cromwell工作流程可重复念珠菌种类分型,装配-连续性评估,抗真菌抗性标记筛选和系统基因组监测。
IF 3.6 Q2 MATHEMATICAL & COMPUTATIONAL BIOLOGY Pub Date : 2026-08-12 eCollection Date: 2026-01-01 DOI: 10.3389/fbinf.2026.1896572
Gerald Mboowa, Ivan Sserwadda, Stephen Kanyerezi, Benson R Kidenya, Jonani Bwambale, Benson Musinguzi
<p><strong>Background: </strong><i>Candida</i> spp. infections are an increasing public health concern, particularly in settings where laboratory mycology, genomic surveillance infrastructure, and antifungal susceptibility testing remain limited. Accurate species identification, reproducible assembly assessment, conservative genomic screening for antifungal-resistance markers, and interpretable phylogenomic outputs are essential for surveillance and outbreak preparedness. However, fungal whole-genome sequencing workflows remain fragmented, difficult to reproduce across computing environments, and insufficiently adapted for implementation in low-resource public health genomics settings.</p><p><strong>Methods: </strong>We developed rMAP-Candida, a modular, Dockerized WDL/Cromwell workflow for paired-end <i>Candida</i> spp. whole-genome sequencing analysis. The workflow performs read quality control and trimming with fastp, Candida-focused species typing using Kraken2/Bracken, <i>de novo</i> assembly with MEGAHIT, assembly-contiguity assessment with QUAST, optional genome-completeness assessment using Compleasm or BUSCO, antifungal-resistance marker screening using ChroQueTas/FungAMR-derived outputs, species-aware core-SNP phylogenomics, pairwise SNP-distance summarization, closest-neighbor analysis, and integrated HTML surveillance reporting. To improve independent reproducibility, the repository includes a quick-start local Cromwell test, a corrected two-sample input JSON, documented checks for public container and database access, and a two-sample reproducibility run.</p><p><strong>Results: </strong>rMAP-Candida generated reproducible species assignments, assembly-contiguity metrics, optional completeness summaries, antifungal-resistance marker outputs, species-aware phylogenomic summaries, pairwise SNP-distance tables, closest-neighbor summaries, and integrated HTML reports. In the Ugandan validation dataset, the workflow identified six principal species groups, dominated by <i>Candida albicans</i>, followed by <i>Candida tropicalis</i>, <i>Pichia kudriavzevii</i>, <i>Nakaseomyces glabratus</i>, <i>Clavispora lusitaniae</i>, and <i>Candida parapsilosis</i>. The integrated report summarized 24 antifungal-resistance marker hits and a median assembly N50 of 35,946 bp. Species-aware phylogenomics was performed for eligible species groups, while ineligible or skipped groups were explicitly reported with reasons. The report also distinguished "no curated genomic antifungal-resistance marker detected" from phenotypic susceptibility, supporting conservative interpretation of resistance-screening outputs.</p><p><strong>Conclusion: </strong>rMAP-Candida provides a portable, reproducible, modular, and surveillance-oriented WDL/Cromwell workflow for <i>Candida</i> spp. genomic analysis. By integrating species identification, assembly-contiguity assessment, optional completeness evaluation, antifungal-resistance marker screening, species-aware phylogenomics,
背景:念珠菌感染是一个日益严重的公共卫生问题,特别是在实验室真菌学、基因组监测基础设施和抗真菌药敏检测仍然有限的环境中。准确的物种鉴定、可重复的组装评估、抗真菌耐药性标记的保守基因组筛选以及可解释的系统基因组输出对于监测和疫情准备至关重要。然而,真菌全基因组测序工作流程仍然是碎片化的,难以跨计算环境重现,并且不足以适应在资源匮乏的公共卫生基因组学环境中实施。方法:我们开发了rMAP-Candida,一个模块化的Dockerized WDL/Cromwell工作流程,用于假丝酵母菌对端全基因组测序分析。该工作流程使用Kraken2/Bracken快速进行以念菌为重点的物种分型,使用MEGAHIT进行重新组装,使用QUAST进行组装连续性评估,使用Compleasm或BUSCO进行可选的基因组完整性评估,使用ChroQueTas/ fungamr衍生的输出进行抗真菌抗性标记筛选,使用物种感知核心snp系统基因组学,成对snp距离总结,最近邻分析和集成HTML监测报告进行阅读质量控制和修剪。为了提高独立的可重复性,存储库包括一个快速启动的本地克伦威尔测试,一个修正的双样本输入JSON,对公共容器和数据库访问的文档检查,以及一个双样本可重复性运行。结果:rMAP-Candida生成了可重复的物种分配、装配-邻近度量、可选的完整性摘要、抗真菌抗性标记输出、物种感知的系统基因组摘要、成对snp距离表、最近邻摘要和集成的HTML报告。在乌干达验证数据集中,该工作流程确定了6个主要物种群,以白色念珠菌为主,其次是热带念珠菌、kudriavzevii毕赤酵母、光秃中丝酵母、卢西塔尼亚Clavispora lusitaniae和副假丝酵母。该综合报告总结了24个抗真菌抗性标记命中,中位组装N50为35,946 bp。对符合条件的物种组进行物种意识系统基因组学,而不符合条件或跳过的组则明确报告并说明原因。该报告还将“未检测到精心设计的基因组抗真菌抗性标记”与表型易感性区分开来,支持对抗性筛选结果的保守解释。结论:rMAP-Candida为念珠菌基因组分析提供了一种便携式、可复制、模块化和面向监测的WDL/Cromwell工作流程。通过整合物种鉴定、装配-邻近性评估、可选完整性评估、抗真菌抗性标记筛选、物种感知系统基因组学、snp距离总结、最近邻报告和HTML报告,该工作流程支持培训、研究和在低资源和其他实施环境中应用真菌基因组监测。
{"title":"rMAP-Candida: a modular Dockerized WDL/Cromwell workflow for reproducible <i>Candida species</i> typing, assembly-contiguity assessment, antifungal-resistance marker screening, and phylogenomic surveillance.","authors":"Gerald Mboowa, Ivan Sserwadda, Stephen Kanyerezi, Benson R Kidenya, Jonani Bwambale, Benson Musinguzi","doi":"10.3389/fbinf.2026.1896572","DOIUrl":"10.3389/fbinf.2026.1896572","url":null,"abstract":"&lt;p&gt;&lt;strong&gt;Background: &lt;/strong&gt;&lt;i&gt;Candida&lt;/i&gt; spp. infections are an increasing public health concern, particularly in settings where laboratory mycology, genomic surveillance infrastructure, and antifungal susceptibility testing remain limited. Accurate species identification, reproducible assembly assessment, conservative genomic screening for antifungal-resistance markers, and interpretable phylogenomic outputs are essential for surveillance and outbreak preparedness. However, fungal whole-genome sequencing workflows remain fragmented, difficult to reproduce across computing environments, and insufficiently adapted for implementation in low-resource public health genomics settings.&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Methods: &lt;/strong&gt;We developed rMAP-Candida, a modular, Dockerized WDL/Cromwell workflow for paired-end &lt;i&gt;Candida&lt;/i&gt; spp. whole-genome sequencing analysis. The workflow performs read quality control and trimming with fastp, Candida-focused species typing using Kraken2/Bracken, &lt;i&gt;de novo&lt;/i&gt; assembly with MEGAHIT, assembly-contiguity assessment with QUAST, optional genome-completeness assessment using Compleasm or BUSCO, antifungal-resistance marker screening using ChroQueTas/FungAMR-derived outputs, species-aware core-SNP phylogenomics, pairwise SNP-distance summarization, closest-neighbor analysis, and integrated HTML surveillance reporting. To improve independent reproducibility, the repository includes a quick-start local Cromwell test, a corrected two-sample input JSON, documented checks for public container and database access, and a two-sample reproducibility run.&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Results: &lt;/strong&gt;rMAP-Candida generated reproducible species assignments, assembly-contiguity metrics, optional completeness summaries, antifungal-resistance marker outputs, species-aware phylogenomic summaries, pairwise SNP-distance tables, closest-neighbor summaries, and integrated HTML reports. In the Ugandan validation dataset, the workflow identified six principal species groups, dominated by &lt;i&gt;Candida albicans&lt;/i&gt;, followed by &lt;i&gt;Candida tropicalis&lt;/i&gt;, &lt;i&gt;Pichia kudriavzevii&lt;/i&gt;, &lt;i&gt;Nakaseomyces glabratus&lt;/i&gt;, &lt;i&gt;Clavispora lusitaniae&lt;/i&gt;, and &lt;i&gt;Candida parapsilosis&lt;/i&gt;. The integrated report summarized 24 antifungal-resistance marker hits and a median assembly N50 of 35,946 bp. Species-aware phylogenomics was performed for eligible species groups, while ineligible or skipped groups were explicitly reported with reasons. The report also distinguished \"no curated genomic antifungal-resistance marker detected\" from phenotypic susceptibility, supporting conservative interpretation of resistance-screening outputs.&lt;/p&gt;&lt;p&gt;&lt;strong&gt;Conclusion: &lt;/strong&gt;rMAP-Candida provides a portable, reproducible, modular, and surveillance-oriented WDL/Cromwell workflow for &lt;i&gt;Candida&lt;/i&gt; spp. genomic analysis. By integrating species identification, assembly-contiguity assessment, optional completeness evaluation, antifungal-resistance marker screening, species-aware phylogenomics,","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1896572"},"PeriodicalIF":3.6,"publicationDate":"2026-08-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13506686/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148835417","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
A microbe-drug association prediction model based on graph attention networks and rotation forest. 基于图关注网络和旋转森林的微生物-药物关联预测模型。
IF 3.6 Q2 MATHEMATICAL & COMPUTATIONAL BIOLOGY Pub Date : 2026-08-12 eCollection Date: 2026-01-01 DOI: 10.3389/fbinf.2026.1871436
Jing Li, Juncai Li, Qijia Chen, Zhong Wang, Xianzhi Liu, Mingmin Liang, Junzhuang Wang, Hongyuan Ding, Bin Zeng, Lei Wang

Background: In recent years, with the diversification and expansion of drug research in the medical field, the widespread use of drugs, particularly antibiotics, has led to increased microbial resistance. Consequently, exploring potential associations between drugs and microbes has become critically important. However, traditional biological experiments are extremely expensive and time-consuming. Therefore, developing more effective computational models for predicting potential associations between microbes and drugs is both essential and challenging.

Results: We proposed GATROF, a hybrid heterogeneous graph-based framework for microbe-drug association prediction. In GATROF, by integrating multiple microbe-drug-disease similarity measures, we first constructed two distinct microbe-drug networks. In addition, based on different features of microbes and drugs, we further constructed two novel microbe-drug feature matrices. On this basis, the microbe-drug networks and the constructed feature matrices were further used in a Graph Attention Network to learn complementary topology-aware representations of microbes and drugs. These GAT-derived representations were then integrated with the constructed drug-side and microbe-side feature matrices and input into a Rotation Forest classifier for final association prediction. Experimental results and case studies demonstrated that GATROF predicts microbe-drug associations more accurately than existing state-of-the-art methods.

Conclusion: GATROF provides a new integrated predictive framework for predicting potential microbe-drug associations. By combining heterogeneous biological information, GAT-based topological representation learning, and Rotation Forest classification, GATROF may help prioritize candidate drug-microbe associations for further biological validation.

背景:近年来,随着医学领域药物研究的多样化和扩大,药物特别是抗生素的广泛使用,导致微生物耐药性增加。因此,探索药物和微生物之间的潜在联系变得至关重要。然而,传统的生物实验非常昂贵且耗时。因此,开发更有效的计算模型来预测微生物和药物之间的潜在关联既必要又具有挑战性。结果:我们提出了一种基于混合异构图的微生物-药物关联预测框架GATROF。在GATROF中,通过整合多种微生物-药物-疾病相似性度量,我们首先构建了两个不同的微生物-药物网络。此外,根据微生物和药物的不同特征,我们进一步构建了两种新的微生物-药物特征矩阵。在此基础上,将微生物-药物网络和构建的特征矩阵进一步应用于图注意网络中,学习微生物和药物的互补拓扑感知表示。然后将这些gat衍生的表示与构建的药物侧和微生物侧特征矩阵集成,并输入到旋转森林分类器中进行最终的关联预测。实验结果和案例研究表明,GATROF比现有的最先进的方法更准确地预测微生物与药物的关联。结论:GATROF为预测微生物与药物的潜在关联提供了新的综合预测框架。通过结合异质生物信息、基于gatt的拓扑表示学习和旋转森林分类,gatt可以帮助确定候选药物-微生物关联的优先级,以进一步进行生物学验证。
{"title":"A microbe-drug association prediction model based on graph attention networks and rotation forest.","authors":"Jing Li, Juncai Li, Qijia Chen, Zhong Wang, Xianzhi Liu, Mingmin Liang, Junzhuang Wang, Hongyuan Ding, Bin Zeng, Lei Wang","doi":"10.3389/fbinf.2026.1871436","DOIUrl":"10.3389/fbinf.2026.1871436","url":null,"abstract":"<p><strong>Background: </strong>In recent years, with the diversification and expansion of drug research in the medical field, the widespread use of drugs, particularly antibiotics, has led to increased microbial resistance. Consequently, exploring potential associations between drugs and microbes has become critically important. However, traditional biological experiments are extremely expensive and time-consuming. Therefore, developing more effective computational models for predicting potential associations between microbes and drugs is both essential and challenging.</p><p><strong>Results: </strong>We proposed GATROF, a hybrid heterogeneous graph-based framework for microbe-drug association prediction. In GATROF, by integrating multiple microbe-drug-disease similarity measures, we first constructed two distinct microbe-drug networks. In addition, based on different features of microbes and drugs, we further constructed two novel microbe-drug feature matrices. On this basis, the microbe-drug networks and the constructed feature matrices were further used in a Graph Attention Network to learn complementary topology-aware representations of microbes and drugs. These GAT-derived representations were then integrated with the constructed drug-side and microbe-side feature matrices and input into a Rotation Forest classifier for final association prediction. Experimental results and case studies demonstrated that GATROF predicts microbe-drug associations more accurately than existing state-of-the-art methods.</p><p><strong>Conclusion: </strong>GATROF provides a new integrated predictive framework for predicting potential microbe-drug associations. By combining heterogeneous biological information, GAT-based topological representation learning, and Rotation Forest classification, GATROF may help prioritize candidate drug-microbe associations for further biological validation.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1871436"},"PeriodicalIF":3.6,"publicationDate":"2026-08-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13506770/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148835425","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Single cell mapping of B and T cell dynamics in breast cancer lymph node metastasis. 乳腺癌淋巴结转移过程中B细胞和T细胞动力学的单细胞定位。
IF 3.6 Q2 MATHEMATICAL & COMPUTATIONAL BIOLOGY Pub Date : 2026-08-12 eCollection Date: 2026-01-01 DOI: 10.3389/fbinf.2026.1833416
Sajid Khan, Sabahat Jamil, Muhammad Hamza, Zarlish Attique, Suping Zhang

Metastatic breast cancer remains difficult to cure, and the way B and T lymphocytes adapt across metastatic niches especially under therapy remains insufficiently defined. Clarifying compartment specific immune remodeling may help explain resistance to PD-1/PD-L1 blockade and identify actionable targets. We performed an integrated meta-analysis of single cell RNA-seq datasets from normal breast tissue, primary tumors, tumor-draining lymph nodes (TLNs), and peripheral blood mononuclear cells (PBMCs), focusing on B and Tcell states. Immune composition differed notably by compartment. Tumors were enriched for effector CD8 states (CD8 cytotoxic 20.1%; CD8 activated 13.5%), whereas TLNs preserved larger naïve and memory reservoirs (CD4 naïve 40.7%; B naïve 11.4%; B memory 12.0%) and contained a higher B cell fraction than tumors (39.6% vs. 19.5%). Post therapy, PBMCs and TLNs showed increased BTLA-HVEM (TNFRSF14) checkpoint signaling and enhanced MIF-CD74 interactions with a shift from CD44 toward CXCR4, consistent with CXCR4 driven migratory and survival programs. In TLNs, TNFRSF14 signaling was unidirectional (B→T), absent in the reverse direction, and not detected in tumors. Clinically, higher tumor CXCR4 combined with lower TNFRSF14 was associated with shorter progression free survival in TCGA-BRCA, most evident in node positive, early stage disease. To target the BTLA-HVEM checkpoint axis, we performed structure guided de novo peptide design using the native HVEM (23-39) peptide as an active structural template, followed by docking and molecular dynamics simulations. The optimized De novo-P2 peptide showed stable and favorable interactions at the BTLA interface, supporting its potential as a competitive modulator of BTLA-HVEM signaling. These data define niche specific lymphocyte remodeling and implicate BTLA-HVEM and CXCL12-CXCR4 as candidate biomarkers and therapeutic targets linked to PD-1/PD-L1 resistance.

转移性乳腺癌仍然难以治愈,B淋巴细胞和T淋巴细胞适应转移性壁龛的方式,特别是在治疗下,仍然没有充分的定义。阐明室特异性免疫重构可能有助于解释PD-1/PD-L1阻断的耐药性,并确定可行的靶点。我们对来自正常乳腺组织、原发肿瘤、肿瘤引流淋巴结(tln)和外周血单核细胞(PBMCs)的单细胞RNA-seq数据集进行了综合荟萃分析,重点关注B细胞和t细胞状态。不同细胞间的免疫组成差异显著。肿瘤富集了CD8效应态(CD8细胞毒性20.1%,CD8活化13.5%),而tln保留了更大的naïve和记忆库(CD4 naïve 40.7%; B naïve 11.4%; B记忆12.0%),并且比肿瘤含有更高的B细胞比例(39.6%比19.5%)。治疗后,pbmc和tln显示BTLA-HVEM (TNFRSF14)检查点信号增加,MIF-CD74相互作用增强,从CD44向CXCR4转移,与CXCR4驱动的迁移和生存计划一致。在tln中,TNFRSF14信号是单向的(B→T),相反方向不存在,在肿瘤中未检测到。在临床上,TCGA-BRCA中,较高的肿瘤CXCR4合并较低的TNFRSF14与较短的无进展生存期相关,在淋巴结阳性的早期疾病中最为明显。为了靶向BTLA-HVEM检查点轴,我们以天然HVEM(23-39)肽作为活性结构模板,进行了结构引导的从头肽设计,然后进行对接和分子动力学模拟。优化后的De novo-P2肽在BTLA界面上表现出稳定和良好的相互作用,支持其作为BTLA- hvem信号传导的竞争性调节剂的潜力。这些数据定义了小生境特异性淋巴细胞重塑,并暗示BTLA-HVEM和CXCL12-CXCR4是与PD-1/PD-L1耐药相关的候选生物标志物和治疗靶点。
{"title":"Single cell mapping of B and T cell dynamics in breast cancer lymph node metastasis.","authors":"Sajid Khan, Sabahat Jamil, Muhammad Hamza, Zarlish Attique, Suping Zhang","doi":"10.3389/fbinf.2026.1833416","DOIUrl":"10.3389/fbinf.2026.1833416","url":null,"abstract":"<p><p>Metastatic breast cancer remains difficult to cure, and the way B and T lymphocytes adapt across metastatic niches especially under therapy remains insufficiently defined. Clarifying compartment specific immune remodeling may help explain resistance to PD-1/PD-L1 blockade and identify actionable targets. We performed an integrated meta-analysis of single cell RNA-seq datasets from normal breast tissue, primary tumors, tumor-draining lymph nodes (TLNs), and peripheral blood mononuclear cells (PBMCs), focusing on B and Tcell states. Immune composition differed notably by compartment. Tumors were enriched for effector CD8 states (CD8 cytotoxic 20.1%; CD8 activated 13.5%), whereas TLNs preserved larger naïve and memory reservoirs (CD4 naïve 40.7%; B naïve 11.4%; B memory 12.0%) and contained a higher B cell fraction than tumors (39.6% vs. 19.5%). Post therapy, PBMCs and TLNs showed increased BTLA-HVEM (TNFRSF14) checkpoint signaling and enhanced MIF-CD74 interactions with a shift from CD44 toward CXCR4, consistent with CXCR4 driven migratory and survival programs. In TLNs, TNFRSF14 signaling was unidirectional (B→T), absent in the reverse direction, and not detected in tumors. Clinically, higher tumor CXCR4 combined with lower TNFRSF14 was associated with shorter progression free survival in TCGA-BRCA, most evident in node positive, early stage disease. To target the BTLA-HVEM checkpoint axis, we performed structure guided <i>de novo</i> peptide design using the native HVEM (23-39) peptide as an active structural template, followed by docking and molecular dynamics simulations. The optimized De novo-P2 peptide showed stable and favorable interactions at the BTLA interface, supporting its potential as a competitive modulator of BTLA-HVEM signaling. These data define niche specific lymphocyte remodeling and implicate BTLA-HVEM and CXCL12-CXCR4 as candidate biomarkers and therapeutic targets linked to PD-1/PD-L1 resistance.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1833416"},"PeriodicalIF":3.6,"publicationDate":"2026-08-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13506802/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148835433","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Correction: Transcriptome-informed metabolic modeling reveals astrocyte-specific vulnerabilities in mild cognitive impairment and Alzheimer's disease progression. 更正:转录组代谢模型揭示了星形胶质细胞在轻度认知障碍和阿尔茨海默病进展中的特异性脆弱性。
IF 3.6 Q2 MATHEMATICAL & COMPUTATIONAL BIOLOGY Pub Date : 2026-08-11 eCollection Date: 2026-01-01 DOI: 10.3389/fbinf.2026.1953979
Andrea Angarita-Rodríguez, Viviana Vargas-López, Andrés Pinzón, Adrián Sandoval-Hernandez, Kai Kang, Leping Li, Jason Papin, Pedro Puentes-Rozo, Andrés Felipe Aristizábal, Janneth González

[This corrects the article DOI: 10.3389/fbinf.2026.1816121.].

[这更正了文章DOI: 10.3389/fbinf.2026.1816121.]。
{"title":"Correction: Transcriptome-informed metabolic modeling reveals astrocyte-specific vulnerabilities in mild cognitive impairment and Alzheimer's disease progression.","authors":"Andrea Angarita-Rodríguez, Viviana Vargas-López, Andrés Pinzón, Adrián Sandoval-Hernandez, Kai Kang, Leping Li, Jason Papin, Pedro Puentes-Rozo, Andrés Felipe Aristizábal, Janneth González","doi":"10.3389/fbinf.2026.1953979","DOIUrl":"https://doi.org/10.3389/fbinf.2026.1953979","url":null,"abstract":"<p><p>[This corrects the article DOI: 10.3389/fbinf.2026.1816121.].</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1953979"},"PeriodicalIF":3.6,"publicationDate":"2026-08-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13504194/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148820441","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
Design and in silico evaluation of Thiazole-Isoxazole hybrids as PPARγ-Targeted antidiabetic agents. 噻唑-异恶唑复合物作为ppar γ靶向降糖药的设计与硅评价。
IF 3.6 Q2 MATHEMATICAL & COMPUTATIONAL BIOLOGY Pub Date : 2026-08-11 eCollection Date: 2026-01-01 DOI: 10.3389/fbinf.2026.1873262
Abhirami Pv, Gupta Dheeraj Rajesh, N V L Sirisha Mulukuri, Ranjitha A, Niyas Rehman, Shiv Basant Kumar, Dileep Kumar, Pankaj Kumar

Introduction: Diabetes mellitus is a chronic metabolic disorder characterised by persistent hyperglycemia resulting from impaired insulin secretion, action, or both. Prolonged hyperglycemia disrupts metabolic homeostasis and increases the risk of complications, including cardiovascular disease, neuropathy, nephropathy, and retinopathy. Current therapies target multiple pathways but are limited by reduced efficacy and adverse effects. Peroxisome proliferator-activated receptor gamma (PPARγ) is a key regulator of glucose and lipid metabolism and an important therapeutic target. Thiazole and isoxazole scaffolds possess antidiabetic potential, and their hybridisation may improve efficacy and safety.

Objective: To design and evaluate a novel series of thiazole-linked isoxazole derivatives targeting the alternate ligand-binding domain of PPARγ for potential antidiabetic activity.

Result: The designed compounds (C1-C210) were docked into PPARγ (Protein Data Bank: 5 GTN), yielding nine candidates with binding affinities of -11.11 to -9.97 kcal/mol. Reference ligands included pioglitazone, Q-35, and S-35, with C139 showing the highest affinity. Molecular dynamics simulations (200 ns) confirmed stability. Advanced Trajectory analyses indicate favourable conformational sampling, correlated residue motions, and a relatively confined conformational space. In silico drug-likeness profiling further supported their activity profile.

Conclusion: In this computational study, the thiazole-linked isoxazole compound C139 demonstrated probable antidiabetic activity against PPARγ. However, these findings remain predictive and require further validation through comprehensive in vitro and in vivo studies.

简介:糖尿病是一种慢性代谢紊乱,其特征是胰岛素分泌、作用受损或两者兼而有之,导致持续高血糖。长期的高血糖会破坏代谢稳态,增加并发症的风险,包括心血管疾病、神经病变、肾病和视网膜病变。目前的治疗方法针对多种途径,但由于疗效降低和不良反应而受到限制。过氧化物酶体增殖物激活受体γ (PPARγ)是葡萄糖和脂质代谢的关键调节因子,也是重要的治疗靶点。噻唑和异恶唑支架具有抗糖尿病潜能,它们的杂交可以提高疗效和安全性。目的:设计并评价一系列新的噻唑-异恶唑衍生物,其靶向PPARγ的替代配体结合结构域,具有潜在的抗糖尿病活性。结果:设计的化合物(C1-C210)与PPARγ(蛋白质数据库:5 GTN)对接,得到9个结合亲和力为-11.11 ~ -9.97 kcal/mol的候选化合物。参考配体包括吡格列酮、Q-35和S-35,其中C139亲和力最高。分子动力学模拟(200 ns)证实了其稳定性。先进的轨迹分析表明有利的构象采样,相关的残余运动和相对有限的构象空间。在计算机上,药物相似性分析进一步支持了他们的活动特征。结论:在这项计算研究中,噻唑连接的异恶唑化合物C139可能具有抗PPARγ的抗糖尿病活性。然而,这些发现仍然具有预测性,需要通过全面的体外和体内研究进一步验证。
{"title":"Design and <i>in silico</i> evaluation of Thiazole-Isoxazole hybrids as PPARγ-Targeted antidiabetic agents.","authors":"Abhirami Pv, Gupta Dheeraj Rajesh, N V L Sirisha Mulukuri, Ranjitha A, Niyas Rehman, Shiv Basant Kumar, Dileep Kumar, Pankaj Kumar","doi":"10.3389/fbinf.2026.1873262","DOIUrl":"10.3389/fbinf.2026.1873262","url":null,"abstract":"<p><strong>Introduction: </strong>Diabetes mellitus is a chronic metabolic disorder characterised by persistent hyperglycemia resulting from impaired insulin secretion, action, or both. Prolonged hyperglycemia disrupts metabolic homeostasis and increases the risk of complications, including cardiovascular disease, neuropathy, nephropathy, and retinopathy. Current therapies target multiple pathways but are limited by reduced efficacy and adverse effects. Peroxisome proliferator-activated receptor gamma (PPARγ) is a key regulator of glucose and lipid metabolism and an important therapeutic target. Thiazole and isoxazole scaffolds possess antidiabetic potential, and their hybridisation may improve efficacy and safety.</p><p><strong>Objective: </strong>To design and evaluate a novel series of thiazole-linked isoxazole derivatives targeting the alternate ligand-binding domain of PPARγ for potential antidiabetic activity.</p><p><strong>Result: </strong>The designed compounds (C1-C210) were docked into PPARγ (Protein Data Bank: 5 GTN), yielding nine candidates with binding affinities of -11.11 to -9.97 kcal/mol. Reference ligands included pioglitazone, Q-35, and S-35, with C139 showing the highest affinity. Molecular dynamics simulations (200 ns) confirmed stability. Advanced Trajectory analyses indicate favourable conformational sampling, correlated residue motions, and a relatively confined conformational space. <i>In silico</i> drug-likeness profiling further supported their activity profile.</p><p><strong>Conclusion: </strong>In this computational study, the thiazole-linked isoxazole compound C139 demonstrated probable antidiabetic activity against PPARγ. However, these findings remain predictive and require further validation through comprehensive <i>in vitro</i> and <i>in vivo</i> studies.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1873262"},"PeriodicalIF":3.6,"publicationDate":"2026-08-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13503372/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148820449","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
PUDU (pipeline for universal diversity unveiling): an accessible end-to-end workflow for taxonomic profiling and ecological visualization of environmental microbiomes across amplicon, shotgun, and long-read sequencing. PUDU (pipeline for universal diversity unveiling):一个可访问的端到端工作流程,用于对环境微生物组进行扩增子测序、鸟枪测序和长读测序的分类分析和生态可视化。
IF 3.6 Q2 MATHEMATICAL & COMPUTATIONAL BIOLOGY Pub Date : 2026-08-11 eCollection Date: 2026-01-01 DOI: 10.3389/fbinf.2026.1909327
Alejandro Medaglia-Mata, Pablo Rojas-Rodríguez, Vojtěch Bystrý, Rossy Guillén-Watson, Olman Gómez-Espinoza, Kattia Núñez-Montero

Background: Environmental microbiome research has advanced through three complementary sequencing modalities, targeted 16S rRNA amplicon sequencing, whole-genome shotgun (WGS) metagenomics, and long-read full-length 16S rRNA profiling, each supported by distinct toolsets with heterogeneous outputs, variable configurations, and different levels of reproducibility documentation. Existing pipelines are typically modality-specific, require substantial configuration expertise, or produce outputs that need further custom scripting before standard ecological analyses can begin. This analytical fragmentation introduces avoidable technical variability and complicates cross-study reproducibility and comparability. PUDU addresses this by integrating all three modalities into a single reproducible workflow with simplified configuration, harmonized outputs across classifiers, and direct compatibility with downstream ecological analysis frameworks.

Results: We present PUDU (Pipeline for Universal Diversity Unveiling), a modular Snakemake workflow that supports amplicon (short-read 16S), shotgun metagenomics (WGS), and long-read 16S analyses from raw reads to standardized outputs for downstream microbial ecology. PUDU performs technology-aware preprocessing and centralized quality control, and integrates established taxonomic approaches, including DADA2 for amplicons, Emu for full-length 16S long reads, and Kraken2/Bracken and Centrifuger for WGS. Across methods, PUDU produces harmonized count and relative-abundance tables at user-defined taxonomic ranks, Krona files, and a standardized Phyloseq-compatible R object to streamline diversity analyses and statistical workflows. PUDU also provides an integrated Shiny interface for metadata-aware alpha/beta diversity, ordination, community composition, and shared-taxa exploration with exportable figures and taxa tables. We demonstrate PUDU on two publicly available environmental datasets spanning rhizosphere WGS and long-read marine sediment 16S, yielding broadly consistent community-level patterns across classifiers (Spearman ρ = 0.936 at phylum level; PERMANOVA R2 = 0.87-0.95) with peak memory below 45 GB on a standard Linux workstation.

Conclusion: PUDU is an end-to-end, reproducible, and extensible framework that enables standardized taxonomic profiling and ecology-oriented analysis across sequencing modalities. By combining harmonized outputs, Phyloseq interoperability, and an integrated visualization layer, PUDU facilitates reproducible, standardized, and comparable environmental microbiome analysis from raw reads to interpretable ecological insights.

背景:环境微生物组研究通过三种互补的测序模式取得进展,即靶向16S rRNA扩增子测序、全基因组霰弹枪(WGS)宏基因组学和长读全长16S rRNA分析,每种模式都有不同的工具集支持,这些工具集具有异质输出、可变配置和不同水平的可重复性文件。现有的管道通常是特定于模式的,需要大量的配置专业知识,或者在开始标准生态分析之前产生需要进一步定制脚本的输出。这种分析碎片化引入了可避免的技术可变性,并使交叉研究的可重复性和可比性复杂化。PUDU通过将所有三种模式集成到一个具有简化配置、跨分类器协调输出以及与下游生态分析框架直接兼容的可重复工作流中来解决这个问题。结果:我们提出了PUDU (Pipeline for Universal Diversity Unveiling),这是一个模块化的Snakemake工作流程,支持扩增子(短读16S)、霰弹枪宏基因组学(WGS)和长读16S分析,从原始reads到下游微生物生态的标准化输出。PUDU执行技术敏感的预处理和集中的质量控制,并集成了已建立的分类方法,包括DADA2扩增子,Emu全长16S长读取,Kraken2/Bracken和离心机用于WGS。在各种方法中,PUDU在用户定义的分类等级、Krona文件和标准化的phyloseq兼容R对象中生成协调的计数和相对丰度表,以简化多样性分析和统计工作流程。PUDU还提供了一个集成的Shiny接口,用于元数据感知的alpha/beta多样性、排序、群落组成和共享分类群探索,以及可导出的图形和分类群表。我们在两个公开的环境数据集上展示了PUDU,这些数据集跨越根际WGS和长读海洋沉积物16S,在标准Linux工作站上,在分类器上产生了广泛一致的群落水平模式(在门水平上,Spearman ρ = 0.936; PERMANOVA R2 = 0.87-0.95),峰值内存低于45 GB。结论:PUDU是一个端到端的、可重复的、可扩展的框架,可以实现标准化的分类分析和跨测序模式的面向生态的分析。通过将协调输出、Phyloseq互操作性和集成的可视化层相结合,PUDU促进了从原始读数到可解释的生态见解的可重复、标准化和可比较的环境微生物组分析。
{"title":"PUDU (pipeline for universal diversity unveiling): an accessible end-to-end workflow for taxonomic profiling and ecological visualization of environmental microbiomes across amplicon, shotgun, and long-read sequencing.","authors":"Alejandro Medaglia-Mata, Pablo Rojas-Rodríguez, Vojtěch Bystrý, Rossy Guillén-Watson, Olman Gómez-Espinoza, Kattia Núñez-Montero","doi":"10.3389/fbinf.2026.1909327","DOIUrl":"10.3389/fbinf.2026.1909327","url":null,"abstract":"<p><strong>Background: </strong>Environmental microbiome research has advanced through three complementary sequencing modalities, targeted 16S rRNA amplicon sequencing, whole-genome shotgun (WGS) metagenomics, and long-read full-length 16S rRNA profiling, each supported by distinct toolsets with heterogeneous outputs, variable configurations, and different levels of reproducibility documentation. Existing pipelines are typically modality-specific, require substantial configuration expertise, or produce outputs that need further custom scripting before standard ecological analyses can begin. This analytical fragmentation introduces avoidable technical variability and complicates cross-study reproducibility and comparability. PUDU addresses this by integrating all three modalities into a single reproducible workflow with simplified configuration, harmonized outputs across classifiers, and direct compatibility with downstream ecological analysis frameworks.</p><p><strong>Results: </strong>We present PUDU (Pipeline for Universal Diversity Unveiling), a modular Snakemake workflow that supports amplicon (short-read 16S), shotgun metagenomics (WGS), and long-read 16S analyses from raw reads to standardized outputs for downstream microbial ecology. PUDU performs technology-aware preprocessing and centralized quality control, and integrates established taxonomic approaches, including DADA2 for amplicons, Emu for full-length 16S long reads, and Kraken2/Bracken and Centrifuger for WGS. Across methods, PUDU produces harmonized count and relative-abundance tables at user-defined taxonomic ranks, Krona files, and a standardized Phyloseq-compatible R object to streamline diversity analyses and statistical workflows. PUDU also provides an integrated Shiny interface for metadata-aware alpha/beta diversity, ordination, community composition, and shared-taxa exploration with exportable figures and taxa tables. We demonstrate PUDU on two publicly available environmental datasets spanning rhizosphere WGS and long-read marine sediment 16S, yielding broadly consistent community-level patterns across classifiers (Spearman ρ = 0.936 at phylum level; PERMANOVA R<sup>2</sup> = 0.87-0.95) with peak memory below 45 GB on a standard Linux workstation.</p><p><strong>Conclusion: </strong>PUDU is an end-to-end, reproducible, and extensible framework that enables standardized taxonomic profiling and ecology-oriented analysis across sequencing modalities. By combining harmonized outputs, Phyloseq interoperability, and an integrated visualization layer, PUDU facilitates reproducible, standardized, and comparable environmental microbiome analysis from raw reads to interpretable ecological insights.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1909327"},"PeriodicalIF":3.6,"publicationDate":"2026-08-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13503585/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148820468","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
LungMicroHostR: an R package for integrated host-microbiome analysis of bronchoalveolar lavage fluid metagenomic sequencing data. LungMicroHostR:用于支气管肺泡灌洗液宏基因组测序数据的综合宿主-微生物组分析的R包。
IF 3.6 Q2 MATHEMATICAL & COMPUTATIONAL BIOLOGY Pub Date : 2026-08-10 eCollection Date: 2026-01-01 DOI: 10.3389/fbinf.2026.1906036
Nan Li, Jing Hu, Wanning Tong, Chengdong Liu, Yun Ding, Ning Li, Zhigang Cai

Introduction: Bronchoalveolar lavage fluid metagenomic next-generation sequencing captures microbial profiles and host-derived molecular measurements from the same respiratory specimen, but downstream analysis requires coordinated handling of low-biomass microbial signals, negative-control information and multiple feature tables.

Methods: We developed LungMicroHostR, an R package for downstream host-microbiome analysis of bronchoalveolar lavage fluid metagenomic sequencing data. The package brings processed microbial profiles, host-derived molecular measurements, sample metadata and negative-control information into a unified R workflow for feature filtering, comparative model evaluation, visualization and reproducible reporting.

Results: Using the public GSE252118 resource comprising 402 samples from patients with lung cancer or pulmonary infections, LungMicroHostR assembled matched microbial, host and clinical feature tables, estimated prevalence in negative controls and compared host transcriptomic, microbial-profile and combined host-microbial models. In the test set, the 10-feature host transcriptome nearest-centroid model achieved an AUC of 0.772 (95% confidence interval, 0.680-0.860), the five-feature RNA microbial logistic model achieved an AUC of 0.745 (0.655-0.832), and the combined host transcriptome-RNA microbial logistic model achieved an AUC of 0.765 (0.655-0.866) with balanced accuracy of 0.720. An external PRJNA714488 BALF shotgun metagenomic dataset was additionally analysed at the mOTU level; LungMicroHostR matched the resulting feature table with phenotype metadata and generated a 388-feature by 26-sample microbial abundance matrix.

Discussion: LungMicroHostR provides documented functions for respiratory metagenomic analyses that require joint evaluation of microbial profiles, host-derived measurements, negative-control information and external microbial feature tables.

新一代宏基因组测序可捕获来自同一呼吸道标本的微生物谱和宿主来源的分子测量,但下游分析需要协调处理低生物量微生物信号、阴性对照信息和多个特征表。方法:我们开发了LungMicroHostR,一个用于支气管肺泡灌洗液下游宿主微生物组分析的R包。该软件包将处理过的微生物剖面、宿主衍生的分子测量、样本元数据和阴性控制信息整合到统一的R工作流中,用于特征过滤、比较模型评估、可视化和可重复报告。结果:利用公共GSE252118资源,包括402例肺癌或肺部感染患者的样本,LungMicroHostR组装了匹配的微生物、宿主和临床特征表,估计了阴性对照组的患病率,并比较了宿主转录组学、微生物谱和宿主-微生物组合模型。在测试集中,10个特征的宿主转录组最近质心模型的AUC为0.772(95%置信区间为0.680-0.860),5个特征的RNA微生物logistic模型的AUC为0.745(0.655-0.832),组合宿主转录组-RNA微生物logistic模型的AUC为0.765(0.655-0.866),平衡精度为0.720。此外,在mOTU水平上分析外部PRJNA714488 BALF霰弹枪宏基因组数据集;LungMicroHostR将得到的特征表与表型元数据进行匹配,并由26个样本生成388个特征的微生物丰度矩阵。讨论:LungMicroHostR为呼吸宏基因组分析提供了记录功能,这些分析需要微生物谱、宿主衍生测量、阴性对照信息和外部微生物特征表的联合评估。
{"title":"LungMicroHostR: an R package for integrated host-microbiome analysis of bronchoalveolar lavage fluid metagenomic sequencing data.","authors":"Nan Li, Jing Hu, Wanning Tong, Chengdong Liu, Yun Ding, Ning Li, Zhigang Cai","doi":"10.3389/fbinf.2026.1906036","DOIUrl":"10.3389/fbinf.2026.1906036","url":null,"abstract":"<p><strong>Introduction: </strong>Bronchoalveolar lavage fluid metagenomic next-generation sequencing captures microbial profiles and host-derived molecular measurements from the same respiratory specimen, but downstream analysis requires coordinated handling of low-biomass microbial signals, negative-control information and multiple feature tables.</p><p><strong>Methods: </strong>We developed LungMicroHostR, an R package for downstream host-microbiome analysis of bronchoalveolar lavage fluid metagenomic sequencing data. The package brings processed microbial profiles, host-derived molecular measurements, sample metadata and negative-control information into a unified R workflow for feature filtering, comparative model evaluation, visualization and reproducible reporting.</p><p><strong>Results: </strong>Using the public GSE252118 resource comprising 402 samples from patients with lung cancer or pulmonary infections, LungMicroHostR assembled matched microbial, host and clinical feature tables, estimated prevalence in negative controls and compared host transcriptomic, microbial-profile and combined host-microbial models. In the test set, the 10-feature host transcriptome nearest-centroid model achieved an AUC of 0.772 (95% confidence interval, 0.680-0.860), the five-feature RNA microbial logistic model achieved an AUC of 0.745 (0.655-0.832), and the combined host transcriptome-RNA microbial logistic model achieved an AUC of 0.765 (0.655-0.866) with balanced accuracy of 0.720. An external PRJNA714488 BALF shotgun metagenomic dataset was additionally analysed at the mOTU level; LungMicroHostR matched the resulting feature table with phenotype metadata and generated a 388-feature by 26-sample microbial abundance matrix.</p><p><strong>Discussion: </strong>LungMicroHostR provides documented functions for respiratory metagenomic analyses that require joint evaluation of microbial profiles, host-derived measurements, negative-control information and external microbial feature tables.</p>","PeriodicalId":73066,"journal":{"name":"Frontiers in bioinformatics","volume":"6 ","pages":"1906036"},"PeriodicalIF":3.6,"publicationDate":"2026-08-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13500579/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148814970","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
引用次数: 0
期刊
Frontiers in bioinformatics
全部 Acc. Chem. Res. ACS Applied Bio Materials ACS Appl. Electron. Mater. ACS Appl. Energy Mater. ACS Appl. Mater. Interfaces ACS Appl. Nano Mater. ACS Appl. Polym. Mater. ACS BIOMATER-SCI ENG ACS Catal. ACS Cent. Sci. ACS Chem. Biol. ACS Chemical Health & Safety ACS Chem. Neurosci. ACS Comb. Sci. ACS Earth Space Chem. ACS Energy Lett. ACS Infect. Dis. ACS Macro Lett. ACS Mater. Lett. ACS Med. Chem. Lett. ACS Nano ACS Omega ACS Photonics ACS Sens. ACS Sustainable Chem. Eng. ACS Synth. Biol. Anal. Chem. BIOCHEMISTRY-US Bioconjugate Chem. BIOMACROMOLECULES Chem. Res. Toxicol. Chem. Rev. Chem. Mater. CRYST GROWTH DES ENERG FUEL Environ. Sci. Technol. Environ. Sci. Technol. Lett. Eur. J. Inorg. Chem. IND ENG CHEM RES Inorg. Chem. J. Agric. Food. Chem. J. Chem. Eng. Data J. Chem. Educ. J. Chem. Inf. Model. J. Chem. Theory Comput. J. Med. Chem. J. Nat. Prod. J PROTEOME RES J. Am. Chem. Soc. LANGMUIR MACROMOLECULES Mol. Pharmaceutics Nano Lett. Org. Lett. ORG PROCESS RES DEV ORGANOMETALLICS J. Org. Chem. J. Phys. Chem. J. Phys. Chem. A J. Phys. Chem. B J. Phys. Chem. C J. Phys. Chem. Lett. Analyst Anal. Methods Biomater. Sci. Catal. Sci. Technol. Chem. Commun. Chem. Soc. Rev. CHEM EDUC RES PRACT CRYSTENGCOMM Dalton Trans. Energy Environ. Sci. ENVIRON SCI-NANO ENVIRON SCI-PROC IMP ENVIRON SCI-WAT RES Faraday Discuss. Food Funct. Green Chem. Inorg. Chem. Front. Integr. Biol. J. Anal. At. Spectrom. J. Mater. Chem. A J. Mater. Chem. B J. Mater. Chem. C Lab Chip Mater. Chem. Front. Mater. Horiz. MEDCHEMCOMM Metallomics Mol. Biosyst. Mol. Syst. Des. Eng. Nanoscale Nanoscale Horiz. Nat. Prod. Rep. New J. Chem. Org. Biomol. Chem. Org. Chem. Front. PHOTOCH PHOTOBIO SCI PCCP Polym. Chem.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
Book学术官方微信
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1