Pub Date : 2026-09-03DOI: 10.1371/journal.pcbi.1014722
Lara M Kösters, Kevin Karbstein, Ladislav Hodač, Laura Albreht, Elvira Sahuquillo Balbuena, Daniel Botello, Olivier Hardy, Phebian Odufuwa, Eva Pardo Otero, Aireen Phang, Manuel Pimentel, Rosalía Piñeiro, James Smith, Peter Wilkie, Patrick Mäder, Jana Wäldchen
In taxonomic research, traditional phylogenetic tree- and structure-based analyses of genetic data are increasingly complemented by machine-learning-based identification and representation learning. Although the amount of DNA data needed to train state-of-the-art machine learning models often exceeds what can realistically be collected and sequenced in biological studies, the number of samples can be extended artificially through data augmentation. Genetic data augmentation usually refers to the introduction of random base variations, translocations, and reverse complementing. These augmentations do not take into account the inherent structures of populations and species, potentially blurring the lines between entities within genetic datasets. Here, we propose DNAInterpolator, an approach based on interpolation of DNA sequences within a given dataset that presents a neighbor-guided alternative to random mutations. We tested interpolation as an augmentation technique using four flowering plant datasets and an artificial neural network trained to predict genetic distances between paired samples. To address unequally distributed distances within our training datasets, we examined the effect of balancing the distance distribution by curating interpolated sequences. We found that balancing helps models capture genetic distances across the full distance range by strengthening performance in underrepresented regions of the distribution. Our new approach leverages the potential of taxonomic DNA datasets for modern machine learning applications.
{"title":"Balanced DNA interpolation improves learning of genetic distance-informed embeddings in plants.","authors":"Lara M Kösters, Kevin Karbstein, Ladislav Hodač, Laura Albreht, Elvira Sahuquillo Balbuena, Daniel Botello, Olivier Hardy, Phebian Odufuwa, Eva Pardo Otero, Aireen Phang, Manuel Pimentel, Rosalía Piñeiro, James Smith, Peter Wilkie, Patrick Mäder, Jana Wäldchen","doi":"10.1371/journal.pcbi.1014722","DOIUrl":"https://doi.org/10.1371/journal.pcbi.1014722","url":null,"abstract":"<p><p>In taxonomic research, traditional phylogenetic tree- and structure-based analyses of genetic data are increasingly complemented by machine-learning-based identification and representation learning. Although the amount of DNA data needed to train state-of-the-art machine learning models often exceeds what can realistically be collected and sequenced in biological studies, the number of samples can be extended artificially through data augmentation. Genetic data augmentation usually refers to the introduction of random base variations, translocations, and reverse complementing. These augmentations do not take into account the inherent structures of populations and species, potentially blurring the lines between entities within genetic datasets. Here, we propose DNAInterpolator, an approach based on interpolation of DNA sequences within a given dataset that presents a neighbor-guided alternative to random mutations. We tested interpolation as an augmentation technique using four flowering plant datasets and an artificial neural network trained to predict genetic distances between paired samples. To address unequally distributed distances within our training datasets, we examined the effect of balancing the distance distribution by curating interpolated sequences. We found that balancing helps models capture genetic distances across the full distance range by strengthening performance in underrepresented regions of the distribution. Our new approach leverages the potential of taxonomic DNA datasets for modern machine learning applications.</p>","PeriodicalId":20241,"journal":{"name":"PLoS Computational Biology","volume":"22 9","pages":"e1014722"},"PeriodicalIF":3.6,"publicationDate":"2026-09-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148887192","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-09-03DOI: 10.1371/journal.pcbi.1014746
Binod Pant, Marko Lalovic, István Z Kiss, Mauricio Santillana
During epidemic outbreaks, populations adapt their behavior in response to disease burden, fundamentally altering transmission dynamics. Despite this, most compartmental models assume constant contact rates throughout outbreaks. To quantify biases from this assumption, we fitted a baseline SEIRD model with constant transmission and three behavioral variants-incorporating mortality-driven transmission reduction via exponential, rational, and mixed functional forms-to COVID-19 mortality data from 20 selected US locations during the first pandemic wave (March-July 2020). All three behavioral models achieved a lower median normalized sum of squared error in at least 18 of 20 locations, and Bayesian model selection favored them in at least 18 of 20 locations. More importantly, we identified systematic biases when behavioral responses are ignored: the baseline model consistently underestimated the basic reproduction number (ℛ0) while paradoxically overestimating the final epidemic size. Median ℛ0 estimates from the behavioral models exceeded the baseline estimates across all 20 locations, yet baseline models predicted larger cumulative infection burdens. Controlled synthetic experiments-where mortality trajectories were generated from behavioral models with known parameters-confirmed these biases result from model misspecification rather than data quality or stochastic variation. We prove analytically that for any fixed ℛ0, the baseline model overestimates cumulative infections compared to behavioral models where mortality reduces transmission, regardless of functional form. This dual bias has potential implications for pandemic response: standard models may simultaneously underestimate pathogen contagiousness, which could contribute to delayed or insufficient early interventions while overestimating infection burden, which could bias planning for later epidemic phases. Our findings across 20 geographically diverse locations demonstrate that incorporating behavioral change substantially improves both model fit and estimation of epidemiological parameters relevant for public health policy.
{"title":"The paradox of neglecting changes in behavior: How standard epidemic models misestimate both transmissibility and final epidemic size.","authors":"Binod Pant, Marko Lalovic, István Z Kiss, Mauricio Santillana","doi":"10.1371/journal.pcbi.1014746","DOIUrl":"https://doi.org/10.1371/journal.pcbi.1014746","url":null,"abstract":"<p><p>During epidemic outbreaks, populations adapt their behavior in response to disease burden, fundamentally altering transmission dynamics. Despite this, most compartmental models assume constant contact rates throughout outbreaks. To quantify biases from this assumption, we fitted a baseline SEIRD model with constant transmission and three behavioral variants-incorporating mortality-driven transmission reduction via exponential, rational, and mixed functional forms-to COVID-19 mortality data from 20 selected US locations during the first pandemic wave (March-July 2020). All three behavioral models achieved a lower median normalized sum of squared error in at least 18 of 20 locations, and Bayesian model selection favored them in at least 18 of 20 locations. More importantly, we identified systematic biases when behavioral responses are ignored: the baseline model consistently underestimated the basic reproduction number (ℛ0) while paradoxically overestimating the final epidemic size. Median ℛ0 estimates from the behavioral models exceeded the baseline estimates across all 20 locations, yet baseline models predicted larger cumulative infection burdens. Controlled synthetic experiments-where mortality trajectories were generated from behavioral models with known parameters-confirmed these biases result from model misspecification rather than data quality or stochastic variation. We prove analytically that for any fixed ℛ0, the baseline model overestimates cumulative infections compared to behavioral models where mortality reduces transmission, regardless of functional form. This dual bias has potential implications for pandemic response: standard models may simultaneously underestimate pathogen contagiousness, which could contribute to delayed or insufficient early interventions while overestimating infection burden, which could bias planning for later epidemic phases. Our findings across 20 geographically diverse locations demonstrate that incorporating behavioral change substantially improves both model fit and estimation of epidemiological parameters relevant for public health policy.</p>","PeriodicalId":20241,"journal":{"name":"PLoS Computational Biology","volume":"22 9","pages":"e1014746"},"PeriodicalIF":3.6,"publicationDate":"2026-09-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148887770","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-09-03DOI: 10.1371/journal.pcbi.1014698
Yixiang Huang, Lei Yang, Jiudong Wang, Xinqi Gong
Protein complexes are molecular machines that execute essential cellular functions, but their computational identification remains challenging. Existing protein complex identification methods largely rely on PPI network topology, functional annotations, or protein-level biochemical evidence. Although these approaches have recovered many biologically meaningful assemblies, they are often sensitive to incomplete or noisy interactomes and provide limited mechanistic insight into the residue- and interface-level determinants of complex formation. In particular, conventional PPI-based graph representations indicate whether proteins are associated, but usually ignore how protein subunits physically interact through spatially organized residues and structural interfaces. These limitations motivate the development of computational frameworks that connect residue-scale structural cues with interactome-scale organization. Here we present PCIPG, a multi-scale probabilistic graph framework that jointly models residues, proteins, interactions and complexes. PCIPG encodes residue-level physicochemical descriptors on intra-chain contact maps, screens informative residues to construct structure-aware protein representations and propagates these representations over the PPI graph to infer a protein-complex membership matrix. To couple complex membership with sparse interaction evidence, PCIPG reconstructs the network using a zero-inflated Bernoulli-Exponential likelihood, providing a principled learning signal under missing-edge and noise regimes. Across five Saccharomyces cerevisiae benchmarks, PCIPG achieved higher average F1 and Acc than the representative baseline methods included in this study, with average improvements of 11.46% and 3.64%, respectively. On the evaluated human interactomes, PCIPG achieved the highest F1 score among the compared methods on HCT116 and HEK293T, whereas its performance on HuRI was below that of AdaPPI and ClusterONE. Embedding-guided interaction completion improved PCIPG's performance relative to its results on the corresponding original human PPI networks. Beyond complex calling, PCIPG supports core-module mining by recovering known cores and delineating coherent accessory modules within assemblies; several predictions match previously reported functional entities, including TRAPPII- and PCNA-loading-factor-related complexes. At the residue level, residues prioritized by PCIPG show increased overlap with experimentally defined protein-binding interfaces in the evaluated structures. In a computational CFTR case study, the model generated state-dependent interaction predictions that partially overlapped with experimentally profiled wild-type and ΔF508 interaction networks. Together, PCIPG bridges residue-scale structural cues with interactome-scale organization to enable interpretable and scalable protein complex identification. Code and data are available at https://github.com/hyx-1/PCIPG.
{"title":"PCIPG: A comprehensive framework for protein complex identification based on a probabilistic graphical model.","authors":"Yixiang Huang, Lei Yang, Jiudong Wang, Xinqi Gong","doi":"10.1371/journal.pcbi.1014698","DOIUrl":"https://doi.org/10.1371/journal.pcbi.1014698","url":null,"abstract":"<p><p>Protein complexes are molecular machines that execute essential cellular functions, but their computational identification remains challenging. Existing protein complex identification methods largely rely on PPI network topology, functional annotations, or protein-level biochemical evidence. Although these approaches have recovered many biologically meaningful assemblies, they are often sensitive to incomplete or noisy interactomes and provide limited mechanistic insight into the residue- and interface-level determinants of complex formation. In particular, conventional PPI-based graph representations indicate whether proteins are associated, but usually ignore how protein subunits physically interact through spatially organized residues and structural interfaces. These limitations motivate the development of computational frameworks that connect residue-scale structural cues with interactome-scale organization. Here we present PCIPG, a multi-scale probabilistic graph framework that jointly models residues, proteins, interactions and complexes. PCIPG encodes residue-level physicochemical descriptors on intra-chain contact maps, screens informative residues to construct structure-aware protein representations and propagates these representations over the PPI graph to infer a protein-complex membership matrix. To couple complex membership with sparse interaction evidence, PCIPG reconstructs the network using a zero-inflated Bernoulli-Exponential likelihood, providing a principled learning signal under missing-edge and noise regimes. Across five Saccharomyces cerevisiae benchmarks, PCIPG achieved higher average F1 and Acc than the representative baseline methods included in this study, with average improvements of 11.46% and 3.64%, respectively. On the evaluated human interactomes, PCIPG achieved the highest F1 score among the compared methods on HCT116 and HEK293T, whereas its performance on HuRI was below that of AdaPPI and ClusterONE. Embedding-guided interaction completion improved PCIPG's performance relative to its results on the corresponding original human PPI networks. Beyond complex calling, PCIPG supports core-module mining by recovering known cores and delineating coherent accessory modules within assemblies; several predictions match previously reported functional entities, including TRAPPII- and PCNA-loading-factor-related complexes. At the residue level, residues prioritized by PCIPG show increased overlap with experimentally defined protein-binding interfaces in the evaluated structures. In a computational CFTR case study, the model generated state-dependent interaction predictions that partially overlapped with experimentally profiled wild-type and ΔF508 interaction networks. Together, PCIPG bridges residue-scale structural cues with interactome-scale organization to enable interpretable and scalable protein complex identification. Code and data are available at https://github.com/hyx-1/PCIPG.</p>","PeriodicalId":20241,"journal":{"name":"PLoS Computational Biology","volume":"22 9","pages":"e1014698"},"PeriodicalIF":3.6,"publicationDate":"2026-09-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148887444","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-09-03DOI: 10.1371/journal.pcbi.1014771
Wenyu Zhang, Yizheng Wang, Yixiao Zhai, Pinglu Zhang, Yijie Ding, Quan Zou
The rapid emergence of drug-resistant pathogens poses a critical threat to global health. With traditional antibiotics losing efficacy, antimicrobial peptides (AMPs) have gained attention for their unique mechanisms and lower resistance potential. We aimed to accelerate AMP discovery by proposing a closed-loop framework that combines AMP-Hunter (a shared-architecture discriminator for AMP classification and MIC prediction that integrates convolutional neural networks with graph neural networks), and AMP-Forge (a generator integrating multiple sequence alignment to select original candidates) and is guided by minimum inhibitory concentration (MIC)for latent space optimization and candidate selection. AMP-Hunter outperformed baseline models in both AMP classification and MIC prediction, achieving 95.82% accuracy and a 95.80% F1 score on the test set for classification, and an R2 of 0.9245 with an MAE of 0.2305 for MIC prediction. Guided by its predictions, AMP-Forge generated peptide sequences with lower MIC values and improved physicochemical properties associated with antimicrobial activity. Molecular dynamics simulations further provided in silico evidence supporting the antimicrobial potential of selected sequences by identifying stable membrane disruption and insertion behaviors consistent with membrane-targeting activity. Thus, the generation-screening-validation workflow enables reliable discovery of potent AMPs, and provides a practical strategy for rational peptide design, rapid prediction, and translational applications.
{"title":"A unified framework for potency-oriented AMP discovery via multi-modal learning and guided sequence synthesis.","authors":"Wenyu Zhang, Yizheng Wang, Yixiao Zhai, Pinglu Zhang, Yijie Ding, Quan Zou","doi":"10.1371/journal.pcbi.1014771","DOIUrl":"https://doi.org/10.1371/journal.pcbi.1014771","url":null,"abstract":"<p><p>The rapid emergence of drug-resistant pathogens poses a critical threat to global health. With traditional antibiotics losing efficacy, antimicrobial peptides (AMPs) have gained attention for their unique mechanisms and lower resistance potential. We aimed to accelerate AMP discovery by proposing a closed-loop framework that combines AMP-Hunter (a shared-architecture discriminator for AMP classification and MIC prediction that integrates convolutional neural networks with graph neural networks), and AMP-Forge (a generator integrating multiple sequence alignment to select original candidates) and is guided by minimum inhibitory concentration (MIC)for latent space optimization and candidate selection. AMP-Hunter outperformed baseline models in both AMP classification and MIC prediction, achieving 95.82% accuracy and a 95.80% F1 score on the test set for classification, and an R2 of 0.9245 with an MAE of 0.2305 for MIC prediction. Guided by its predictions, AMP-Forge generated peptide sequences with lower MIC values and improved physicochemical properties associated with antimicrobial activity. Molecular dynamics simulations further provided in silico evidence supporting the antimicrobial potential of selected sequences by identifying stable membrane disruption and insertion behaviors consistent with membrane-targeting activity. Thus, the generation-screening-validation workflow enables reliable discovery of potent AMPs, and provides a practical strategy for rational peptide design, rapid prediction, and translational applications.</p>","PeriodicalId":20241,"journal":{"name":"PLoS Computational Biology","volume":"22 9","pages":"e1014771"},"PeriodicalIF":3.6,"publicationDate":"2026-09-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148887205","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-09-03DOI: 10.1371/journal.pcbi.1014713
Sebastian Towers, Jessica James, Harrison Steel, Idris Kempf
Directed evolution is a method for engineering biological systems or components, such as proteins, wherein desired traits are optimised through iterative rounds of mutagenesis and selection of fit variants. The process of protein directed evolution can be envisaged as navigation over high-dimensional optimisation landscapes with numerous local maxima. The performance of any strategy in navigating such a landscape is dependent on the ruggedness of that landscape. However, this information is generally unavailable at the outset of an experiment. Here we propose SLIDE, Sequence-free Landscape Inference for Directed Evolution, which consists of two parts. First, SLIDE provides an estimation of landscape ruggedness from a mutating population using only population-level phenotypic data and an estimate of the mutation rate. Such ruggedness information in itself is valuable in protein design, for instance in predicting evolutionary stability. Second, SLIDE offers a framework for using the estimated ruggedness metric to identify high-performing selection strategies for directed evolution. Using theoretical NK landscapes and four empirical protein fitness landscapes, we demonstrate consistent in silico improvement upon the performance of fixed-parameter strategies, using a pipeline that could also be combined with emerging AI-based methods for driving directed evolution.
定向进化是一种用于工程生物系统或组件(如蛋白质)的方法,其中通过反复的诱变和选择合适的变体来优化所需的特征。蛋白质定向进化的过程可以设想为在具有许多局部最大值的高维优化景观上的导航。导航这种地形的任何策略的性能都取决于地形的坚固性。然而,在实验开始时,这些信息通常是不可获得的。在此,我们提出了SLIDE (Sequence-free Landscape Inference for Directed Evolution),它由两部分组成。首先,SLIDE仅使用种群水平的表型数据和突变率估计,提供了突变种群的景观坚固性估计。这种坚固性信息本身在蛋白质设计中是有价值的,例如在预测进化稳定性方面。其次,SLIDE提供了一个框架,用于使用估计的坚固度度量来确定定向进化的高性能选择策略。使用理论NK景观和四个经验蛋白质适应度景观,我们证明了固定参数策略性能的一致的硅改进,使用的管道也可以与新兴的基于人工智能的方法相结合,以驱动定向进化。
{"title":"Sequence-free landscape inference for directed evolution.","authors":"Sebastian Towers, Jessica James, Harrison Steel, Idris Kempf","doi":"10.1371/journal.pcbi.1014713","DOIUrl":"https://doi.org/10.1371/journal.pcbi.1014713","url":null,"abstract":"<p><p>Directed evolution is a method for engineering biological systems or components, such as proteins, wherein desired traits are optimised through iterative rounds of mutagenesis and selection of fit variants. The process of protein directed evolution can be envisaged as navigation over high-dimensional optimisation landscapes with numerous local maxima. The performance of any strategy in navigating such a landscape is dependent on the ruggedness of that landscape. However, this information is generally unavailable at the outset of an experiment. Here we propose SLIDE, Sequence-free Landscape Inference for Directed Evolution, which consists of two parts. First, SLIDE provides an estimation of landscape ruggedness from a mutating population using only population-level phenotypic data and an estimate of the mutation rate. Such ruggedness information in itself is valuable in protein design, for instance in predicting evolutionary stability. Second, SLIDE offers a framework for using the estimated ruggedness metric to identify high-performing selection strategies for directed evolution. Using theoretical NK landscapes and four empirical protein fitness landscapes, we demonstrate consistent in silico improvement upon the performance of fixed-parameter strategies, using a pipeline that could also be combined with emerging AI-based methods for driving directed evolution.</p>","PeriodicalId":20241,"journal":{"name":"PLoS Computational Biology","volume":"22 9","pages":"e1014713"},"PeriodicalIF":3.6,"publicationDate":"2026-09-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148887506","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-09-03eCollection Date: 2026-09-01DOI: 10.1371/journal.pcbi.1014672
Teresa Berther, Elio Balestrieri, Martina Saltafossi, Laura Bock Paulsen, Lau M Andersen, Daniel S Kluger
The rapidly developing research field of brain-body neuroscience faces methodological challenges, as analysts continue to develop new analysis strategies in the absence of established best practices. This quest for valid methods is further complicated by the (naturally) circular data involved in the study of phase-locked effects, e.g., in respiration-brain coupling. Various available approaches for phase extraction, constructing adequate surrogate data for statistical comparison, and accounting for the circularity of respiratory data lead to poor cross-study generalisability of results. Interpretation of effects is particularly affected by the problem of multiple comparisons in phase-related inferential statistics. In this tutorial, we propose a robust pipeline for respiration phase-related analyses based on a novel circular extension of cluster-based permutation testing. We highlight and offer guidance on critical parameters in the analysis, systematically compare various approaches being used in the field today, and provide open-access software code for flexible use and future development of our proposed pipeline.
{"title":"Robust circular cluster-based statistics for respiration-brain coupling.","authors":"Teresa Berther, Elio Balestrieri, Martina Saltafossi, Laura Bock Paulsen, Lau M Andersen, Daniel S Kluger","doi":"10.1371/journal.pcbi.1014672","DOIUrl":"10.1371/journal.pcbi.1014672","url":null,"abstract":"<p><p>The rapidly developing research field of brain-body neuroscience faces methodological challenges, as analysts continue to develop new analysis strategies in the absence of established best practices. This quest for valid methods is further complicated by the (naturally) circular data involved in the study of phase-locked effects, e.g., in respiration-brain coupling. Various available approaches for phase extraction, constructing adequate surrogate data for statistical comparison, and accounting for the circularity of respiratory data lead to poor cross-study generalisability of results. Interpretation of effects is particularly affected by the problem of multiple comparisons in phase-related inferential statistics. In this tutorial, we propose a robust pipeline for respiration phase-related analyses based on a novel circular extension of cluster-based permutation testing. We highlight and offer guidance on critical parameters in the analysis, systematically compare various approaches being used in the field today, and provide open-access software code for flexible use and future development of our proposed pipeline.</p>","PeriodicalId":20241,"journal":{"name":"PLoS Computational Biology","volume":"22 9","pages":"e1014672"},"PeriodicalIF":3.6,"publicationDate":"2026-09-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13541119/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148887487","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-09-03DOI: 10.1371/journal.pcbi.1014709
Ivan L A Spirandelli, Arnur Nigmetov, Dmitriy Morozov, Myfanwy E Evans
The simulated assembly of molecular building blocks into functional complexes is central to computational biology and materials science. Protein-assembly simulations, driven by short-range nonpolar interactions, can in principle reach their biologically correct structures, but rugged energy landscapes often trap simulations in non-functional local minima. We introduce a long-range topological potential, quantified by weighted total persistence, and combine it with the morphometric approach to solvation free energy. Across four protein systems, this combination increases assembly success rates by up to sixteen-fold and enables assembly in cases that otherwise fail. Unlike previous topology-based approaches, our method uses topological measures as an active energetic bias rather than a descriptive tool. Depending only on atom geometry, the method extends in principle to other self-assembling systems, offering a general strategy for overcoming kinetic barriers in molecular simulations.
{"title":"Topological potentials guiding protein self-assembly.","authors":"Ivan L A Spirandelli, Arnur Nigmetov, Dmitriy Morozov, Myfanwy E Evans","doi":"10.1371/journal.pcbi.1014709","DOIUrl":"https://doi.org/10.1371/journal.pcbi.1014709","url":null,"abstract":"<p><p>The simulated assembly of molecular building blocks into functional complexes is central to computational biology and materials science. Protein-assembly simulations, driven by short-range nonpolar interactions, can in principle reach their biologically correct structures, but rugged energy landscapes often trap simulations in non-functional local minima. We introduce a long-range topological potential, quantified by weighted total persistence, and combine it with the morphometric approach to solvation free energy. Across four protein systems, this combination increases assembly success rates by up to sixteen-fold and enables assembly in cases that otherwise fail. Unlike previous topology-based approaches, our method uses topological measures as an active energetic bias rather than a descriptive tool. Depending only on atom geometry, the method extends in principle to other self-assembling systems, offering a general strategy for overcoming kinetic barriers in molecular simulations.</p>","PeriodicalId":20241,"journal":{"name":"PLoS Computational Biology","volume":"22 9","pages":"e1014709"},"PeriodicalIF":3.6,"publicationDate":"2026-09-03","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148887730","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-09-02DOI: 10.1371/journal.pcbi.1014733
Dennis Vetter, Muhammad Ahsan, Diana Delicado, Thomas A Neubauer, Thomas Wilke, Gemma Roig
Cryptic species complexes pose fundamental challenges to biologists, as species exhibit minimal morphological differences that require integrating morphology, genetics, and biogeography for identification. Here, we present a deep learning approach to support species identification in the freshwater snail genus Radomaniola (Hydrobiidae), a morphologically cryptic group from the Balkans. Our approach mirrors the integrative workflow of expert taxonomists by combining shell images, morphometric measurements, and collection‑site metadata, with optional phylogenetic information. Despite being trained on fewer than 700 specimens across 20 visually similar species with strongly imbalanced class sizes, the system achieved high identification performance. Careful control of spurious correlations, such as those arising from site‑specific imaging conditions or overly precise geographic metadata, was essential to ensure that the network learned biologically meaningful features. Across all experiments, integrating multiple data types and jointly optimizing meaningful embeddings and classification consistently improved performance over image‑only and classification‑only baselines. On specimens from collection sites seen during training we achieved a macro-averaged F1 score of 0.93. Even though this dropped as low as 0.14 when evaluating on specimens from previously unsampled localities, it could be rapidly recovered by retraining with 2-3 newly labeled specimens. Additionally, model top-3 accuracy stayed consistently above 80% in all settings. These results show that relatively lightweight deep learning models can provide practical decision support in real taxonomic workflows.
{"title":"Speeding up taxonomy in the digital age: A deep learning approach for identifying cryptic freshwater snails.","authors":"Dennis Vetter, Muhammad Ahsan, Diana Delicado, Thomas A Neubauer, Thomas Wilke, Gemma Roig","doi":"10.1371/journal.pcbi.1014733","DOIUrl":"https://doi.org/10.1371/journal.pcbi.1014733","url":null,"abstract":"<p><p>Cryptic species complexes pose fundamental challenges to biologists, as species exhibit minimal morphological differences that require integrating morphology, genetics, and biogeography for identification. Here, we present a deep learning approach to support species identification in the freshwater snail genus Radomaniola (Hydrobiidae), a morphologically cryptic group from the Balkans. Our approach mirrors the integrative workflow of expert taxonomists by combining shell images, morphometric measurements, and collection‑site metadata, with optional phylogenetic information. Despite being trained on fewer than 700 specimens across 20 visually similar species with strongly imbalanced class sizes, the system achieved high identification performance. Careful control of spurious correlations, such as those arising from site‑specific imaging conditions or overly precise geographic metadata, was essential to ensure that the network learned biologically meaningful features. Across all experiments, integrating multiple data types and jointly optimizing meaningful embeddings and classification consistently improved performance over image‑only and classification‑only baselines. On specimens from collection sites seen during training we achieved a macro-averaged F1 score of 0.93. Even though this dropped as low as 0.14 when evaluating on specimens from previously unsampled localities, it could be rapidly recovered by retraining with 2-3 newly labeled specimens. Additionally, model top-3 accuracy stayed consistently above 80% in all settings. These results show that relatively lightweight deep learning models can provide practical decision support in real taxonomic workflows.</p>","PeriodicalId":20241,"journal":{"name":"PLoS Computational Biology","volume":"22 9","pages":"e1014733"},"PeriodicalIF":3.6,"publicationDate":"2026-09-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148881328","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-09-02eCollection Date: 2026-09-01DOI: 10.1371/journal.pcbi.1014699
Nandakishor Krishnan, István Zachar, Ádám Kun, Chaitanya S Gokhale, József Garay
Microbial symbiosis is widespread among metabolically coupled cells; it presumably gave rise to mitochondria. However, how such symbioses emerge, evolve, and stabilize are unknown, particularly in the prokaryotic domain where endosymbiosis is virtually nonexistent. Yet there is growing evidence suggesting that mitochondria originated from such a metabolically driven prokaryotic partnership rather than phagocytotic predation. While prokaryotes almost ubiquitously engage in metabolic syntrophy, it is unknown whether syntrophy alone can enable stable physical associations that could pave the road toward physical integration. Here, we tested the hypothesis that syntrophy can transition into stable ectosymbiosis, using an ecological mathematical model. Starting from an existing syntrophic partnership between free-living hosts and symbionts, we demonstrate that population-level obligate ectosymbiosis can emerge and stabilize, even in unilateral syntrophy where only the symbiont consumes a host-produced metabolite. A key assumption is that the hosts' by-product inhibits their growth when it accumulates. By consuming the toxic by-product, the symbiont locally reduces hosts' self-inhibition at the contact surface, manifesting as a private benefit providing selective advantage. Our results show that due to the direct and indirect benefits, the ectosymbiotic consortium is stable against free-living forms and the consortial cooperation is ecologically selected for. Furthermore, solid metabolic coupling promotes population-level obligacy, ultimately excluding free-living individuals under stricter conditions. Our results support the hypothesis that cooperative, syntrophic microbes (particularly prokaryotes) are capable of forming stable, physical, and species-specific ectosymbiosis through inhibition reduction, providing a plausible first step toward potential, gradual endosymbiotic integration. Our work bridges the gap between models of microbial cooperation between free-living species and models that assume already-concluded, fully integrated endosymbiosis under multilevel selection.
{"title":"Host-initiated microbial association leads to stable ectosymbiosis in an ecological model.","authors":"Nandakishor Krishnan, István Zachar, Ádám Kun, Chaitanya S Gokhale, József Garay","doi":"10.1371/journal.pcbi.1014699","DOIUrl":"10.1371/journal.pcbi.1014699","url":null,"abstract":"<p><p>Microbial symbiosis is widespread among metabolically coupled cells; it presumably gave rise to mitochondria. However, how such symbioses emerge, evolve, and stabilize are unknown, particularly in the prokaryotic domain where endosymbiosis is virtually nonexistent. Yet there is growing evidence suggesting that mitochondria originated from such a metabolically driven prokaryotic partnership rather than phagocytotic predation. While prokaryotes almost ubiquitously engage in metabolic syntrophy, it is unknown whether syntrophy alone can enable stable physical associations that could pave the road toward physical integration. Here, we tested the hypothesis that syntrophy can transition into stable ectosymbiosis, using an ecological mathematical model. Starting from an existing syntrophic partnership between free-living hosts and symbionts, we demonstrate that population-level obligate ectosymbiosis can emerge and stabilize, even in unilateral syntrophy where only the symbiont consumes a host-produced metabolite. A key assumption is that the hosts' by-product inhibits their growth when it accumulates. By consuming the toxic by-product, the symbiont locally reduces hosts' self-inhibition at the contact surface, manifesting as a private benefit providing selective advantage. Our results show that due to the direct and indirect benefits, the ectosymbiotic consortium is stable against free-living forms and the consortial cooperation is ecologically selected for. Furthermore, solid metabolic coupling promotes population-level obligacy, ultimately excluding free-living individuals under stricter conditions. Our results support the hypothesis that cooperative, syntrophic microbes (particularly prokaryotes) are capable of forming stable, physical, and species-specific ectosymbiosis through inhibition reduction, providing a plausible first step toward potential, gradual endosymbiotic integration. Our work bridges the gap between models of microbial cooperation between free-living species and models that assume already-concluded, fully integrated endosymbiosis under multilevel selection.</p>","PeriodicalId":20241,"journal":{"name":"PLoS Computational Biology","volume":"22 9","pages":"e1014699"},"PeriodicalIF":3.6,"publicationDate":"2026-09-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13537694/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148881371","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-09-02DOI: 10.1371/journal.pcbi.1014729
Rupal Chauhan, Biswajit Das, Ajeet K Sharma
E. coli relies on the heat shock response (HSR) to preserve protein homeostasis under stress, through three feedback modules: feedforward translational control, chaperone-mediated sequestration and targeted degradation. Although previous studies have highlighted how this layered architecture ensures rapid and robust protection compared to simpler designs, not much attention is paid to how these modules interact. Moreover, how do interactions among the three modules balance performance trade-offs, where gains in one module may come at the expense of another, yet together yield an optimal overall response? We address this using a mathematical model that integrates protein folding with σ32 regulation. We show that the feedback modules both cooperate and compete, giving rise to nonmonotonic dynamics that govern HSR performance. Specifically, increasing feedforward strength does accelerate response, but beyond a threshold, despite increasing chaperone levels, it paradoxically slows recovery. Similarly, while sequestration enhances relative chaperone production and per-chaperone efficiency, when excessive, it traps σ32 in inactive complexes, prolonging recovery and delaying shutdown. Mapping the parameter space reveals regimes of synergy as well as trade-offs between speed and efficiency, with wild-type parameters lying near the optimal region. These results reveal design principles that produces a robust and efficient heat shock response.
{"title":"Synergies and trade-offs in the heat shock response mechanism.","authors":"Rupal Chauhan, Biswajit Das, Ajeet K Sharma","doi":"10.1371/journal.pcbi.1014729","DOIUrl":"https://doi.org/10.1371/journal.pcbi.1014729","url":null,"abstract":"<p><p>E. coli relies on the heat shock response (HSR) to preserve protein homeostasis under stress, through three feedback modules: feedforward translational control, chaperone-mediated sequestration and targeted degradation. Although previous studies have highlighted how this layered architecture ensures rapid and robust protection compared to simpler designs, not much attention is paid to how these modules interact. Moreover, how do interactions among the three modules balance performance trade-offs, where gains in one module may come at the expense of another, yet together yield an optimal overall response? We address this using a mathematical model that integrates protein folding with σ32 regulation. We show that the feedback modules both cooperate and compete, giving rise to nonmonotonic dynamics that govern HSR performance. Specifically, increasing feedforward strength does accelerate response, but beyond a threshold, despite increasing chaperone levels, it paradoxically slows recovery. Similarly, while sequestration enhances relative chaperone production and per-chaperone efficiency, when excessive, it traps σ32 in inactive complexes, prolonging recovery and delaying shutdown. Mapping the parameter space reveals regimes of synergy as well as trade-offs between speed and efficiency, with wild-type parameters lying near the optimal region. These results reveal design principles that produces a robust and efficient heat shock response.</p>","PeriodicalId":20241,"journal":{"name":"PLoS Computational Biology","volume":"22 9","pages":"e1014729"},"PeriodicalIF":3.6,"publicationDate":"2026-09-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148881290","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}