Pub Date : 2026-09-02DOI: 10.1038/s41597-026-08252-6
Markus Lang, Lotta Mayer
Far-right activism materializes in music events, marches, and violence. Yet, the empirical links between these types of activism remain unclear. Quantifying these manifestations is difficult because of fragmented reporting and inconsistent geocoding at the local level. We introduce the Far-Right Activism (FARAC) dataset, a panel (2013-2024) that aggregates German federal parliamentary inquiries and civil society records at the county level. The dataset addresses the problem of regional assignment by standardizing all event records in accordance with 2024 administrative boundaries. This harmonized structure permits the first systematic, local-level comparison of far-right activism. It uses standard administrative identifiers, enabling direct integration with electoral results, sociodemographic indicators, and other county-level datasets.
{"title":"Far-right music events, protests, and violence: a county-level dataset for Germany, 2013-2024.","authors":"Markus Lang, Lotta Mayer","doi":"10.1038/s41597-026-08252-6","DOIUrl":"10.1038/s41597-026-08252-6","url":null,"abstract":"<p><p>Far-right activism materializes in music events, marches, and violence. Yet, the empirical links between these types of activism remain unclear. Quantifying these manifestations is difficult because of fragmented reporting and inconsistent geocoding at the local level. We introduce the Far-Right Activism (FARAC) dataset, a panel (2013-2024) that aggregates German federal parliamentary inquiries and civil society records at the county level. The dataset addresses the problem of regional assignment by standardizing all event records in accordance with 2024 administrative boundaries. This harmonized structure permits the first systematic, local-level comparison of far-right activism. It uses standard administrative identifiers, enabling direct integration with electoral results, sociodemographic indicators, and other county-level datasets.</p>","PeriodicalId":21597,"journal":{"name":"Scientific Data","volume":"13 1","pages":""},"PeriodicalIF":7.2,"publicationDate":"2026-09-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13538395/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148881490","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-09-01DOI: 10.1038/s41597-026-08216-w
Rico Hiemann, Dirk Reinhold, Karsten Conrad, Stefan Rödiger, Peter Schierack, Dirk Roggenbuck
We introduce a large-scale single cell dataset for cell cycle phases. The dataset includes single cell images of 4',6-diamidino-2-phenylindole-stained human epithelioma-2 (HEp-2) cells grown on microscopic glass slides. Greyscale images of adherent HEp-2 cells were acquired at 20x magnification by an automated, inverse microscope system. Cell images were automatically segmented with 256 × 256 pixel patch size including surrounding area. Initially, 10,000 images were manually classified by experts into 5 cell-cycle groups (interphase, prophase, metaphase, anaphase, telophase) and two additional groups (artefact, triple). Based on these initial single cell-image classifications, a convolutional neural network (CNN) model (VGG16) was trained. Subsequently, additional segmented single cell images were pre-classified by the CNN and double-blinded reviewed as well as corrected by two experts to build an algorithm-aided dataset with 100,000 images. The dataset provides a resource for development of biomedical and deep learning applications and, thus, enables training and assessment of cell-cycle detection algorithms.
{"title":"A human epithelial cell cycle dataset.","authors":"Rico Hiemann, Dirk Reinhold, Karsten Conrad, Stefan Rödiger, Peter Schierack, Dirk Roggenbuck","doi":"10.1038/s41597-026-08216-w","DOIUrl":"10.1038/s41597-026-08216-w","url":null,"abstract":"<p><p>We introduce a large-scale single cell dataset for cell cycle phases. The dataset includes single cell images of 4',6-diamidino-2-phenylindole-stained human epithelioma-2 (HEp-2) cells grown on microscopic glass slides. Greyscale images of adherent HEp-2 cells were acquired at 20x magnification by an automated, inverse microscope system. Cell images were automatically segmented with 256 × 256 pixel patch size including surrounding area. Initially, 10,000 images were manually classified by experts into 5 cell-cycle groups (interphase, prophase, metaphase, anaphase, telophase) and two additional groups (artefact, triple). Based on these initial single cell-image classifications, a convolutional neural network (CNN) model (VGG16) was trained. Subsequently, additional segmented single cell images were pre-classified by the CNN and double-blinded reviewed as well as corrected by two experts to build an algorithm-aided dataset with 100,000 images. The dataset provides a resource for development of biomedical and deep learning applications and, thus, enables training and assessment of cell-cycle detection algorithms.</p>","PeriodicalId":21597,"journal":{"name":"Scientific Data","volume":"13 1","pages":""},"PeriodicalIF":7.2,"publicationDate":"2026-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13534591/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148876053","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-31DOI: 10.1038/s41597-026-08133-y
Malte Brammerloh, Anneke Alkemade, Pierre-Louis Bazin, Caroline Jantzen, Carsten Jäger, Aurora Gasparello, Mikhail Zubkov, Puneet Talwar, Gilles Vandewalle, Andreas Herrler, Kerrin J Pine, Markus Morawski, Rawien Balesar, Katrin Amunts, Birte U Forstmann, Nikolaus Weiskopf, Evgeniya Kirilina
Nigrosomes are formed by clusters of pigmented dopaminergic cells in the substantia nigra that critically contribute to dopaminergic function. The ever-increasing resolution of ultra-high-field MRI brings clinical imaging of these clusters into reach, promising unprecedented insight into the functional role of the nigrosomes and their early degeneration in Parkinson's disease. However, due to the nigrosomes' small extents and intricate shapes, they are not included in current MRI brain atlases, preventing nigrosome-specific MRI data analysis. We provide a comprehensive 3D histological atlas of the five nigrosomes co-aligned to the widely-used MNI152 2009b space. This atlas is based on 3D-reconstructed, ultra-high-resolution block-face images and gold-standard nigrosome delineations in calbindin-D28K immunohistochemistry. We validated the atlas's accuracy using the multimodal ultra-high-resolution post mortem BigBrain dataset and demonstrated its consistency with qualitative nigrosome atlases based on classical 2D histology. We provide detailed usage instructions for applying our atlas to ultra-high-resolution and -field MRI data. The openly available atlas enables neuroimaging studies of the nigrosomes, opening a new avenue toward understanding the differential involvement of the nigrosomes in the healthy and diseased brain and the development of neuroimaging biomarkers of dopaminergic neurodegeneration.
{"title":"A neuroimaging atlas of the nigrosomes in the substantia nigra based on 3D histology.","authors":"Malte Brammerloh, Anneke Alkemade, Pierre-Louis Bazin, Caroline Jantzen, Carsten Jäger, Aurora Gasparello, Mikhail Zubkov, Puneet Talwar, Gilles Vandewalle, Andreas Herrler, Kerrin J Pine, Markus Morawski, Rawien Balesar, Katrin Amunts, Birte U Forstmann, Nikolaus Weiskopf, Evgeniya Kirilina","doi":"10.1038/s41597-026-08133-y","DOIUrl":"10.1038/s41597-026-08133-y","url":null,"abstract":"<p><p>Nigrosomes are formed by clusters of pigmented dopaminergic cells in the substantia nigra that critically contribute to dopaminergic function. The ever-increasing resolution of ultra-high-field MRI brings clinical imaging of these clusters into reach, promising unprecedented insight into the functional role of the nigrosomes and their early degeneration in Parkinson's disease. However, due to the nigrosomes' small extents and intricate shapes, they are not included in current MRI brain atlases, preventing nigrosome-specific MRI data analysis. We provide a comprehensive 3D histological atlas of the five nigrosomes co-aligned to the widely-used MNI152 2009b space. This atlas is based on 3D-reconstructed, ultra-high-resolution block-face images and gold-standard nigrosome delineations in calbindin-D28K immunohistochemistry. We validated the atlas's accuracy using the multimodal ultra-high-resolution post mortem BigBrain dataset and demonstrated its consistency with qualitative nigrosome atlases based on classical 2D histology. We provide detailed usage instructions for applying our atlas to ultra-high-resolution and -field MRI data. The openly available atlas enables neuroimaging studies of the nigrosomes, opening a new avenue toward understanding the differential involvement of the nigrosomes in the healthy and diseased brain and the development of neuroimaging biomarkers of dopaminergic neurodegeneration.</p>","PeriodicalId":21597,"journal":{"name":"Scientific Data","volume":"13 1","pages":""},"PeriodicalIF":7.2,"publicationDate":"2026-08-31","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13529644/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148866493","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-29DOI: 10.1038/s41597-026-08179-y
Ji Shao, Jing Cao, Changjun Wang, Peifang Xu, Xuan Zhang, Yiming Sun, Pengjie Chen, Ningxin Dai, Lixia Lou, Juan Ye
Abnormal eyelid position and morphology can cause visual dysfunction, ocular surface damage, and facial deformities, highlighting the need for accurate eyelid assessment. Advances in artificial intelligence (AI) have enabled automated analysis of eyelid abnormalities using external eye images, but development is limited by the lack of publicly available datasets with multi-category disorders and standardized structural annotations. To address this gap, we constructed a clinically annotated eye image dataset of 1,414 images from eight common eyelid disorders and a normal control group. Each image includes diagnostic labels and manual annotations of three key periocular structures: eyelid fissure, cornea, and eyebrow. Image quality evaluation and expert verification were performed to ensure annotation reliability. Inter- and intra-annotator consistency demonstrated excellent agreement. A baseline Attention 2D U-Net segmentation model trained on the dataset achieved Dice coefficients of 0.93, 0.96, and 0.89 for the eyelid fissure, cornea, and eyebrow, respectively. This dataset provides a valuable resource for automated segmentation, quantitative periocular measurement, and AI-assisted eyelid analysis, supporting standardized and reproducible approaches for eyelid assessment.
{"title":"A multi-category eye image dataset for AI-based segmentation and analysis of eyelid disorders.","authors":"Ji Shao, Jing Cao, Changjun Wang, Peifang Xu, Xuan Zhang, Yiming Sun, Pengjie Chen, Ningxin Dai, Lixia Lou, Juan Ye","doi":"10.1038/s41597-026-08179-y","DOIUrl":"https://doi.org/10.1038/s41597-026-08179-y","url":null,"abstract":"<p><p>Abnormal eyelid position and morphology can cause visual dysfunction, ocular surface damage, and facial deformities, highlighting the need for accurate eyelid assessment. Advances in artificial intelligence (AI) have enabled automated analysis of eyelid abnormalities using external eye images, but development is limited by the lack of publicly available datasets with multi-category disorders and standardized structural annotations. To address this gap, we constructed a clinically annotated eye image dataset of 1,414 images from eight common eyelid disorders and a normal control group. Each image includes diagnostic labels and manual annotations of three key periocular structures: eyelid fissure, cornea, and eyebrow. Image quality evaluation and expert verification were performed to ensure annotation reliability. Inter- and intra-annotator consistency demonstrated excellent agreement. A baseline Attention 2D U-Net segmentation model trained on the dataset achieved Dice coefficients of 0.93, 0.96, and 0.89 for the eyelid fissure, cornea, and eyebrow, respectively. This dataset provides a valuable resource for automated segmentation, quantitative periocular measurement, and AI-assisted eyelid analysis, supporting standardized and reproducible approaches for eyelid assessment.</p>","PeriodicalId":21597,"journal":{"name":"Scientific Data","volume":"13 1","pages":""},"PeriodicalIF":7.2,"publicationDate":"2026-08-29","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148892392","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-29DOI: 10.1038/s41597-026-07899-5
Markus Glaß, Stefan Hüttelmaier
RNA binding proteins (RBPs) are key post-transcriptional regulators controlling every aspect of the RNA life cycle from synthesis to decay. We extracted RBP target genes from publicly available enhanced cross-linking and immuno-precipitation followed by sequencing (eCLIP-seq) data of 168 RBPs and assembled a gene set collection that can be used to examine gene lists for enriched RBP targets via functional enrichment analysis methods like over-representation analysis (ORA), gene set enrichment analysis (GSEA) or gene set variation analysis (GSVA).
{"title":"A gene set collection of human RBP target genes.","authors":"Markus Glaß, Stefan Hüttelmaier","doi":"10.1038/s41597-026-07899-5","DOIUrl":"10.1038/s41597-026-07899-5","url":null,"abstract":"<p><p>RNA binding proteins (RBPs) are key post-transcriptional regulators controlling every aspect of the RNA life cycle from synthesis to decay. We extracted RBP target genes from publicly available enhanced cross-linking and immuno-precipitation followed by sequencing (eCLIP-seq) data of 168 RBPs and assembled a gene set collection that can be used to examine gene lists for enriched RBP targets via functional enrichment analysis methods like over-representation analysis (ORA), gene set enrichment analysis (GSEA) or gene set variation analysis (GSVA).</p>","PeriodicalId":21597,"journal":{"name":"Scientific Data","volume":"13 1","pages":""},"PeriodicalIF":7.2,"publicationDate":"2026-08-29","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13526010/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148857633","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-29DOI: 10.1038/s41597-026-08151-w
Helge Bruelheide, Karl Andraczek, Aline B Bombo, Gabriella Damasceno, Ying Fan, Grégoire T Freschet, Lena Hartmann, Justus Hennecke, Cody Coyotee Howard, Saheed Olaide Jimoh, Jitka Klimešová, Daniel C Laughlin, Liesje Mommer, Tsumbedzo Ramalevha, Frances Siebert, Shersingh Joseph Tumber-Dávila, Alexandra Weigelt, David Atkins, Gianmaria Bonari, Trevor A Carter, Jiří Doležal, Timothy Harris, Borja Jiménez-Alfaro, Austin R Kelly, F Curtis Lubbe, Jana Martínková, Mathieu Millan, Jacqueline Ott, Steve Orzell, Jianqiang Qian, Alice E Stears, Dennis Whigham, Xin Jing, Joana Bergmann, Alessandra Fidelis
Belowground functional diversity is relevant to numerous ecosystem functions and to ecosystem resilience under global change, yet quantitative information on belowground plant traits is disparate and poorly integrated across fields of research. So far, belowground plant traits have been studied in three disparate research domains: (i) fine root traits in the context of belowground resource economics and symbioses; (ii) maximum rooting depth and lateral extent of the root system in the context of overall plant allometry and resource uptake; (iii) clonal organs and bud banks with a focus on plant and community resilience. However, there remains a major disconnection among these three fields of research. A prerequisite for linking them together is the creation of a comprehensive, curated, open-access dataset available for comprehensive analyses. Here, we compiled and harmonised such a trait dataset for 10,453 vascular plant species and 19 belowground plant traits, which we named UNDERPLOT. Based on these data, we defined two integrative indices, a Belowground Persistence Type (BPT) and a Clonal Spread Index (CSI), which are suitable trait-based indicators for predicting resistance and resilience of communities to disturbance. Further applications will increase our knowledge of the dimensionality of the belowground trait space, adding new insights into trait-environment relationships, vegetation responses to climate change and disturbance, as well as a deeper understanding of evolutionary first principles.
{"title":"A global dataset of key traits for plant belowground functioning.","authors":"Helge Bruelheide, Karl Andraczek, Aline B Bombo, Gabriella Damasceno, Ying Fan, Grégoire T Freschet, Lena Hartmann, Justus Hennecke, Cody Coyotee Howard, Saheed Olaide Jimoh, Jitka Klimešová, Daniel C Laughlin, Liesje Mommer, Tsumbedzo Ramalevha, Frances Siebert, Shersingh Joseph Tumber-Dávila, Alexandra Weigelt, David Atkins, Gianmaria Bonari, Trevor A Carter, Jiří Doležal, Timothy Harris, Borja Jiménez-Alfaro, Austin R Kelly, F Curtis Lubbe, Jana Martínková, Mathieu Millan, Jacqueline Ott, Steve Orzell, Jianqiang Qian, Alice E Stears, Dennis Whigham, Xin Jing, Joana Bergmann, Alessandra Fidelis","doi":"10.1038/s41597-026-08151-w","DOIUrl":"10.1038/s41597-026-08151-w","url":null,"abstract":"<p><p>Belowground functional diversity is relevant to numerous ecosystem functions and to ecosystem resilience under global change, yet quantitative information on belowground plant traits is disparate and poorly integrated across fields of research. So far, belowground plant traits have been studied in three disparate research domains: (i) fine root traits in the context of belowground resource economics and symbioses; (ii) maximum rooting depth and lateral extent of the root system in the context of overall plant allometry and resource uptake; (iii) clonal organs and bud banks with a focus on plant and community resilience. However, there remains a major disconnection among these three fields of research. A prerequisite for linking them together is the creation of a comprehensive, curated, open-access dataset available for comprehensive analyses. Here, we compiled and harmonised such a trait dataset for 10,453 vascular plant species and 19 belowground plant traits, which we named UNDERPLOT. Based on these data, we defined two integrative indices, a Belowground Persistence Type (BPT) and a Clonal Spread Index (CSI), which are suitable trait-based indicators for predicting resistance and resilience of communities to disturbance. Further applications will increase our knowledge of the dimensionality of the belowground trait space, adding new insights into trait-environment relationships, vegetation responses to climate change and disturbance, as well as a deeper understanding of evolutionary first principles.</p>","PeriodicalId":21597,"journal":{"name":"Scientific Data","volume":"13 1","pages":""},"PeriodicalIF":7.2,"publicationDate":"2026-08-29","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13526019/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148857589","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-29DOI: 10.1038/s41597-026-08215-x
Markus Erhard Schorn, Jürgen Bauhus, Christian Wirth, Nadja Rüger
This dataset contains forest inventory data from four censuses (1999, 2007, 2013 and 2023) from a single ~28 ha permanent monitoring plot in a beech-dominated forest that has been unmanaged since 1965 at Hainich National Park, Thuringia, Germany. For each census, the location, species identity and diameter at breast height (dbh) of all living trees with a dbh ≥1 cm was recorded. Additional attributes for living trees include vitality, crown coverage by neighboring trees, sociological class, and observed damages. For trees that died since the previous census, information on cause of mortality, degree of decay and stem integrity is provided for some censuses. This dataset enables analyses of the dynamics of unmanaged Central European beech-dominated forests over more than two decades.
{"title":"Repeated large-scale forest inventory data from an unmanaged European beech forest at Hainich National Park, Germany.","authors":"Markus Erhard Schorn, Jürgen Bauhus, Christian Wirth, Nadja Rüger","doi":"10.1038/s41597-026-08215-x","DOIUrl":"10.1038/s41597-026-08215-x","url":null,"abstract":"<p><p>This dataset contains forest inventory data from four censuses (1999, 2007, 2013 and 2023) from a single ~28 ha permanent monitoring plot in a beech-dominated forest that has been unmanaged since 1965 at Hainich National Park, Thuringia, Germany. For each census, the location, species identity and diameter at breast height (dbh) of all living trees with a dbh ≥1 cm was recorded. Additional attributes for living trees include vitality, crown coverage by neighboring trees, sociological class, and observed damages. For trees that died since the previous census, information on cause of mortality, degree of decay and stem integrity is provided for some censuses. This dataset enables analyses of the dynamics of unmanaged Central European beech-dominated forests over more than two decades.</p>","PeriodicalId":21597,"journal":{"name":"Scientific Data","volume":"13 1","pages":""},"PeriodicalIF":7.2,"publicationDate":"2026-08-29","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13526015/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148857602","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-26DOI: 10.1038/s41597-026-07803-1
Stefano M Iacus, Devika Jain, Andrea Nasuto, Giuseppe Porro, Marcello Carammia, Andrea Vezzulli
Quantifying human flourishing-a multidimensional construct including happiness, health, purpose, virtue, relationships, and financial stability-is critical for understanding societal well-being beyond economic indicators. Existing measures often lack fine spatial and temporal resolution. Here we introduce the Human Flourishing Geographic Index (HFGI), derived from analyzing approximately 2.6 billion geolocated U.S. tweets (2013-2023) using fine-tuned large language models to classify expressions across 48 indicators aligned with Harvard's Global Flourishing Study framework plus attitudes towards migration and perception of corruption. The dataset offers monthly and yearly county- and state-level indicators of flourishing-related discourse, validated to confirm that the measures accurately represent the underlying constructs and show expected correlations with established indicators. This resource enables multidisciplinary analyses of well-being, inequality, and social change at unprecedented resolution, offering insights into the dynamics of human flourishing as reflected in social media discourse across the United States over the past decade.
{"title":"The Human Flourishing Geographic Index: A County-Level Dataset for the United States, 2013-2023.","authors":"Stefano M Iacus, Devika Jain, Andrea Nasuto, Giuseppe Porro, Marcello Carammia, Andrea Vezzulli","doi":"10.1038/s41597-026-07803-1","DOIUrl":"10.1038/s41597-026-07803-1","url":null,"abstract":"<p><p>Quantifying human flourishing-a multidimensional construct including happiness, health, purpose, virtue, relationships, and financial stability-is critical for understanding societal well-being beyond economic indicators. Existing measures often lack fine spatial and temporal resolution. Here we introduce the Human Flourishing Geographic Index (HFGI), derived from analyzing approximately 2.6 billion geolocated U.S. tweets (2013-2023) using fine-tuned large language models to classify expressions across 48 indicators aligned with Harvard's Global Flourishing Study framework plus attitudes towards migration and perception of corruption. The dataset offers monthly and yearly county- and state-level indicators of flourishing-related discourse, validated to confirm that the measures accurately represent the underlying constructs and show expected correlations with established indicators. This resource enables multidisciplinary analyses of well-being, inequality, and social change at unprecedented resolution, offering insights into the dynamics of human flourishing as reflected in social media discourse across the United States over the past decade.</p>","PeriodicalId":21597,"journal":{"name":"Scientific Data","volume":"13 1","pages":""},"PeriodicalIF":7.2,"publicationDate":"2026-08-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13518848/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148832058","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Pub Date : 2026-08-26DOI: 10.1038/s41597-026-07806-y
Brett Pike, Anders Gonçalves da Silva, Wilson Terán
We present 4 haplotype-resolved, chromosome scale diploid assemblies of cannabis, assembled from ONT R9.4.1 reads via a novel pipeline based on Hi-C phasing. These assemblies, while low in QV, offer contiguity and genic content comparable to recent HiFi assemblies. Along with a trio-binned assembly previously produced by us and 56 haplotypes selected from the Salk Institute Cannabis Pangenome project, we use these assemblies to create a reference-free pangenome graph. Within a total length of 6.48 Gb, it contains 162.14 M nodes, 228.27 M edges, 14.87 M SNPs, and 6.40 M indels. By optimizing parameters within the Pangenome Graph Builder (PGGB), we avoid many spurious connections among repeat elements, reduce processing time, and arrive at a data structure that visibly recapitulates the linear nature of plant chromosomes. Via k-mer analysis, we corroborate that more genotypes are needed to close the cannabis pangenome, and that, in particular, the region of origin likely remains undersampled.
{"title":"Four haplotype-resolved genome assemblies and a reference-free 66-haplotype pangenome graph for Cannabis sativa.","authors":"Brett Pike, Anders Gonçalves da Silva, Wilson Terán","doi":"10.1038/s41597-026-07806-y","DOIUrl":"10.1038/s41597-026-07806-y","url":null,"abstract":"<p><p>We present 4 haplotype-resolved, chromosome scale diploid assemblies of cannabis, assembled from ONT R9.4.1 reads via a novel pipeline based on Hi-C phasing. These assemblies, while low in QV, offer contiguity and genic content comparable to recent HiFi assemblies. Along with a trio-binned assembly previously produced by us and 56 haplotypes selected from the Salk Institute Cannabis Pangenome project, we use these assemblies to create a reference-free pangenome graph. Within a total length of 6.48 Gb, it contains 162.14 M nodes, 228.27 M edges, 14.87 M SNPs, and 6.40 M indels. By optimizing parameters within the Pangenome Graph Builder (PGGB), we avoid many spurious connections among repeat elements, reduce processing time, and arrive at a data structure that visibly recapitulates the linear nature of plant chromosomes. Via k-mer analysis, we corroborate that more genotypes are needed to close the cannabis pangenome, and that, in particular, the region of origin likely remains undersampled.</p>","PeriodicalId":21597,"journal":{"name":"Scientific Data","volume":"13 1","pages":""},"PeriodicalIF":7.2,"publicationDate":"2026-08-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13518974/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148832016","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}
Mastigias papua is a representative species of the Scyphozoa class. Its outer umbrella surface is covered with white spots and it is mainly distributed in the Pacific and Indian Ocean. In this study, we generated a haplotype-resolved, chromosome-scale genome assembly for M. papua using PacBio HiFi long reads and Hi-C technology. The resulting assembly contains two haplotypes (Hap A and Hap B) with sizes of 379.81 MB (contig N50 = 13.03 Mb) and 344.57 MB (contig N50 = 14.24 Mb), and both anchored to 21 chromosomes with anchor ratio of 94.35% and 97.59%, respectively. The sequencing depth, mapping coverage, contig continuity, and BUSCO assessment collectively indicate a high-quality haplotype-resolved genome assembly. This high-quality genome assembly provides a valuable resource for further genetic studies and genetic improvement of the group of M. papua. BUSCO assessment indicated that the completeness of the two haploid genomes was 93.3% (HapA) and 92.2% (HapB), respectively, demonstrating excellent assembly quality. Annotation results identified 27,401 protein-coding genes and 182.83 Mb of repetitive sequences (accounting for 48.14% of the genome) in HapA, while HapB contained 23,905 protein-coding genes and 158.00 Mb of repetitive sequences (representing 45.85% of the genome). Functional annotation revealed that over 88% of the genes could be matched to public databases such as NR, KEGG, and GO. These phased genome assemblies and gene annotations provide a quality-controlled genomic resource for Mastigias papua.
{"title":"Haplotype-resolved chromosomal-level genome assembly of Mastigias papua.","authors":"Bingbing Li, Xiuxiu Wang, Shirong Lv, Xuelu Yu, Wenqi Zhu, Xue Wang, Yingying Feng, Nanmei Liu, Ying He, Jishun Yang","doi":"10.1038/s41597-026-07804-0","DOIUrl":"10.1038/s41597-026-07804-0","url":null,"abstract":"<p><p>Mastigias papua is a representative species of the Scyphozoa class. Its outer umbrella surface is covered with white spots and it is mainly distributed in the Pacific and Indian Ocean. In this study, we generated a haplotype-resolved, chromosome-scale genome assembly for M. papua using PacBio HiFi long reads and Hi-C technology. The resulting assembly contains two haplotypes (Hap A and Hap B) with sizes of 379.81 MB (contig N50 = 13.03 Mb) and 344.57 MB (contig N50 = 14.24 Mb), and both anchored to 21 chromosomes with anchor ratio of 94.35% and 97.59%, respectively. The sequencing depth, mapping coverage, contig continuity, and BUSCO assessment collectively indicate a high-quality haplotype-resolved genome assembly. This high-quality genome assembly provides a valuable resource for further genetic studies and genetic improvement of the group of M. papua. BUSCO assessment indicated that the completeness of the two haploid genomes was 93.3% (HapA) and 92.2% (HapB), respectively, demonstrating excellent assembly quality. Annotation results identified 27,401 protein-coding genes and 182.83 Mb of repetitive sequences (accounting for 48.14% of the genome) in HapA, while HapB contained 23,905 protein-coding genes and 158.00 Mb of repetitive sequences (representing 45.85% of the genome). Functional annotation revealed that over 88% of the genes could be matched to public databases such as NR, KEGG, and GO. These phased genome assemblies and gene annotations provide a quality-controlled genomic resource for Mastigias papua.</p>","PeriodicalId":21597,"journal":{"name":"Scientific Data","volume":"13 1","pages":""},"PeriodicalIF":7.2,"publicationDate":"2026-08-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC13518791/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"148832134","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}