-
[hal-05105798] Metabolic Flux Inference in a Cheese Microbial Community via comFI: a Biology-informed Approach for Time-resolved Multi-omics Integration
Microbial communities play a central role in many bioprocesses with key applications in food fermentation, waste treatment, human and animal well-being, plant protection or metabolite transformation in industrial bioprocesses. However, the metabolic microbial interactions driving the community dynamics remain difficult to characterize because of their complexity and their temporal variability. Recent advances in sequencing and analytical technologies now provide time-resolved multi-omics data at the community scale providing key insights into the mechanisms shaping the community dynamics. However, integrating these heterogeneous data in an interpretable way to decipher species-specific metabolic activity and microbial interactions remains a major challenge in the study of microbial communities. We introduce the community metabolic flux inference (comFI) method, a mathematical framework for inferring the metabolic fluxes of individual microorganisms from community-level longitudinal data. The method formulates flux estimation as a biology-informed constrained inference problem that combines observed microbial abundances and extracellular metabolite exchange data, with metabolic constraints, encoded in a metabolic model, and transcriptomic-based lasso regularization terms. We evaluated comFI on synthetic datasets generated from dynamic models of microbial communities involving three Escherichia coli mutant strains. The comFI method showed a very good reconstruction accuracy for exchange fluxes, intracellular metabolic fluxes distribution, metabolic pathway activation patterns and strain contribution. We also applied the method to experimental cheese fermentation data involving three bacteria (Lactococcus lactis, Lactobacillus plantarum, and Propionibacterium freudenreichii ), combining abundance measurements, targeted metabolomics and transcriptomics data. The comFI framework enabled to recover previously identified interaction patterns, and to reconstruct latent intracellular flux states for individual microorganisms alongside with their respective metabolic contributions within the community, consistently with the omics data. All together, we demonstrate that comFI provides a practical framework for recovering the metabolic activity of individual microorganisms from community-scale multi-omics time-resolved data.
ano.nymous@ccsd.cnrs.fr.invalid (Sthyve Junior Tatho Djeanou) 24 May 2026
https://hal.science/hal-05105798v2
-
[hal-05708825] Metabolic Flux Inference in a Cheese Microbial Community via comFI: a Biology-informed Approach for Time-resolved Multi-omics Integration
Microbial communities play a central role in many bioprocesses with key applications in food fermentation, waste treatment, human and animal well-being, plant protection or metabolite transformation in industrial bioprocesses. However, the metabolic microbial interactions driving the community dynamics remain difficult to characterize because of their complexity and their temporal variability. Recent advances in sequencing and analytical technologies now provide time-resolved multi-omics data at the community scale providing key insights into the mechanisms shaping the community dynamics. However, integrating these heterogeneous data in an interpretable way to decipher species-specific metabolic activity and microbial interactions remains a major challenge in the study of microbial communities. We introduce the community metabolic flux inference (comFI) method, a mathematical framework for inferring the metabolic fluxes of individual microorganisms from community-level longitudinal data. The method formulates flux estimation as a biology-informed constrained inference problem that combines observed microbial abundances and extracellular metabolite exchange data, with metabolic constraints encoded in a metabolic model and transcriptomic-based lasso regularization terms. We evaluated comFI on synthetic datasets generated from dynamic models of microbial communities involving three Escherichia coli mutant strains. The comFI method showed a very good reconstruction accuracy for exchange fluxes, intracellular metabolic fluxes distribution, metabolic pathway activation patterns and strain contribution. We also applied the method to experimental cheese fermentation data involving three bacteria (Lactococcus lactis, Lactobacillus plantarum, and Propionibacterium freudenreichii), combining abundance measurements, targeted metabolomics and metatranscriptomics data. The comFI framework enabled to recover previously identified interaction patterns, and to reconstruct latent intracellular flux states for individual microorganisms alongside with their respective metabolic contributions within the community, consistently with the omics data and known physiology. All together, we demonstrate that comFI provides a practical framework for recovering the metabolic activity of individual microorganisms from community-scale multi-omics time-resolved data.
ano.nymous@ccsd.cnrs.fr.invalid (Sthyve Junior Tatho Djeanou) 31 Jul 2026
https://hal.science/hal-05708825v1
-
[hal-05736639] On the role of geometric representations for ultra-fine-grained crop variety identification
Distinguishing plant varieties within a single species, an Ultra-Fine-Grained Visual Categorization (Ultra-FGVC) problem, remains largely unsolved under real-world agricultural conditions. While foundation models have significantly improved species-level plant identification, their ability to capture the subtle phenotypic variations that define cultivars remains unclear.<p>In this work, we investigate which representation learning paradigm is best suited for infraspecific plant identification. We compare two approaches: (i) geometric self-supervised learning (DINOv2/v3), which learns representations by enforcing consistency across different sub-regions (patches) of the same image, and (ii) semantic alignment (BioCLIP 2), which aligns images with taxonomic textual descriptions. We further evaluate a hybrid strategy combining geometric pre-training with species-level supervised fine-tuning (Pl@ntNet), aiming to leverage the complementary strengths of both approaches. Experiments are conducted across three complementary datasets: a controlled market garden dataset (22 varieties), a participatory banana dataset (87 classes), and a grape variety dataset exhibiting a strong domain shift (34 varieties).</p><p>Our results show that geometric self-supervised representations provide the most suitable foundation for ultra-fine-grained recognition, consistently outperforming semantic alignment approaches. While species-level supervision yields modest improvements, its contribution remains limited compared to the gains obtained from geometric pre-training, indicating that the discriminative signal at the varietal level is primarily morphological rather than semantic. Furthermore, domain-specific adaptation provides small but consistent gains, and naïve multi-organ aggregation can dilute discriminative signals in asymmetric taxa such as Musa (bananas).</p><p>These findings demonstrate that ultra-fine-grained crop variety identification relies primarily on geometric sensitivity, and highlight the importance of designing models and aggregation strategies that preserve fine-scale morphological information for agrobiodiversity monitoring.</p>
ano.nymous@ccsd.cnrs.fr.invalid (Charlotte Tibi) 03 Sep 2026
https://inria.hal.science/hal-05736639v1
-
[hal-05652708] Text-to-MDX: LLM-assisted generation of MDX queries from user questions
MDX (MultiDimensional Expressions) is the standard language for querying multidimensional data in OLAP systems, but its complex syntax poses challenges for non-expert users. While a lot of research has focused on natural language interfaces for SQL, little attention has been given to MDX. This paper explores the potential of Large Language Models (LLMs), specifically GPT-4o, in translating natural language questions into MDX statements. We investigate whether LLMs can act as full MDX query generators or assistants, and study how the writing style of questions affects output correctness. Through four research questions, we evaluate ChatGPT’s basic capabilities and the effectiveness of prompt engineering in improving text-to-MDX performance. Our evaluation confirms that, with ad-hoc prompt engineering, GPT-4o is indeed able to generate complex MDX queries —particularly when the natural language question is given a structured formulation.
ano.nymous@ccsd.cnrs.fr.invalid (Sandro Bimonte) 10 Jun 2026
https://hal.inrae.fr/hal-05652708v1
-
[hal-05780588] Overview of PestCLEF 2026: Knowledge Graph Extraction from Plant Health Literature LifeCLEF at CLEF 2026
<div><p>The first edition of PestCLEF, organized within LifeCLEF, challenges participants to extract knowledge graphs from Web documents in the domain of plant health. Understanding and monitoring crop disease transmission is essential for ensuring food security, economic stability, and environmental sustainability, yet relevant knowledge is dispersed across diverse sources that use varying terminology and perspectives. PestCLEF aims to advance largescale, standardized extraction of crop disease knowledge from documents, enabling improved data integration, epidemiological surveillance, and cross-disciplinary understanding of plant disease mechanisms.</p><p>The PestCLEF data was built upon the EPOP corpus of 247 documents collected and annotated by the French plant epidemiomonitoring platform. The participants were required to extract the knowledge graph representing the domain knowledge and events, such as the cause of diseases, affected crops, or places and dates of occurrences of pest and vector species. PestCLEF pushes forward Information Extraction challenges with complex entities and long-range relations. The submissions were evaluated with the F-score computed by comparing the reference and predicted knowledge graph edges. The matching of edges in the evaluation process ensures that predicted edges are accounted for regardless of the node forms chosen by generative models.</p><p>Twelve teams participated in the PestCLEF challenge of which four outperformed the provided baseline. This paper describes 1) a detailed description of the challenge motivation, data, evaluation protocol, and baseline, 2) a description and analysis of the participating systems, and 3) a discussion on PestCLEF outcomes and future.</p></div>
ano.nymous@ccsd.cnrs.fr.invalid (Robert Bossy) 06 Oct 2026
https://hal.inrae.fr/hal-05780588v1
-
[hal-05778771] FrontVeg V2: A Training-Free Software Framework for Foreground-Aware Zero-Shot Plant Trait Segmentation in High-Resolution Images of Trellised Crops
FrontVeg V2 is an open-source, training-free software framework for foregroundaware zero-shot segmentation of plant traits in high-resolution images of trellised crops. The pipeline combines monocular depth estimation, automatic foreground extraction using Valley-Aware Depth Thresholding, tiled zero-shot segmentation, Graph-Based Mask Assembly, and geometry-aware fusion. This design enables plant organs and disease symptoms to be segmented while reducing detections arising from neighboring vegetation rows. The current implementation integrates Depth Anything V2 (DAV2) and SAM3 and can be used through both command-line batch processing and a Napari graphical interface. FrontVeg V2 provides a reusable framework for multi-crop, multi-trait digital phenotyping without task-specific model retraining.
ano.nymous@ccsd.cnrs.fr.invalid (Abdoul Djalil Ousseini Hamza) 05 Oct 2026
https://hal.inrae.fr/hal-05778771v1
-
[hal-05774005] Supervised NMF to unravel signature-class associations in microbiome composition: application to microbial signature of downy mildew epidemology in vineyards
<div><p>Microbial communities play a crucial role in plant protection, so that microbial signatures associated with high and low disease pressure could be detected in microbiome data and serve for epidemiosurveillance. In this paper, we build upon non negative matrix factorization (NMF) techniques, that have been previously used to decipher microbial signatures in the human gut, to analyze metabarcoding data of vineyard soils. These soils originate from plots characterized by high or low levels of incidence and severity of downy mildew (DM) symptoms, a major grapevine disease, according to a long-term epidemiological survey. In this dataset, the microbiome varies strongly among vineyard plots, as they are cultivated with different grapevine cultivars and management practices under different climatic conditions. Therefore, dedicated methodology must be deployed to discard this confounding geographical signal to target epidemiological signal with higher accuracy. We introduce two different NMF methods aiming to capture epidemiological signal: simultaneous and sequential multi-task supervised NMF. These methods are compared to classical NMF to predict epidemiological outcome from microbiological external data. Multi-task NMF is the most efficient method to predict epidemiological outcome and capture microbial signatures associated with low DM pressure, partially recovered by the other methods. These findings open new perspectives for microbiome-based epidemiosurveillance of crop health.</p></div>
ano.nymous@ccsd.cnrs.fr.invalid (Alioune Badara Diouf) 01 Oct 2026
https://hal.inrae.fr/hal-05774005v1
-
[hal-05757637] AI in Food Systems
Food is about people—flavour, nutrition, health, relationships, community, joy, and pleasure, all contributing to a flourishing life. The role of technology, in particular artificial intelligence (AI), is to improve each of these benefits and minimize contradictions and frictions between them as much as possible. Food is also about the environment, energy, economics, and governance. In fact, it is the latter, with support from AI, that are important in realizing the former. Alongside the need for a conducive natural environment and climate, food constitutes the very basis both of humanity’s existential survival and of our mosaic of national and local cuisines and cultures. Given AI’s transformative impact and its huge future potential, how can it enable much better functioning food systems, and what are the main challenges involved? This chapter directly addresses these issues. First, by outlining the status of the global food system providing the context for the different roles AI can play. Second, by examining how AI supports different parts of the food value chain. Third, by proposing the circular and local food system as an important generic food strategy enabled by AI. Finally, a conclusion reflects on the main issues and takes a forward look.
ano.nymous@ccsd.cnrs.fr.invalid (Jeremy Millard) 21 Sep 2026
https://hal.science/hal-05757637v1
-
[hal-05758414] Semantic mapping for safe navigation in agriculture
Context: To support agro‑ecological practices, agricultural robots must travel autonomously and safely from their base to the fields, while adapting their behavior to a changing environment. At present, autonomous navigation is limited to the interior of the fields assuming action on a single type of vegetation and relying on human supervision. Moreover, row detection, obstacle avoidance and traversability analysis tend to be treated as separate problems. Hence, the adaptation of treatment/behavior with respect to environmental conditions and properties or unexpected hazards is not addressed, although it is essential for real‑world deployment. Proposed approach: In this paper, an integrated approach is proposed to construct a semantic map allowing a robot to self reconfigure its behaviour, with respect to environment recognition capabilities. It is based on a perception strategy usable for many outdoor conditions (urban as well as off‑road) that mix 3D sensors with a regular camera. A 3D semantic map can then be generated relying on deep learning, thanks to an innovative approach, shown in Figure 1, decomposed as follows: 1. Robust 2D semantic segmentation based on RGB camera – a deep learning model is trained on several open‑source off‑road image datasets selected for diversity, providing a broad visual basis for agricultural contexts. 2. 3D occupancy mapping – lidar, IMU, GPS and odometry data are fused to generate an egocentric 3D occupancy map of the robot’s surroundings, using ICP (Iterative Closest Point) and dynamic object removal. 3. Semantic projection – the 2D semantic labels are projected onto the 3D map and temporally aggregated, yielding a clean, traversability‑aware 3D semantic map. A dedicated ontology is proposed to merge the various datasets which defines hierarchical traversability levels (non‑traversable obstacles, lightly touchable vegetation, fully traversable terrain) that can be exploited to select the appropriate behaviour to be adopted as well as the robot configuration. Results: The efficiency of the proposed approach is highlighted through actual experiments obtained in a mixed experimental area (vegetation/urban). We demonstrated that combining multiple datasets during training improves the model’s ability to generalize across diverse environments while preserving predictive performance on each individual dataset. The high‑quality semantic predictions enable the construction of accurate 3D semantic maps even in previously unseen domains, without requiring complex spatio‑temporal aggregation. Conclusion: The proposed pipeline demonstrates that semantic mapping can provide a reliable safety layer for agricultural robots, facilitating fully autonomous navigation from the base to the fields. It opens further the way to the necessary adaptation of a robot’s behavior in different types of area, allowing a discriminate treatment of the environment and fields pending on their recognized properties/configurations.
ano.nymous@ccsd.cnrs.fr.invalid (Alexis Frizot) 22 Sep 2026
https://hal.science/hal-05758414v1
-
[hal-05747907] Optimal Transport Prototypes and Symbolic Logic Combination for Self-explainable Graph Neural Networks
Graph Neural Networks achieve strong performance across diverse graph-based tasks, but their black-box nature remains a major obstacle in domains where interpretability is crucial. Prototype-based self-explainable GNNs mitigate this issue by grounding predictions in comparison with learned patterns. However, existing methods either restrict prototypes to training subgraphs, thereby limiting expressiveness, or learn latent prototypes that lack interpretable counterparts in graph space. We introduce a hybrid self-explainable graph classifier that integrates optimal transport, prototype grounding, and symbolic logic to produce faithful and human-readable explanations. We learn prototypes as point clouds in the embedding space and compare them to input graphs using partial optimal transport. To recover semantic meaning in graph space, each prototype point is anchored to a representative ego-network extracted from the training graphs. These similarity scores are then processed by a transparent logic layer that distills the model's decisions into human-interpretable logical rules. Experiments on molecular and synthetic benchmarks show that our method achieves competitive predictive performance while producing interpretable, example-based explanations grounded in real graph structures.
ano.nymous@ccsd.cnrs.fr.invalid (Elouan Vincent) 12 Sep 2026
https://hal.science/hal-05747907v1
-
[hal-05357866] Adopting graph-based pangenomics for agronomy and biodiversity studies: current resources and challenges
Pangenomics is transforming the way genomics is done. By using assemblies of multiple complete genomes for the same species, it is possible to depart from a biased, reference-centric, point of view. Represented as graphs, these pangenomes open new paths for genome understanding and improvement of species of agricultural interest. They have, for instance, been used to discover new structural variants linked with traits of interest, and increased the heritability prediction. However, transitioning to a graph-based approach is not straightforward, and many tools are still needed for this shift. While numerous papers present both methodological and applied approaches on pangenomics, and many reviews present advantages of these methods, very few resources outline what remains to be done. In this paper, we would like to list methods that are still required to fully exploit pangenomes, and favor the widespread adoption of this approach.
ano.nymous@ccsd.cnrs.fr.invalid (Stéphanie Bocs) 23 Sep 2026
https://hal.inrae.fr/hal-05357866v7
-
[hal-05743930] Generation of metabolomic-informed models of metabolism in microbial communities
Presentation for the ISGSB 2026: The generation of genome-wide metabolic networks has become a routine analysis for individual organisms or communities communities. However, these automatically generated metabolic networks are incomplete because they are constructed by based on the combination of gene annotation and reactions available in generic available in generic databases (Metacyc, BIGG, ModelSEED...). These are oriented towards well-known organisms or organisms or model organisms and miss out on important functions secondary metabolism. We propose to combine metabolomic data analysis, metabolic modelling and annotation metabolic modelling and annotation mining to build high-quality models of high quality models of microbial metabolism with the long-term aim of better understanding of microbial communities. In terms of application of the methods to plant microbial communities, we hope that the plant microbial communities, we hope that the newly developed models will provide a better understanding of the process of microbial recruitment by the plant: metabolic functions involved, micro-organisms associated with these functions.
ano.nymous@ccsd.cnrs.fr.invalid (Coralie Muller) 09 Sep 2026
https://inria.hal.science/hal-05743930v1
-
[hal-05761385] Pareto Suboptimal Resource Allocation and Microbial Growth
Abstract Microbial growth has often been analyzed under the assumption that microorganisms have evolved to optimize phenotypic characteristics of interest, such as growth rate and growth yield. This assumption has been useful, for example, for the genome-scale modeling of metabolism and the study of the allocation of cellular resources to physiological processes. In many experimental situations of interest, however, microorganisms are found to be suboptimal with respect to phenotypic characteristics that are thought to be favorable in that situation. Whereas a strong theoretical framework exists to mathematically relate cellular resource allocation strategies to Pareto optimality of microbial growth and other biological processes, much less is known about the consequences of Pareto suboptimality. We extend the framework to the latter case and show that a given Pareto suboptimal phenotype can be explained by a range of underlying resource allocation strategies, each corresponding to a different growth physiology and biomass composition. We test the predictions with the help of a coarse-grained model of microbial growth and published experimental data, which relate Pareto suboptimal rate-yield phenotypes of Escherichia coli and the microalga Tisochrysis lutea to the macromolecular composition of the cells (storage metabolite and total protein contents). Changing the focus from Pareto optimality to suboptimality provides interesting leads to exploring the diversity of growth strategies that can support a given phenotype. This change of perspective is of practical interest, because some of the growth physiologies within this range may be important for biotechnological applications. Author summary Many methods for the analysis of metabolism and growth are based on the assumption that, under the pressure of natural selection, microorganisms have evolved towards phenotypes that are optimal with respect to ecologically important characteristics. In mathematical terms, this assumption can be formulated as Pareto optimality of microbial growth and related to underlying resource allocation strategies that shape cellular physiology. Based on the observation that this assumption is often not satisfied in practice, we investigate the consequences of adopting the opposite assumption, namely that microbial growth and cellular resource allocation are suboptimal. We extend an existing mathematical framework to the case of Pareto suboptimality, which allows us to make predictions on the diversity of the macromolecular composition of microbial cells. These predictions are shown to correspond with experimental data on growth rate, growth yield, and macromolecular composition of the model bacterium Escherichia coli and the microalga Tisochrysis lutea . Our results thus associate suboptimality with observed metabolic diversity in microorganisms.
ano.nymous@ccsd.cnrs.fr.invalid (Valentina Baldazzi) 23 Sep 2026
https://hal.inrae.fr/hal-05761385v1
-
[hal-05726408] PLEX: Edge-Aware Logic-Based Self-Explainable Graph Neural Networks with Polyadic Predicates
Deploying Graph Neural Networks in safety-critical domains is hindered by their lack of interpretability, despite their state-of-theart performance across numerous domains. Self-explainable architectures grounded in formal logic have emerged as a promising path toward transparency, but existing approaches are confined to unary predicates, making them blind to relational properties between nodes and the information carried by edges. We introduce PLEX (Polyadic Logic EXplainable GNN), a self-explainable GNN built upon Polyadic Graded Modal Logic, designed to overcome these limitations. PLEX adopts a hierarchical duallogic architecture that learns ternary predicates, enabling explicit modeling of edge attributes and pairwise node interactions, capabilities entirely absent in monadic predecessors. We evaluate PLEX across graph classification, node classification, and link prediction benchmarks, where it matches the predictive accuracy of black-box models and consistently surpasses monadic logic-based models. Moreover, the richer logical language does not come at the cost of interpretability: we show that greater expressivity can actually reduce the number of logical rules needed, yielding models that are simultaneously more accurate and more transparent.
ano.nymous@ccsd.cnrs.fr.invalid (Simone Palumbo) 25 Aug 2026
https://hal.science/hal-05726408v1
-
[hal-05682039] Reconstruction of ancestral plant genomes for inter-crop translational research
We present an initial exploratory framework, Ancestral Genome Reconstruction (AGR), to automatically infer ‘paleogenomes’ from large comparative datasets. By analyzing 84 extant angiosperm species, we reconstructed 10 key ancestral plant genomes of millions of years old. These reconstructed ancestors were instrumental in (i) estimating when emerged the angiosperms as well as major botanical families as well as ancestral shared whole genome duplication events, (ii) tracing the evolutionary trajectories of ancestral chromosomes as well as genes, especially those that may have driven the emergence of key life history traits (exemplifies with woody vs. herbaceous, aquatic vs. terrestrial, C3 vs. C4 and symbiotic root nodulators vs. non-nodulators species). We demonstrate that these paleogenomes serve as tractable backbones for inter-crop translational research, in delivering though an open access web tool (OrthoViewer) genes that have conserved the same ancestral genomic context favoring the identification of ‘phenologs’ -genes underlying similar phenotypes, traits, or processes across species- exemplified with FUWA for yield components, FLC for flowering time, and DDM1 for DNA methylation. Taken together, this study provides a testable paleogenomics workflow, opening novel avenues to integrate evolutionary genomics data into modern climate-smart breeding and support the agroecological transition.
ano.nymous@ccsd.cnrs.fr.invalid (Cléa Siguret) 18 Sep 2026
https://hal.science/hal-05682039v2
-
[hal-05743184] Contributions of physiological sensors and AI-assisted videos to the study of physiological and behavioural adaptative mechanisms in growing pigs facing a heat wave
Heat stress (HS) is a growing concern due to global warming, especially for livestock species with limited thermoregulatory capacities such as pigs. Several physiological and behavioural mechanisms are triggered to facilitate pig’s adaptation to elevated temperature, which is however generally associated to reduced growth performance. Moreover, there is a wide individual variation in how quickly and effectively pigs may adapt to heat. The aim of this study was to describe the real-time diurnal and nocturnal variations in some physiological and behavioural indicators in growing pigs facing a heatwave. Twelve crossbred (Large White × Landrace) × Piétrain) male pigs were reared in individual indoor pens, with an ambient temperature set to 22°C for 18 days (thermoneutral conditions: [TN]), and then at 32°C for the subsequent 6 days (heat stress challenge [HS]). All pigs were fitted with an indwelled venous catheter, and time-series blood samples were collected every hour at one day during TN and at one day during HS, in order to study the circadian variations of plasma glucose levels. Salivary cortisol levels were measured at two days during TN and each day of HS. Continuous monitoring of core temperature (Tcore) were obtained from sensors (Anipill, France) implanted intramuscularly. A continuous glucose monitoring (CGM) system (Dexcom G6, Dexcom Clarity, US) was inserted subcutaneously the day before HS, and rest during the following 6 days. Measurements were wirelessly transmitted to dedicated recorders. Pigs were fed ad libitum thanks to automated electronic feeders. Pens were video recorded by infrared cameras. We developed AI-based assisted classification of pig behaviours using a convolutional neural network (YOLOv11) and a minute frequency sampling rate, to extract the following behavioural traits (frequency of occurrence): visit to the drinker, head in the feeder, lying down, standing up, sitting, moving, sniffing or nosing the littermate in the neighbouring pen. Data were analysed byANOVA to assess the effects of the period (TN, HS) and circadian rhythm (nocturnal, diurnal). This reveals that salivary cortisol levels were superimposed on the circadian rhythm and, on average, were greater in HS than in TN (0.97 + 0.80 vs. 0.68 + 0.63 ng/mL, P &lt; 0.001). On average, Tcore was higher under HS than under TN (39.6 + 0.8 vs. 39.1 + 0.4°C, P&lt; 0.001), which indicates perturbations of thermoregulatory mechanisms. There was no difference in mean Tcore between diurnal and nocturnal periods. Daily mean plasma glucose concentrations were similar between TN and HS. The CGM data reveal a large inter-individual variability in metric for peaks and nadirs of glucose levels, in nocturnal as well as in diurnal times. Glycemic excursions and Tcore dynamics were generally synchronized in the HS period. The AI-assisted video classifications revealed behavioural shifts between HS and TN periods, with more frequent visits to drinkers, less frequent visits to feeders, more sitting and reduced social interactions (P &lt; 0.001). Data from the automatic feeder confirmed that pigs under HS had a lower feed intake (P &lt; 0.02) and shorter duration of meal ingestion (P = 0.01) than pigs under TN. In conclusion, sensor-based monitoring systems are helpful to detect (a) synchronic deviations of physiology under heat stress in growing pigs, whereas AI-assisted videos may be reliably used to detect warning signals of health and welfare deviations. This work was funded by the French National Research Agency under the project Wait4 (ANR-22-PEAE-0008).
ano.nymous@ccsd.cnrs.fr.invalid (Caroline Xavier) 10 Sep 2026
https://hal.inrae.fr/hal-05743184v1
-
[hal-05744249] Leveraging human-trained neural networks for cross-species chromatin regulation annotations
Analogous to the Encyclopedia of DNA Elements (ENCODE) project, the Functional Annotation of ANimal Genomes (FAANG) consortium has produced chromatin annotations for domesticated animals, albeit in smaller amounts. Although acquiring experimental data is more accessible and affordable for many species, human and mouse organisms will remain the reference. Classical methods based on sequence conservation can be used to infer missing annotations, but are inappropriate for non-conserved sequences. While regulatory sequences share low to moderate conservation, they have retained their regulatory function during the evolution process. Here, we take advantage of three neural networks (DeepBind, DeepSEA, and Enformer) trained with human and mouse ENCODE data to infer chromatin annotations (transcription factors binding, chromatin accessibility, and histone marks) in cattle, pig, chicken, and European seabass. For this purpose, we comprehensively assessed the quality of predictions using experimental data from FAANG, through AUC-ROC and AUC-PR metrics. Our results showed low variability between tissues and similar performances for various annotations in mammals and chicken, with AUC-PR ranging from 0.663 ± 0.010 (chicken) to 0.765 ± 0.009 (pig) for the best predicted experiment, H3K4me3 (all tissues grouped), but lower (0.238 ± 0.005) in fish. Further analyses focused on pigs highlighted (i) accurate predictions even for non-conserved sequences, and (ii) variable predictions depending on genomic feature annotations. Our results advocate the widespread use of human-trained neural networks as a first step in cross-species genome annotation before training species-specific models.
ano.nymous@ccsd.cnrs.fr.invalid (Noémien Maillard) 25 Sep 2026
https://hal.inrae.fr/hal-05744249v2
-
[hal-05760239] A loss-of-function allele fixed during domestication reshapes nectar chemistry, microbial diversity, and pollinator visits
Nectar is a hub for plant-pollinator interactions, yet gene-level causal links between plant genetic variation, pollinator foraging, and nectar microbial assembly remain poorly resolved. Using near-isogenic lines, innovative field time-lapse monitoring of pollinator visits, and long-read amplicon sequencing of nectar microbiota, we show that a natural single-nucleotide variant at a cell-wall invertase gene (HaCWINV2) controls sunflower nectar chemistry and influences both pollinators and microbes. Plants homozygous for a loss-of-function HaCWINV2 allele produce sucrose-rich nectar, resulting in fewer bee visits under field conditions. In pollinator-excluded flowers, invertase-deficient plants harbored greater fungal diversity and compositionally distinct communities, indicating that nectar sugar profiles act as ecological filters shaping the nectar microbiome. This loss-of-function allele is rare in wild sunflowers, but fixed in 35% of cultivated lines, indicating positive selection during domestication. Our findings establish a causal link between a single gene and nectar chemistry, with cascading ecological effects in a plant-pollinator system, thus illustrating how subtle genetic changes scale up to alter nectar traits, microbial assembly, and pollinator foraging behavior.
ano.nymous@ccsd.cnrs.fr.invalid (Guillaume Tueux) 23 Sep 2026
https://hal.science/hal-05760239v1
-
[hal-05709904] Leveraging hologenomic data for phenotypic prediction: potential and pitfalls
The microbiota is increasingly recognized as an active component of host biology, influencing various host phenotypes. Advances in high-throughput sequencing and the emergence of the holobiont perspective have raised expectations regarding hologenomic-informed prediction. Yet, whether and under which conditions integrating microbiota and genomic data meaningfully improves phenotypic prediction remains unclear. The biological characteristics of the microbiota, including but not limited to transmission mechanisms, environmental effects and interactions with host genetics, complicate their integration into classical evaluation frameworks. In addition, microbiota datasets are high-dimensional, highly dispersed, sparse and compositional. Finally, analytical choices such as the taxonomic granularity considered for aggregation or the similarity matrix used in prediction models may impact downstream inference and prediction accuracy. Here we explore these challenges using a comprehensive set of transgenerational hologenomic simulations. By generating controlled and contrasted biological scenarios across a broad parameter space, we examine how microbiota granularity, variance structure and host modulation influence (i) the estimation of variance components and (ii) the accuracy of phenotypic prediction. We show that the added value of hologenomic, compared to genomic prediction, is highly context dependent. Our results provide a structured framework to interrogate when and how integrating microbiota may enhance phenotypic prediction in breeding applications.
ano.nymous@ccsd.cnrs.fr.invalid (Solène Pety) 25 Aug 2026
https://hal.science/hal-05709904v2
-
[hal-05230510] Seed Inference in Interacting Microbial Communities Using Combinatorial Optimization
The behaviour of microorganisms and microbial communities can be abstracted by models combining a description of their metabolic capabilities as metabolic networks, and suitable computational or mathematical paradigms that further integrate simulation conditions. A major component of the latter is the composition of the environment or growth medium that can be referred to as seeds. Predicting the seeds from the metabolic network and an expected behaviour is an inverse problem that can be addressed with linear programming or logic paradigms such as Answer Set Programming (ASP). Here, we formalise seed prediction for microbial communities, taking into account that their members may interact positively through metabolite transfers, which may reduce the need for external seed metabolites. We address the problem with ASP and add a hybrid component ensuring the satisfiability of linear constraints. We explore the subset-minimality solving heuristic of the Clingo solver and develop two heuristics supporting priority of seeds over transfers. We present a proof of concept of seed inference in small-scale communities, and assess the scalability of the three heuristics at genome-scale. Overall, our work introduces a hybrid logic-linear model for seed inference in interacting microbial communities, and new heuristics for the exploration of the solution space with subset minimality optimisations.
ano.nymous@ccsd.cnrs.fr.invalid (Chabname Ghassemi Nedjad) 29 Aug 2025
https://inria.hal.science/hal-05230510v1
-
[hal-05708804] comFI: Inferring metabolic fluxes in microbial communities from time-resolved multi-omics data
Microbial communities are complex ecosystems that are actively studied for the multiple beneficial services they provide, spanning from food fermentation and wastewater treatment, to biotechnology, medicine, crop protection and animal health. Understanding the functioning of microbial communities remains challenging due to the complexity of metabolic interactions and community assembly mechanisms. Recent advances in high-throughput sequencing and analytical technologies enable the collection of longitudinal multi-omics data at the community scale, including population counts, metabolomics, and metatranscriptomics. However, integrative modeling frameworks capable of exploiting such heterogeneous data to infer community-level metabolic activity remain limited. Here, we introduce the community Metabolic Flux Inference (comFI) method, a biology-informed inference framework that enables the integration of multi-omics time series to infer metabolic fluxes at both the community and species levels. The method aims to (i) quantify the contribution of each community member to targeted extracellular metabolite consumption and production, and (ii) infer intracellular flux distributions consistent with transcriptomics and observed community-scale metabolite dynamics. We extensively benchmarked comFI robustness on synthetic datasets for varying ecological interactions, community richnesses, observation noises or transcriptomic depths. The method provides accurate reconstructions of metabolic fluxes across benchmarks, with an average goodness-of-fit R2 of 94.16% for observed metabolic exchanges and 94,83 % for unobserved intracellular metabolic flux reconstruction, achieving very low false positive and negative rates of predicted activated reactions compared to metatranscriptomics (2 % and 5 % respectively). We then applied comFI on two experimental datasets illustrative of different ecosystems and applications: (i) a denitrifying synthetic community, representative of natural soil microbiota and of direct relevance for global N- cycle ; and (ii) a controlled cheese fermentation community involving three species engaged in cooperative and competitive metabolic interactions during lactose-to-lactate conversion and subsequent ripening stages. The comFI method recovered known interactions, while providing quantitative insights on metabolic exchanges and microbial physiology, consistent with metabolomics and metatranscriptomics data.
ano.nymous@ccsd.cnrs.fr.invalid (Sthyve Junior Tatho Djeanou) 07 Aug 2026
https://hal.science/hal-05708804v2
-
[hal-05760845] Integrating metagenome-scale metabolic models and metabolomics to explore candidate biochemical interactions in cultivated <i>Microcystis</i> phycospheres
Favored by global changes, freshwater cyanobacterial harmful blooms generate major ecological, economic, and public health challenges. Microcystis, one of the most widespread cyanobacterial genera, grows within a phycosphere where specialized interactions with its microbiome occur, that are suspected to influence bloom appearance and its potential toxicity. Using a combination of metagenomics, metabolomics, and metabolic modeling, we characterized the culture-associated phycospheres of 12 Microcystis strains isolated from a French pond. The distribution of metabolic reactions within Microcystis was consistent with their genospecies, whereas the metabolic landscape at the community level diverged from cyanobacterial phylogeny, indicating partial functional decoupling between cyanobacteria and their associated microbiomes. Bacteria associated with the simplified phycospheres substantially expanded the metabolic repertoire of the system, while maintaining functional redundancy within and across communities. On the other hand, endometabolomic profiles were largely driven by cyanobacterial metabolic outputs, whereas exometabolomic analysis did not reveal metabolites involved in exchange processes. Metabolic modeling, together with the identification of toxic specialized metabolites produced by specific biosynthetic gene clusters, further highlighted differences in metabolic potential among phycospheres. Together, these findings deepen the understanding of Microcystis' phycosphere functioning and demonstrate the value of multi-omics systems biology approaches, while suggesting that metabolic complementarity between species and across phycospheres could play a role in bloom-associated microbiome structure.
ano.nymous@ccsd.cnrs.fr.invalid (Juliette Audemard) 23 Sep 2026
https://inria.hal.science/hal-05760845v1
-
[hal-05656546] A soil-type-specific stratified approach to bare soil mosaicking and SOC prediction from Sentinel-2 time series
Accurate, high-resolution mapping of soil organic carbon (SOC) is essential for environmental modelling and sustainable land management, yet its prediction based on satellite imagery is often affected by vegetation and moisture, possibly causing generalised models to fail in landscapes with heterogeneous soils. To address this, we developed a stratified framework that tailors bare soil thresholding and SOC modelling to specific soil types. Using a 7-year Sentinel-2 time series and 414 soil samples over the Centre-Val de Loire region in France, we first identified the Visible and Shortwave Infrared Drought Index (VSDI) as an effective moisture proxy, avoiding the need for availability-limited external moisture data. We then optimised NDVI, NBR2, and VSDI thresholds individually for each major soil type to filter bare soil observations. Finally, soil-type-specific partial least squares regression (PLSR) models were built and compared against a single generalised model.<p>Our results showed that our soil-type-specific strategy substantially outperformed the generalised model (e.g., RPIQ increased from 0.68 to 2.38 for Brunisols eutriques). The primary spectral predictors for SOC were highly variable, varying from visible to SWIR bands according to the soil type. The optimised VSDI filter was also critical for this improvement, reducing prediction RMSE by nearly 50% for loamy-texture soils like Luvisols. This study demonstrates that, for accurate SOC prediction at a very large regional scale, context (soil-type)-specific stratification of bare soil thresholding and SOC modelling is critical, serving as a framework for integrating pedological knowledge into SOC prediction and subsequent digital soil mapping workflow.
ano.nymous@ccsd.cnrs.fr.invalid (Qianqian Chen) 14 Jun 2026
https://hal.inrae.fr/hal-05656546v1
-
[hal-05657073] A free time machine? Milking robots and transformations in the temporal regime of dairy farmers in France
This article examines the effects of the large-scale diffusion of milking robots on dairy farmers’ lifestyle, with a special attention to the temporal patterns of their daily life. The study is based on quantitative data from a questionnaire completed by 831 respondents and qualitative data from 43 interviews with dairy farmers. It critically questions prevailing narratives depicting milking robots as technologies that liberate farmers' time and modernize their lifestyles. Our findings show instead that the robot reconfigures labour along principles of fluidity, and further blurs the already fuzzy boundaries between work and the domestic sphere. Farmers' daily lives remain tied to a largely unchanged temporal regime defined by long working hours. The reduction of total work hour is on average of 6% and mostly concerns evening hours. The morphology of work time, as well as the patterns of articulation of social times within farming households, are only marginally altered by the adoption of milking robots.
ano.nymous@ccsd.cnrs.fr.invalid (Nicolas Deffontaines) 15 Jun 2026
https://u-picardie.hal.science/hal-05657073v1
-
[hal-05719117] A synopsis of neotropical parasitoid wasps of Diaphania hyalinata (Lepidoptera, Crambidae), with new records from Guadeloupe
Diaphania hyalinata Linnaeus, 1767 (Lepidoptera, Crambidae), commonly known as the melonworm, is a major pest of cucurbits in the Neotropics. Yet, a comprehensive review of the parasitoid wasps that may act as biocontrol agents of this pest is still lacking. An extensive literature review was conducted to compile the current knowledge on the identity, distribution, and biology of parasitoid wasps associated with D. hyalinata across its range. Eggs, larvae, and pupae of D. hyalinata were collected in Guadeloupe and reared to document the local parasitoid diversity. Among the 37 parasitoids identified in our review, eight were recovered during our field survey, all of which are new records for the island. These parasitoids were photographed and DNA-barcoded to facilitate their potential use in future biological control programs targeting D. hyalinata. The species Trichogramma pretiosa Riley, 1879 (Trichogrammatidae; egg parasitoid) and Schoenlandella montserratensis Kang, 2021 (Braconidae; larval parasitoid) appear to be promising candidates for the biological control of D. hyalinata in Guadeloupe. It may be also worth evaluating the effectiveness of two additional larval parasitoids, Eiphosoma dentator Fabricius, 1804 (Ichneumonidae) and Apanteles impiger Muesebeck, 1958 (Braconidae), while assessing possible non-target effects on local fauna. However, it is important to consider that the presence of the hyperparasitoid Aphanogmus fijiensis Ferrière, 1933 (Ceraphronidae) could undermine biocontrol efforts. Overall, it would be valuable to gain a better understanding of the biological and environmental factors (such as agricultural practices and landscape features) that enhance the natural regulatory activities of this parasitoid community.
ano.nymous@ccsd.cnrs.fr.invalid (Margot Gumbau) 17 Aug 2026
https://hal.inrae.fr/hal-05719117v1
-
[hal-05704872] SelNeTime
[...]
ano.nymous@ccsd.cnrs.fr.invalid (Mathieu Uhl) 27 Jul 2026
https://hal.inrae.fr/hal-05704872v1
-
[hal-05701255] Information criteria exploiting latent structure for model selection in Structural Equation Models
Structural equation models (SEM) are widely used to describe dependency structures between latent variables, making model selection a key issue in many applications. Existing information criteria are generally based on the integrated observed-data likelihood and therefore do not explicitly account for the latent structure of the model. In this paper, we propose two new information criteria derived from the integrated complete-data likelihood. The first adapts the Integrated Completed Likelihood criterion to Gaussian SEM, while the second proposes an alternative approach to approximating the integrated observed-data log-likelihood by incorporating latent structural information and using an importance sampling strategy. Their performance is assessed through an extensive simulation study covering null, direct, indirect and complete latent structures under different sample sizes and signal strengths. The results show that the proposed importance sampling strategy provides robust and competitive model selection across a wide range of scenarios, whereas the proposed ICL criterion is particularly effective for recovering latent dependency structures when the latent variables are accurately estimated. These findings demonstrate the potential benefits of explicitly exploiting the latent structure when developing information criteria for structural equation models.
ano.nymous@ccsd.cnrs.fr.invalid (Marion Naveau) 22 Jul 2026
https://hal.science/hal-05701255v1
-
[hal-05697386] Modeling breeding programs considering social behavior in large groups of farmed fish
Breeding programs are essential in aquaculture, improving economically and environmentally important traits. In aquaculture systems, animals are raised in large groups, where social interactions are frequent and can influence individual performance. In these circumstances, indirect genetic effects can play an important role in the response to selection, and consequently, their effects on selection outcomes must be analyzed. This study aimed to evaluate the implications of heterogeneous social interaction effects on fish breeding programs using stochastic simulations. We simulated a fish breeding program with 2000 selection candidates from 1000 families formed by a partial mating design of 100 males and 100 females. Social interactions were simulated, affected by the target phenotype and two latentpersonality traits. We investigated how genetic gains and phenotypic variances are affected by the magnitude and direction of social interaction effects on the target phenotype, different selection strategies, and the genetic correlations between the target phenotype and personality traits. Our results showed that increased social interaction effects lead to greater phenotypic variability in the target trait. Under mass selection, the genetic means of personality traits change, and these changes depend on the strength and direction of genetic correlations between the focal and personality traits. To achieve similar genetic gain under group selection, many groups with few families are needed. Phenotypic variability of the target trait and genetic means of the personality traits remain constant under group selection. However, this strategy increases the rate of inbreeding per generation, but in a similar amount across the different group designs.
ano.nymous@ccsd.cnrs.fr.invalid (Gabriel Rovere) 19 Jul 2026
https://hal.science/hal-05697386v1
-
[hal-05692432] Text guidance is powerful but prompt-sensitive for weakly-supervised leaf symptom segmentation
Accurate segmentation of plant disease symptoms is essential for crop monitoring and phenotyping, yet it typically requires costly pixel-level annotations. Weakly supervised semantic segmentation (WSSS) alleviates this burden by relying on cheaper forms of annotation, such as image-level labels, bounding boxes, or point annotations, to produce supervision (pseudo-mask) for training a segmentation model. We investigate whether text-guided segmentation with the Segment Anything Model 3 (SAM3) can serve as an alternative as it directly segments regions matching a text prompt description of the disease symptoms.Three pseudo-mask generation strategies are compared: (i) class activation maps (CAMs) refined with SAM or SAM3, (ii) zero-shot text-guided SAM3, and (iii) a hybrid approach combining weak spatial cues with text prompts. The resulting pseudo-masks are used to train a fully supervised model (DeepLabV3+). Text guidance alone matches or outperforms conventional WSSS, achieving up to 0.46 IoU without spatial supervision and 0.61 IoU on a public dataset, although performance is sensitive to text prompt formulation. The hybrid strategy improves robustness, reaching 0.50 IoU on the primary dataset and 0.58 IoU on the additional dataset while reducing prompt sensitivity. Overall, text guidance is a promising alternative to conventional weak supervision, while hybrid approaches provide a more robust solution for plant disease segmentation.
ano.nymous@ccsd.cnrs.fr.invalid (Romane Dubois) 11 Sep 2026
https://hal.science/hal-05692432v1
-
[hal-05688477] Ancient wheat DNA: How and what for?
In complement to classical approaches currently used in agricultural science in characterizing and exploiting in breeding traits- or phenotype-driving genes and genetic variants from modern genetic diversity, knowledge of past evolution and adaptation processes from ancient plant DNA can be mobilized to develop varieties adapted to current environmental challenges. The article reviews the current scientific and technical insights from the investigation of ancient DNA from plant remains with a focus in wheat.
ano.nymous@ccsd.cnrs.fr.invalid (Caroline Pont) 10 Jul 2026
https://hal.science/hal-05688477v1
-
[hal-05710282] Using Energy Flow Analyses (EFA) in a local participatory process for energy policy-making: lessons learned
Over the past few decades, conflicts have emerged at the local scale over the prioritisation of sustainability issues. In the Grand Briançonnais, a region of the French Alps with approximately 34,500 inhabitants across 36 municipalities, these conflicts focus on the construction of new micro-hydroelectric power plants. Although these facilities produce low-carbon energy, they also affect ecosystems and water sports through the artificialisation of river sections. These conflicts have led various local stakeholders to take legal action against such projects. The strong opposition between proponents and opponents of hydropower is partly rooted in significant differences in values, perceptions and understandings of sustainability in the region. In response, the local agency coordinating several inter-municipal authorities—the PETR du Grand Briançonnais —announced in 2022 its intention to develop a coherent local energy policy co-produced with citizens and local energy stakeholders. The objective was to reopen dialogue by debating new micro-hydroelectric projects while broadening the discussion to the entire regional energy policy and opening participation to all interested inhabitants. To achieve this objective, the president of the PETR contacted scientists who decided to help to create the required participatory process and to launch an action research project aimed at evaluating the use of energy flow analyses as a basis for systemic representations of energy issues to support local participatory processes. This research objective has been further detailed into three research intentions : (1) to test the feasibility of local-scale energy flow modelling, detailed enough to be relevant to stakeholders; (2) to assess citizens’ understanding of energy flow analyses and the learning outcomes associated with different formats, namely facilitated workshops and a serious game; and (3) to assess the perceived relevance of energy flow analyses as scientific support for the different phases of the participatory process. The evaluation followed a twofold approach. First, the impact of the use of energy flow analyses was assessed through observations and participant questionnaires. Second, a global evaluation of their mobilisation throughout the participatory process was conducted using participant questionnaires and an ex-post self-reflexive approach. Depending on the criteria considered, between approximately 20 and 50 participants were engaged in the evaluation, within a process involving 330 inhabitants. The results show that local-scale modelling of energy flows is feasible, although it presents several challenges. Data availability, dispersion and relevance, as well as the articulation between scientific modelling and territory-specific perceptions, are key factors in producing meaningful support for local policy-making. Energy flow analyses were generally understood by participants not necessarily familiar with energy issues, provided that appropriate support was offered. Learning outcomes differed depending on the format. Facilitated workshops mainly improved cognitive knowledge of the region’s physical dynamics and helped prioritise actions for energy policy formulation. In contrast, the serious game fostered a deeper understanding of socio-ecological complexity and highlighted the role of governance and relational capacities in achieving sustainability. Overall, energy flow analyses were considered relevant and useful in the consultation process, particularly during the diagnostic phase and issue prioritisation. However, identifying and operationalising effective levers for action remained a major challenge.
ano.nymous@ccsd.cnrs.fr.invalid (Emmanuel Krieger) 03 Aug 2026
https://hal.science/hal-05710282v1
-
[hal-05710002] On Computing and Exploring Multiple Consistent Metabolisms in Material Flow Analysis
This communication focuses on how to compute and explore bio-physically consistent metabolisms for foresight scenarios using Material Flow Analysis (MFA). MFA tools usually output a single solution or a posterior distribution. We consider another option, particularly for the iterative construction of foresight scenarios, by proposing to compute and then explore several close-to-optimal but widely different solutions. This work is embedded in the development of the STAX software.
ano.nymous@ccsd.cnrs.fr.invalid (Thibaut Coudroy) 03 Aug 2026
https://inria.hal.science/hal-05710002v1
-
[hal-05712379] Simulating population pangenomes under coalescent demographic models with MSpangenome
Abstract Motivation Pangenome variation graphs (PVGs) are increasingly used to represent genomic diversity, yet there is currently no general framework for generating population pangenomes directly from explicit evolutionary histories. Existing simulators typically focus on individual classes of variation and do not integrate these variations within a genealogy-aware framework driven by explicit demographic histories. As a result, evaluating pangenome methods in realistic population-genetic settings remains challenging, and benchmark datasets with known evolutionary ground truth are scarce. Results We present MSpangenome , a genealogy-aware framework that bridges coalescent population genetic simulations and pangenome graph analyses. The pipeline combines ancestry simulation with msprime and a de novo graph construction algorithm to generate PVGs directly from simulated genealogies. By explicitly modeling recombination, demographic history and incomplete lineage sorting, MSpangenome produces structurally complex pangenomes in which nested and overlapping structural variants emerge naturally from the underlying genealogies, while their evolutionary history and graph topology remain known by construction. This provides a general framework for generating realistic population pangenomes and establishing ground-truth datasets for methodological evaluation. We demonstrate its utility by generating population-scale pangenomes and using them as controlled references to benchmark the widely used graph construction tools, PGGB and Minigraph-Cactus . Our analyses reveal contrasting performance regimes across levels of sequence diversity, sample sizes and classes of structural variation, highlighting the value of simulation-based benchmarking for identifying reconstruction errors that are hard to detect using empirical datasets alone. Availability and implementation MSpangenome is implemented in Python, fully containerized, freely available at https://forge.inrae.fr/pangepop/MSpangepop and mirrored at https://github.com/inrae/MSpangepop .
ano.nymous@ccsd.cnrs.fr.invalid (Lucien Piat) 06 Aug 2026
https://hal.inrae.fr/hal-05712379v1
-
[hal-05780972] Genomic forecasting for climate‐resilient fruit trees
Fruit trees – long‐lived perennial crops cultivated for their edible fruits or nuts and frequently propagated clonally – are increasingly exposed to climate extremes that threaten their productivity and survival. Yet their capacity to adapt to rapid environmental change remains poorly understood. We argue that fruit trees and their wild relatives are powerful but underused systems for advancing genomic forecasting in perennials, with a focus on genomic offset analyses. Genomic offset estimates the mismatch between current genomic variation and that predicted to be optimal under future climates, offering a promising framework to anticipate maladaptation and guide conservation, breeding, and management strategies. Although its application is expanding rapidly in annual crops and forest trees, its interpretation and predictive value remain actively debated and require stronger empirical validation. Fruit trees are particularly well suited to address these challenges as they combine distinctive biology – including long generation times, clonal propagation and intensive management practices – with expanding genomic resources and common garden networks. Using emblematic Mediterranean and temperate species, we outline a roadmap that combines genomic offset with common‐garden networks, high‐resolution climate data, and trait‐based fitness proxies. Together, these resources position fruit trees as an powerful model to evaluate and refine genomic forecasting into a practical tool for biodiversity‐informed breeding and conservation under global change, and better understand plant adaptation and maladaptation processes.
ano.nymous@ccsd.cnrs.fr.invalid (Maxime Criado) 06 Oct 2026
https://hal.inrae.fr/hal-05780972v1
-
[hal-05694907] How Metagenomic Analysis Strategy Shapes Functional Inference? Metabolic Landscapes from Le French Gut Cohort
Background The human gut microbiome plays a key role in host health by contributing to digestion, immune regulation, and metabolic processes. Alterations in microbial community composition have been linked to numerous diseases, motivating large-scale metagenomic studies to better characterise microbial diversity and function. Projects such as Le French Gut aim to describe gut microbiota structure at the population level (currently n = 10,000 individuals) using shotgun metagenomic sequencing. However, different analytical strategies may strongly influence genome reconstruction, genome-scale metabolic network (GSMN) inference, and downstream analyses. Results We analyzed 105 fecal metagenomes from the Le French Gut cohort using both a reference-based profiling strategy and a de novo assembly-based protocol to retrieve gene catalogues and metagenome-assembled genomes (MAGs). Each sample was characterised as a collection of reference genomes, de novo MAGs, and as a gene catalogue. We first assessed the impact of sequencing depth on GSMN reconstruction quality, then compared three reconstruction tools — CarveMe, gapseq and Pathway Tools — and finally evaluated different metagenomic units, including reference genomes, gene catalogues and (MAGs). Higher sequencing depth consistently improved GSMN completeness by increasing the number of reconstructed reactions and accessible metabolites. Although no reconstruction tool performs best in all contexts, gapseq provided the best balance between metabolic coverage and genetic support for the reconstructed GSMNs. Finally, the gene catalogue approach offered a good balance between functional completeness and biological realism, whereas compartmentalised approaches, particularly MAG-based reconstructions, remain essential to preserve species-level organization and investigate potential metabolic interactions within microbial communities. Conclusions Our results demonstrate that the choice of metagenomic analysis strategy significantly impacts functional inference. De novo and reference-based approaches provide complementary views of the gut microbiota. We further analysed reconstructed metabolic profiles in the context of the extensive health, lifestyle, and dietary metadata associated with the dataset. Overall, our findings highlight the importance of methodological choices in metagenomic studies and provide a framework for integrating microbial genomic data with clinical, dietary, and lifestyle information in large population cohorts.
ano.nymous@ccsd.cnrs.fr.invalid (Sarah Toubal) 16 Jul 2026
https://hal.science/hal-05694907v1
-
[hal-05697497] Assessing Dorado pseudourydilation RNA modification prediction on Arabidopsis thaliana ribosomal RNA
[...]
ano.nymous@ccsd.cnrs.fr.invalid (Emma Rodriguez) 19 Jul 2026
https://hal.science/hal-05697497v1
-
[hal-05686148] PanGen1C: a Snakemake Workflow for k-mers based genotyping of large cohorts on pangenome graph
Advances in long read sequencing have shown previously unstudied structural and sequence variation, often poorly represented by a single linear reference genome. Pangenome graphs address this limitation by integrating multiple haplotypes into a reference structure, more suited to represent the genomic diversity of a given species. Graph-based genotyping uses these references to improve read mapping and variant detection, particularly in highly variable regions, but are intensive both computationally and on storage. K-mer based approaches allow the use of short-read data sequencing data without graph alignment, enabling efficient and scalable detection of both small and structural variants in population-scale datasets. We present PanGen1C, a Pangenie-based workflow to facilitate genotyping large-scale short read sequencing data.
ano.nymous@ccsd.cnrs.fr.invalid (Martin Racoupeau) 08 Jul 2026
https://hal.science/hal-05686148v1
-
[hal-05683977] Assessing the structure of DNA representation spaces using graph-based comparisons
Many models have been proposed to create embeddings of DNA sequences. While these models are typically evaluated using downstream tasks such as species prediction, such evaluations offer limited insight into the intrinsic differences between their embedding spaces. To address this gap, we focus on direct comparison of the geometric and topological properties of embeddings generated by those models. We consider five models (TNF, dna2vec, DNABERT-S, DNABERT-2 and HyenaDNA) chosen to represent a broad range of embedding techniques: from simple k-mer counts to state-of-the-art transformer architectures. Our comparison centers on two key questions. First, how do these models organize sequences from the same species in their latent spaces ? Second, how do they differ in terms of local and global topological properties ? We first evaluated whether sequences from the same species cluster together in each model’s embedding space. PCA projections revealed that, while all models grouped sequences by species, the degree of separation varied. Strikingly, the order of species along the first principal component was consistent across models, as were the relative distances between clusters. This suggests that, despite wide architectural differences, the models capture a shared global biological signal in their embeddings. Three models (TNF, dna2vec, and HyenaDNA) produced tightly clustered species groups, yielding similar silhouette scores. In contrast, DNABERT-S and DNABERT-2 exhibited more dispersed clusters: DNABERT-S achieved a higher silhouette score due to better separation, while DNABERT-2’s lower score reflected greater intra-cluster variance. This divergence may stem from the transformer architectures’ ability to capture more nuanced sequence features, albeit at the cost of cluster compactness. To further probe the embedding spaces using quantitative metrics, we constructed k-nearest neighbors (K-NN) graphs and applied two complementary analyses: (i) Jaccard distance to quantify local neighborhood similarities and (ii) a permutation test based on Random Dot Product Graphs (RDPG) to compare global topological properties. These methods enabled us to assess both fine-grained and large-scale differences between embedding spaces. Our analysis revealed that both Jaccard and RDPG comparisons are sensitive to hyperparameter choices. For K-NN graphs, the distance metric and number of neighbors (k) significantly impacted results, as did the size and composition of the input dataset. The RDPG framework introduced an additional hyperparameter: the latent dimensionality (d) for spectral embedding. Surprisingly, it also resulted in strong asymmetry in model comparisons (A vs. B ≠ B vs. A). This asymmetry poses a challenge for aggregating results and drawing robust conclusions. In summary, our study establishes Jaccard distance and RDPG on k-NN graphs as a unified framework for comparing DNA sequence embeddings, offering insights into both local and global properties of latent spaces. While methodological challenges, in particular hyperparameter sensitivity and comparison asymmetry, remain, addressing them could provide a way toward more biologically interpretable embeddings and deepen our understanding of what genomic models actually learn.
ano.nymous@ccsd.cnrs.fr.invalid (Juliette Francis) 07 Jul 2026
https://hal.science/hal-05683977v1
-
[hal-05754376] AphiDetect: open-by-design equipment and computer vision to characterize plant-aphid interactions in controlled conditions
Aphids exemplify adaptive evolution in response to the strong anthropogenic pressures exerted by poorly diversified and chemically protected agroecosystems. On the one hand, more sustainable agriculture would benefit from integrated pest management practices, which require a deeper understanding of plant-aphid interactions. On the other, increased cultivation of legumes (Fabaceae) could help reduce greenhouse gas emissions and enhance biodiversity in agroecosystems. However, existing legume collections remain partially characterized, and traditional phenotyping methods are laborious, destructive, and prone to human error. To study the genetic and molecular determinants of plant-pest interactions and describe key traits such as plant growth and aphid fecundity, the AphiDetect system was developed. Initially tested on faba bean (Vicia faba) inoculated with the pea aphid (Acyrthosiphon pisum) under controlled conditions, AphiDetect uses an open-by-design, reproducible, and low-cost phenotyping booth to count and classify aphids by developmental stage. For each inoculated plant, the system acquires a series of 360° images. The YOLO11l model, trained on single-view images, achieves a mean absolute error (MAE) of 0.61 and a relative root mean square error (rRMSE) of 0.09, with an F1-score of 0.96 for classifying larval and adult stages. An explainability analysis (XAI) using Grad-CAM suggests that the model employs morphological criteria similar to those used by experts for stage classification. The number of aphids in each image of the series, combined with mathematical estimators, provides an estimate of the total aphid count per plant. As a modular and open system, AphiDetect is adaptable to other legumes and pests. Future integration with multi-omics data could refine our understanding of plant resistance mechanisms, paving the way for new data-driven breeding strategies and a renewed understanding of plant-aphid interactions.
ano.nymous@ccsd.cnrs.fr.invalid (Océane Gourdin) 18 Sep 2026
https://hal.inrae.fr/hal-05754376v1
-
[hal-05673352] EweAcT: Ewe behaviour aligned to accelerometer data for activity monitoring in extensive grazing systems.
Monitoring livestock behaviour under extensive conditions would provide valuable insights to assess animal adaption to environmental perturbations in agroecological systems (e.g., heat waves, parasitism, predator attacks). Animal behaviour can be monitored using accelerometer data collected from neck-collars combined with artificial intelligence models. However, large amounts of accelerometer data aligned with annotated behaviours are necessary to develop accurate models of behaviour prediction. In particular, developing reliable models for extensive systems requires data collected across a wide range of representative conditions. The dataset includes 79 hours of tri-axial accelerometer data aligned with behaviours manually annotated from video recordings for 120 Romane ewes born between 2021 and 2024. The ewes were derived from two divergent genetic lines after three and four generations of selection started 10 years ago: low and high social attractiveness, noted S-and S+, and low and high tolerance towards humans, noted H-and H+. They were reared under the extensive system applied to the Experimental Unit of La Fage (UEF, INRAE, Saint-Jean-et-saint Paul, Aveyron) where 250 sheep were reared exclusively outdoors on 280 hectares of rangeland in southern France. First batch of data was collected on March, June and July 2024 at the UEF under a range of extensive conditions, including sloping pastures and heat-wave periods. Ewes were equipped with accelerometer neck-collars specifically designed for young sheep on pasture. They were grouped on experimental paddocks for 4 to 8 hours and provided with fresh grass and ad libitum access to water. The animals were simultaneously video-recorded using an elevated CCTV camera. Behaviour annotation was carried out using Behavioral Observation Research Interactive Software focusing on the main behaviours on pasture: Grazing, Ruminating, Resting, Moving, and "Other", grouping all remaining activities. Annotations and corresponding accelerometer sequences were aligned using Python language, based on a time synchronization procedure. A second batch of data was acquired on November 2025 to supplement the dataset with the moving activity. For that purpose, ewes were equipped with the accelerometer collars and moved on tracks from the housing area to the pastures, corresponding to an approximately 10 minute-walk. The start and end times of the moves for each ewe were used to align the corresponding accelerometer data with the moving activity. These data were then merged with the dataset from the first batch. The resulting dataset is ready to use for applying artificial intelligence models to classify the 5 main behaviours of sheep under extensive grazing systems from accelerometer data.
ano.nymous@ccsd.cnrs.fr.invalid (Lucile Riaboff) 01 Jul 2026
https://hal.inrae.fr/hal-05673352v1
-
[hal-05688139] Joint inference of selection and demography from genomic time series
[...]
ano.nymous@ccsd.cnrs.fr.invalid (Paul Bunel) 14 Jul 2026
https://hal.science/hal-05688139v1
-
[hal-05686843] Lossless compression of k-mer matrices enabling random row access
<div><p>Genomic search engines such as Logan-Search index petabytes of sequencing data as large binary matrices, called k-mer matrices, where each row encodes the presence of a k-mer across thousands to millions of genomic samples. Logan-Search contains a petabyte of binary matrices, and storing them is expensive, yet compression must not prevent fast random access to any matrix row at query time. We present kmcomp, a lossless compression method for k-mer matrices that satisfies these competing requirements. Block compression partitions the matrix into fixed-size row blocks, each compressed independently; block start positions are stored in an Elias-Fano encoded array, enabling O(1) random access to any block. To improve compressibility without introducing additional decompression steps, we introduce the π-compression: a column reordering that groups similar samples together by solving the Traveling Salesman Problem via a nearest-neighbor heuristic. We accelerate this heuristic with a novel variant of the vantage-point tree, the masked vp-tree, which dynamically prunes nearest-neighbor search space. On three (meta)genomic datasets, kmcomp achieves compression ratios of 1.3 to 5.4; π-compression further improves these to 1.5 to 51.3. Applied to the Logan-Search petabyte-scale index, compression reduces storage by approximately half, and π-compression adds a further 13% gain. Query overhead remains modest: queries of hundreds of nucleotides incur an absolute latency increase of ≈ 100 ms, and highly compressed indexes can match uncompressed query times thanks to reduced disk reads.</p></div>
ano.nymous@ccsd.cnrs.fr.invalid (Alix Regnier) 09 Jul 2026
https://hal.science/hal-05686843v1
-
[hal-05665434] Horloges épigénétiques pour la longévité fonctionnelle des vaches laitières
Chez les bovins lait, la longévité fonctionnelle est un caractère d’intérêt majeur sur les plans économique, environnemental et du bien-être animal mais elle n’est mesurable que tardivement dans la vie des individus. L’objectif de ce stage était d’évaluer si l’âge épigénétique, prédit à partir de la méthylation de l’ADN pouvait constituer un biomarqueur précoce du vieillissement et de la longévité fonctionnelle. Le méthylome sanguin de 4 751 vaches Holstein a été analysé à l’aide de la puce RUMIGEN EpiChip, ciblant 44 053 sites CpGs. Plusieurs modèles de prédiction ont été comparés afin de construire une horloge épigénétique. Le meilleur modèle a prédit l’âge épigénétique des animaux avec une précision de 129 jours. L’écart entre l’âge épigénétique et l’âge réel a été calculé pour chaque animal. Dans les modèles analysés, cet écart présentait une héritabilité modérée et était associé à plusieurs régions génomiques où plusieurs gènes candidats ont été identifiés (dont DNMT3B et RCAN2). De plus, cet écart a permis d’observer que l’accélération du vieillissement épigénétique était associée à une réforme précoce des animaux. Ces résultats montrent qu’une horloge épigénétique peut être construite chez la vache laitière et que l’âge épigénétique pourrait constituer un éventuel biomarqueur précoce de la longévité fonctionnelle. Ces approches sont très prometteuses pour être appliquées à d’autres contextes et à différentes échelles d’étude.
ano.nymous@ccsd.cnrs.fr.invalid (Margaux Gaury) 22 Jun 2026
https://hal.inrae.fr/hal-05665434v1
-
[hal-05643765] DRL-Based Pose Control for Double-Ackermann Robots Under Actuation Uncertainties
Robust deployment of deep reinforcement learning (DRL) policies on real robots remains challenging due to discrepancies between simulation and real-world dynamics. We address this issue in the context of maneuvering with double-Ackermann-steering mobile robots, which introduce additional constraints due to their non-holonomic nature. Building upon the DRL framework ManeuverNet, we extend its objective from position control to full pose control, resulting in a more challenging task. We further investigate the impact of actuationrelated uncertainties on policy transfer. The use of simplified actuation models during training of the extended policy can lead to poor generalization, shown by a success rate drop from 100% in PyBullet to 25% in Gazebo under stricter evaluation conditions. To address this limitation, we adopt a sim-to-simto-real approach, where actuation effects observed in Gazebo are incorporated into the PyBullet training environment. Using multi-environment DRL with SAC and CrossQ, we learn policies that remain robust despite modeling inaccuracies. This approach can significantly reduce the performance gap across simulators, achieving up to 92% success rate in Gazebo and maintaining 69% under stricter thresholds, with successful transfer to a real robot without additional tuning.
ano.nymous@ccsd.cnrs.fr.invalid (Oussama Zaim) 04 Jun 2026
https://hal.science/hal-05643765v1
-
[hal-05578828] Use of Surface Water and Ocean Topography (SWOT) observations to support Land Use/Land Cover (LULC) change products: the case of the pacific coast of Ecuador
<div><p>Radar altimetry has been used to characterize land surfaces. However, the nadir configuration of the radar altimeter sensor and its coarse spatial resolution were limiting factor. The Surface Water and Ocean Topography (SWOT) mission overcomes these limitations through its Ka-band Radar Interferometer (KaRIn), a synthetic aperture radar (SAR) system, providing high spatial resolution and accurate surface height measurements. Initially used for hydrology and oceanography, this study explores an innovative use of SWOT to analyse changes in Land Use/Land Cover (LULC). To do this, three study areas located on the Pacific Coast of Ecuador were considered. The area in the south (A) is characteristic of cultivated areas, while the area in the center (B) presents a landscape mosaic and the area in the north (C) hosts tropical rainforests. For each study area, the SWOT backscatter coefficient (sig0) was analysed for the year 2024 from the raster product at 100m spatial resolution. We calculated the number of occurrences and the sig0 average from the raster product over each pixel. The spatial patterns obtained from these two variables enabled us to assign a LULC class (city, water, road, crop, or no forest, depending on the study area) to each pixel, using a Support Vector Machine (SVM). The assigned LULC classes depend on the partial spatial coverage of the SWOT data, which does not allow representing all the LULC classes. The classification results were compared with the LULC map provided by the Ministry of the Environment using a confusion matrix and obtained an accuracy greater than 0.87 and an F1 score greater than 0.89 for the three study areas. In the forest area (C), the SWOT observations were also compared to two change detection products: RAdar for Detecting Deforestation (RADD) alerts and detections by the Cumulative Sum (CuSum) method. 39% of the SWOT observations were in areas identified as forest by these products but classified as no forest or water in our SWOT classification. By detecting small streams (areas A, B and C), roads (area A), the boundaries of agricultural plots and the state of cultivated land (area A) as well as recent forms of deforestation (zone C), SWOT was found to be a complementary source of information for LULC change products.</p></div>
ano.nymous@ccsd.cnrs.fr.invalid (Valentine Sollier) 03 Apr 2026
https://hal.inrae.fr/hal-05578828v1
-
[hal-05643469] OrthoViewer
OrthoViewer (interactive exploration of orthologous gene families). OrthoViewer is a publicly accessible web platform that enables real-time exploration of orthologous gene families across 84 angiosperm species. It integrates phylogenetic trees, eggNOG-mapper functional annotations (GO, SO, CoG namespaces) and a manually curated layer of reference genes with associated literature, into a single interactive query environment.
ano.nymous@ccsd.cnrs.fr.invalid (Raphaël Flores) 17 Jun 2026
https://hal.inrae.fr/hal-05643469v1
-
[hal-05651322] Spécificité des associations génétiques en contexte multi-populations : une étude par simulation
Les études d'association à l'échelle du génome (GWAS) sont cruciales pour comprendre les liens entre variations génétiques et phénotypes complexes d'intérêt. En agriculture, la sélection a induit la structuration des espèces agricoles en sous-populations (races ou écotypes) hautement spécialisées. À ce jour, peu d'études ont évalué l'impact conjoint de cette structuration et du choix du modèle sur les résultats d'un GWAS. Dans cette étude, nous proposons un cadre de simulation réaliste dans le but d'évaluer la performance des méthodes existantes pour détecter des associations spécifiques à une population ou capturer des effets partagés ou non entre sous-populations. Les résultats obtenus dans cette étude mettent en évidence les limites des approches classiques fondées sur des modèles linéaires mixtes, notamment pour la détection de variants présentant des effets hétérogènes ou opposés entre populations.
ano.nymous@ccsd.cnrs.fr.invalid (Kossi Julien Kowou) 10 Jun 2026
https://hal.science/hal-05651322v1
-
[hal-05665454] Epigenetics clocks for functional longevity of dairy cows
Chez les bovins laitiers la longévité fonctionnelle se définit comme la capacité d’un individu à rester productif dans le troupeau tout en maintenant un bon état de santé et des performances satisfaisantes au cours de sa vie productive (Schuster et al, 2020). L’augmentation de la durée de vie fonctionnelle des individus permettrait de diminuer les besoins de renouvellement du troupeau et de réduire les impacts environnementaux et économiques associés. La longévité fonctionnelle n’est mesurable que tard dans la vie d’un animal et disposer de biomarqueurs précoces permettrait d’accélérer le progrès génétique sur ce caratère. On appelle âge épigénétique l’âge d’un animal prédit via une horloge épigénétique, un algorithme utilisant des données de méthylations de l’ADN comme variables prédictives (Horvath, 2013). L’âge épigénétique pourrait-il constituer un biomarqueur précoce de la longévité fonctionnelle chez les bovins laitiers ? Pour construire l’horloge épigénétique, nous avons analysé le méthylome sanguin de 4 751 vaches Holstein, à l’aide d’une puce de méthylation bovine (RUMIGEN EpiChip), mesurant la méthylation de 44 053 sites CpG. Nous avons comparé plusieurs méthodes d’apprentissage en variant : (1) la nature des variables prédictives (matrices de méthylation de l’ADN ou valeurs propres issues d’une analyse en composante principale, données corrigées de co-variables ou non), et (2) la méthode de prédiction (régressions linéaires pénalisées ou forêts aléatoires). Les modèles les plus performant prédisent un âge épigénétique fortement corrélé à l’âge réel des animaux lors du prélèvement (RMSE = 128.9 jours, R² = 0.95). Nous avons prédit un âge épigénétique pour l’ensemble des échantillons du dispositif, en utilisant les paramètres des meilleurs modèles et en utilisant différentes répartitions des animaux dans les jeux de données d’apprentissage et de test. Cela nous a permis de calculer des écarts d’âges entre l’âge prédit et l’âge réel pour chaque animal de la cohorte. L’héritabilité des écarts d’âges est modérée (h² > 0,2) et plusieurs régions du génome (QTLs) associées à ce phénotype ont été détectés. L’identification des gènes candidats dans chacun de ces QTLs est en cours. L’étape suivante consistera à estimer la corrélation entre cet écart d’âge et différents phénotypes de production, de santé, et de longévité mesurés sur les mêmes animaux. Ces travaux confirment qu’il est possible d’entraîner des horloges épigénétiques chez la vache laitière avec une bonne précision, et nous espérons qu’ils nous apportent des connaissances nouvelles sur l’efficacité de ce biomarqueur pour prédire la longévité fonctionnelle des animaux.
ano.nymous@ccsd.cnrs.fr.invalid (Margaux Gaury) 22 Jun 2026
https://hal.inrae.fr/hal-05665454v1
-
[hal-05682255] Towards lucerne varieties used as living mulch for cereal crops in agroecological systems
Lucerne, a perennial legume known for its nitrogen fixation, persistence, and soil-covering capacity, shows strong potential as a living mulch for cereal cropping. However, its vigorous growth often results in excessive competition with cash crops. The selection of lucerne varieties adapted to living mulch could be a solution to reduce this competition. We synthetize the state of the art on this subject. Wheat–lucerne interactions occur from the earliest stages of wheat cycle until its harvest and are mainly driven by lucerne morphological and phenological traits. Autumn dormancy, growth habit, height, and cover state of lucerne determine the trade-off between reducing competition with wheat and maintaining the ecosystem services provided by lucerne. An intermediate dormancy, combined with moderate height and upright cover, appears to provide the most favourable balance. Genetic correlations between traits measured in spaced plants and living mulch conditions reveal that some traits, such as height, remain stable across designs, whereas others are highly design-dependent. This supports a two-step breeding strategy combining early indirect selection in nursery of spaced plants with an indirect selection under living mulch conditions. Finally, molecular markers used for genomic prediction could accelerate the identification of genotypes suited for living mulch systems. This knowledge can be used to create dedicated varieties.
ano.nymous@ccsd.cnrs.fr.invalid (Zineb El Ghazzal) 06 Jul 2026
https://hal.inrae.fr/hal-05682255v1
-
[hal-05639710] Modelling soil microbial functions at large spatial scale based on metagenomic dimensionality reduction
[...]
ano.nymous@ccsd.cnrs.fr.invalid (Emna Stambouli) 01 Jun 2026
https://inria.hal.science/hal-05639710v1