EventsThe 1st International Online Conference on Earth Science
Published
This submission belongs to the session S1. AI and Big Data in Earth Science of the event The 1st International Online Conference on Earth Science
Published date
31 Aug, 2026
Academic Editor
author-avatarEliseo Clementini
Citation
Daniel Keim Almeida, Gustavo Taboada Soldati, Leonardo Goliatt da Fonseca, Large Language Models as Tools for Biodiversity Traceability and Genetic Resource Monitoring in Brazil, in Proceedings of The 1st International Online Conference on Earth Science, 2 September–4 September 2026, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

Large Language Models as Tools for Biodiversity Traceability and Genetic Resource Monitoring in Brazil

Gustavo Taboada Soldati 2
image
1. Federal University of Juiz de Fora, Juiz de Fora, Brazil
2. Postgraduate Program in Biodiversity and Nature Conservation, Federal University of Juiz de Fora, Juiz de Fora, Brazil
3. Department of Applied and Computational Mechanics , Federal University of Juiz de Fora, Juiz de Fora, Brazil
Abstract

Introduction. The Convention on Biological Diversity, sustained by its Nagoya Protocol, defines that interstate genetic patrimony exchange relies on both national and international processes, considering accordance and traceability to regulate access and ensure that derived benefits are fairly shared with provider countries and indigenous people and local communities. In Brazil, genetic patrimony international shipment to another countries requires registration on the National System for the Management of Genetic Heritage and Associated Traditional Knowledge (SisGen), from Brazil's Environmental and Climate Change Ministry, where researchers formalize all activities related to access, shipment, or economic use of Brazilian genetic resources. However, these records are self-declared and some information is added without constraints, increasing complexity comprehension and frequently leading to incomplete descriptions of resource access. Over recent years, thousands of records piled up, making manual classification and information structuring unfeasible, as it would require excessive time and human resources.\vspace{0.18cm}

Methods. This study aims to investigate the application of Large Language Models to convert unconstrained SisGen text fields into structured, machine-readable information. These models were evaluated using a reference dataset of manually annotated records, enabling a comparative assessment of their ability to identify key attributes related to material use, destination, and research intent. \vspace{0.18cm}

Results. Results indicate that the models differ in performance significantly, particularly with noisy or incomplete text. While some models achieved robust and consistent extraction, others were more sensitive to ambiguity and missing context. \vspace{0.18cm}

Conclusion. Based on these findings, LLMs can provide a viable and scalable approach to converting unstructured regulatory text data into structured formats at the national level. This solution supports biodiversity monitoring, enhances environmental governance, and enables large-scale data-driven analyzes and policy support for the national system.

Keywords
Large Language Models
Genetic Heritage
Biodiversity Traceability
Natural Language Processing
Automated data extraction
Multi-target Evolutionary Hyperparameter Optimization for Multi-task Predictive Modeling in Biomass Pyrolysis
Predicting Soil Organic Carbon Response to Biochar Using Parsimonious Machine Learning Models