Researchers Released Large Dataset of Asymmetric Reactions
The 4,015-reaction dataset provides machine-readable data to improve AI-driven prediction of enantioselectivity.
Updated on Sept. 25, 2026 in Chemistry

Researchers have published a new machine-readable dataset containing 4,015 asymmetric organocatalytic Mannich reactions. This resource compiles data from 153 scientific publications spanning 2000 to 2025 to support machine learning models.
Why it matters
The scarcity of structured experimental data has historically hindered the development of machine learning methods for predicting enantioselectivity. This dataset aims to accelerate progress in computational chemistry by providing standardized inputs for training predictive models.
The dataset includes 4,015 reaction entries featuring reagents, catalysts, and reaction conditions alongside target values for enantiomeric excess and Gibbs free energy difference. This provides a significantly larger structured baseline than the isolated data points found in previous literature.
The details
The researchers manually extracted reaction data from 153 individual scientific studies to create a machine-readable format suitable for AI training. By standardizing reaction variables—such as the specific catalyst, reagents, and conditions—the dataset allows models to better correlate chemical inputs with target outputs like enantiomeric excess (a measure of purity for chiral molecules) and Gibbs free energy difference (the energy change that determines reaction spontaneity).
Timeline
2000-2025: Range of scientific publications used for data extraction.
September 25, 2026: Official publication date of the research dataset.
The Tech Race
This release marks a significant move to formalize experimental protocols for computational use, following a pattern set by the broader push toward AI-driven materials discovery. By bridging the gap between historical bench research and modern computation, it aims to outpace current manual trial-and-error workflows.
This dataset is immediately available for computational chemists and software developers building predictive models for chemical synthesis. It removes the need for individual teams to manually parse thousands of papers, likely streamlining the development of future high-precision chemical catalysts.
The takeaway
This publication provides a critical foundation for chemists to train more reliable, data-driven predictive tools for complex chemical synthesis. Researchers should monitor future publications for model benchmarks that utilize this specific 4,015-reaction set as a training input.
Further reading
For more developments in this field, visit the Chemistry section.
Source note: This article includes information reported by Nature.






