Filtern
Dokumenttyp
- Forschungsdatensatz (12)
- Zeitschriftenartikel (10)
- Preprint (5)
- Posterpräsentation (3)
- Vortrag (1)
- Sonstiges (1)
Schlagworte
- Automation (12)
- Database (12)
- Bonding analysis (11)
- Computational Chemistry (11)
- Materials Informatics (11)
- Machine Learning (8)
- Materials Design (5)
- Phonons (4)
- Bonding Analysis (3)
- Materials Acceleration Platforms (3)
Organisationseinheit der BAM
- VP Vizepräsident (29)
- VP.1 eScience (29)
- 6 Materialchemie (24)
- 6.0 Abteilungsleitung und andere (18)
- 6.6 Digitale Materialchemie (7)
- 1 Analytische Chemie; Referenzmaterialien (2)
- 1.7 Organische Spuren- und Lebensmittelanalytik (2)
- 6.5 Synthese und Streuverfahren nanostrukturierter Materialien (2)
- 1.8 Umweltanalytik (1)
- 7 Bauwerkssicherheit (1)
Eingeladener Vortrag (wissenschaftliche Konferenzen)
- nein (1)
Understanding the chemistry and nature of individual chemical bonds is essential for materials design. Bonding analysis via the LOBSTER software package has provided valuable insights into the properties of materials for thermoelectric and catalysis applications. Thus, the data generated from
bonding analysis becomes an invaluable asset that could be utilized as features in large-scale data analysis and machine learning of material properties. However, no systematic studies exist that conducted high-throughput materials simulations to curate and validate bonding data obtained from LOBSTER. Here we present an approach to constructing such a large database consisting of quantum-chemical bonding information.
We demonstrate a strategy for simulating wide-range X-ray scattering patterns, which spans the small- and wide scattering angles as well as the scattering angles typically used for Pair Distribution Function (PDF) analysis. Such simulated patterns can be used to test holistic analysis models, and, since the diffraction intensity is on the same scale as the scattering intensity, may offer a novel pathway for determining the degree of crystallinity.
The "Ultima Ratio" strategy is demonstrated on a 64-nm Metal Organic Framework (MOF) particle, calculated from Q < 0.01 1/nm up to Q < 150 1/nm, with a resolution of 0.16 Angstrom. The computations exploit a modified 3D Fast Fourier Transform (3D-FFT), whose modifications enable the transformations of matrices at least up to 8000^3 voxels in size. Multiple of these modified 3D-FFTs are combined to improve the low-Q behaviour. The resulting curve is compared to a wide-range scattering pattern measured on a polydisperse MOF powder. While computationally intensive, the approach is expected to be useful for simulating scattering from a wide range of realistic, complex structures, from (poly-)crystalline particles to hierarchical, multicomponent structures such as viruses and catalysts.
Most machine learning models for materials science rely on descriptors based on materials compositions and structures, even though the chemical bond has been proven to be a valuable concept for predicting materials properties. Over the years, various theoretical frameworks have been developed to characterize bonding in solid‐state materials. However, integrating bonding information from these frameworks into machine learning pipelines at scale has been limited by the lack of a systematically generated and validated database. Recent advances in high‐throughput bonding analysis workflows have addressed this issue, and our previously computed Quantum‐Chemical Bonding Database for Solid‐State Materials was extended to include approximately 13,000 materials. This database is then used to derive a new set of quantum‐chemical bonding descriptors. A systematic assessment is performed using statistical significance tests to evaluate how the inclusion of these descriptors influences the performance of machine‐learning models that otherwise rely solely on structure‐ and composition‐derived features. Models are built to predict elastic, vibrational, and thermodynamic properties typically associated with chemical bonding in materials. The results demonstrate that incorporating quantum‐chemical bonding descriptors not only improves predictive performance but also helps identify intuitive expressions for properties such as the projected force constant and lattice thermal conductivity via symbolic regression.
A high-resolution spatiotemporal wildfire propagation dataset for the Mediterranean and Europe
(2026)
Wildfires are becoming more frequent and severe under the influence of climate change, posing increasing risks to ecosystems, human health, and infrastructure. Accurate spatiotemporal data on wildfire propagation is essential for advancing fire behavior modeling, improving management strategies, and mitigating future impacts. However, existing datasets with both high spatial and temporal resolution are rare, costly, and time-consuming to produce. To address this gap, we present FireSpread_MedEU, a dataset comprising 320 consecutive burned area maps from 103 wildfire events across the Mediterranean and Europe between 2017 and 2023. Burned areas were derived from high-resolution Planet optical satellite imagery (~3 m spatial, mostly daily temporal resolution) using a semi-automated workflow, followed by manual refinement to ensure highest accuracy. Each dataset entry is enriched with detailed metadata and a subjective quality assessment. With its high level of spatiotemporal precision, FireSpread_MedEU provides essential data for the development and validation of machine learning models or wildfire simulation models. It opens new research opportunities in wildfire behavior analysis, risk assessment, and predictive modeling.
Machine-learning interatomic potentials are widely used as computationally efficient surrogates for density functional theory in atomistic simulations, enabling large-scale, long-time modeling of materials systems. We investigate how different fine-tuning strategies influence the prediction of harmonic phonon band structures, thermal properties, and the potential energy surface along imaginary phonon modes. We achieve substantial accuracy improvements with minimal additional data, with as few as 10 additional training structures already yielding significant gains. In addition to existing approaches, we introduce Equitrain, a finetuning framework that implements LoRA-based adaptation. Across 53 materials systems, we show that fine-tuned models consistently outperform both the underlying pretrained model and models trained from scratch. Equitrain achieves the best overall performance, and our results demonstrate that fine-tuning enables accurate phonon predictions.
Most machine learning models for materials science rely on descriptors based on materials compositions and structures, even though the chemical bond has been proven to be a valuable concept for predicting materials properties. Over the years, various theoretical frameworks have been developed to characterize bonding in solid-state materials. However, integrating bonding information from these frameworks into machine learning pipelines at scale has been limited by the lack of a systematically generated and validated database. Recent advances in high-throughput bonding analysis workflows have addressed this issue, and our previously computed Quantum-Chemical Bonding Database for Solid-State Materials was extended to include approximately 13,000 materials. This database is then used to derive a new set of quantum-chemical bonding descriptors. A systematic assessment is performed using statistical significance tests to evaluate how the inclusion of these descriptors influences the performance of machine-learning models that otherwise rely solely on structure- and composition-derived features. Models are built to predict elastic, vibrational, and thermodynamic properties typically associated with chemical bonding in materials. The results demonstrate that incorporating quantum-chemical bonding descriptors not only improves predictive performance but also helps identify intuitive expressions for
properties such as the projected force constant and lattice thermal conductivity via symbolic regression.
Phonon calculations with ab-initio methods are computationally expensive. The use of universal machine learning models reduces the cost, but raises concerns about prediction quality. Fine-tuning with only a few structures, improves predictions of phonons, thermal properties and especially diffusive thermal conductivity, while reducing computational cost by a factor of 10 in average compared to DFT methods.
The rapid adoption of machine learning in various scientific domains calls for the development of best practices and community agreed-upon benchmarking tasks and metrics. We present Matbench Discovery as an example evaluation framework for machine learning energy models, here applied as pre-filters to first-principles computed data in a high-throughput search for stable inorganic crystals. We address the disconnect between (1) thermodynamic stability and formation energy and (2) retrospective and prospective benchmarking for materials discovery. Alongside this paper, we publish a Python package to aid with future model submissions and a growing online leaderboard with adaptive user-defined weighting of various performance metrics allowing researchers to prioritize the metrics they value most. To answer the question of which machine learning methodology performs best at materials discovery, our initial release includes random forests, graph neural networks, one-shot predictors, iterative Bayesian optimizers and universal interatomic potentials. We highlight a misalignment between commonly used regression metrics and more task-relevant classification metrics for materials discovery. Accurate regressors are susceptible to unexpectedly high false-positive rates if those accurate predictions lie close to the decision boundary at 0 eV per atom above the convex hull. The benchmark results demonstrate that universal interatomic potentials have advanced sufficiently to effectively and cheaply pre-screen thermodynamic stable hypothetical materials in future expansions of high-throughput materials databases.
Material discovery is a cornerstone of modern science, driving advancements in diverse disciplines from biomedical technology to climate solutions. Predicting synthesizability, a critical factor in realizing novel materials, remains a complex challenge due to the limitations of traditional heuristics and thermodynamic proxies. While stability metrics such as formation energy offer partial insights, they fail to account for kinetic factors and technological constraints that influence synthesis outcomes. These challenges are further compounded by the scarcity of negative data, as failed synthesis attempts are often unpublished or context-specific. We present SynCoTrain, a semi-supervised machine learning model designed to predict the synthesizability of materials. SynCoTrain employs a co-training framework leveraging two complementary graph convolutional neural networks: SchNet and ALIGNN. By iteratively exchanging predictions between classifiers, SynCoTrain mitigates model bias and enhances generalizability. Our approach uses Positive and Unlabeled (PU) learning to address the absence of explicit negative data, iteratively refining predictions through collaborative learning. The model demonstrates robust performance, achieving high recall on internal and leave-out test sets. By focusing on oxide crystals, a well-characterized material family with extensive experimental data, we establish SynCoTrain as a reliable tool for predicting synthesizability while balancing dataset variability and computational efficiency. This work highlights the potential of co-training to advance high-throughput materials discovery and generative research, offering a scalable solution to the challenge of synthesizability prediction.