Filtern
Dokumenttyp
- Zeitschriftenartikel (1)
- Posterpräsentation (1)
- Preprint (1)
Sprache
- Englisch (3)
Schlagworte
- Synthesizability (3)
- Materials Acceleration Platforms (2)
- Materials Design (2)
- Materials Discovery (2)
- Accelerated Materials Design (1)
- Cheminformatics (1)
- Co-training (1)
- Cotraining (1)
- Machine Learning (1)
- New Materials (1)
Organisationseinheit der BAM
Material discovery is a cornerstone of modern science, driving advancements in diverse disciplines from biomedical technology to climate solutions. Predicting synthesizability, a critical factor in realizing novel materials, remains a complex challenge due to the limitations of traditional heuristics and thermodynamic proxies. While stability metrics such as formation energy offer partial insights, they fail to account for kinetic factors and technological constraints that influence synthesis outcomes. These challenges are further compounded by the scarcity of negative data, as failed synthesis attempts are often unpublished or context-specific. We present SynCoTrain, a semi-supervised machine learning model designed to predict the synthesizability of materials. SynCoTrain employs a co-training framework leveraging two complementary graph convolutional neural networks: SchNet and ALIGNN. By iteratively exchanging predictions between classifiers, SynCoTrain mitigates model bias and enhances generalizability. Our approach uses Positive and Unlabeled (PU) learning to address the absence of explicit negative data, iteratively refining predictions through collaborative learning. The model demonstrates robust performance, achieving high recall on internal and leave-out test sets. By focusing on oxide crystals, a well-characterized material family with extensive experimental data, we establish SynCoTrain as a reliable tool for predicting synthesizability while balancing dataset variability and computational efficiency. This work highlights the potential of co-training to advance high-throughput materials discovery and generative research, offering a scalable solution to the challenge of synthesizability prediction.
In the pursuit of discovering materials with desirable properties, extending the available material libraries is crucial. High-throughput simulations have become an integral part in designing new materials in the past decades. However, there is no straightforward way of distinguishing synthesizable materials from all the proposed candidates. This project focuses on employing AI-driven methods to estimate synthesizability of materials.
Up to now, material scientists and engineers have relied on domain knowledge as well as empirical heuristics to guess the stability and synthesizability of molecules and crystals. The famous Pauling rules of crystal stability are an example of such heuristics. However, after the accelerating material discovery in all the years since Pauling, these rules now fail to account for the stability of most known crystals. A new predictive set of heuristics for crystal stability/synthesizability is unlikely to be uncovered by human perception, given the magnitude and dimensionality of crystallographic data. Hence, a data-driven approach should be proposed to find a predictive model or set of heuristics which differentiate synthesizable crystal structures from the rest. The main challenge of this research problem is the lack of a negative set for classification. Here, there are two classes of data: the positive class which contains synthesizable materials and the negative class which contains materials which are not synthesizable. While the data from the positive class is simply the data of crystals which have been experimentally synthesized, we do not have access to data points which are certainly unsynthesizable. Strictly speaking, if an attempt of synthesizing a crystal fails, it does not necessarily follow that the crystal is not synthesizable. Also, there is no database available which contains the intended crystal structures of unsuccessful synthesis attempts.
This project proposes a semi-supervised learning scheme to predict crystal synthesizability. The ML model is trained on experimental and theoretical crystal data. The initial featurization focuses on local environments which is inspired by the Pauling Rules. The experimental data points are downloaded through the Pymatgen API from the Materials Project database which contains relaxed structures recorded in Inorganic Crystal Structure Database – ICSD. The theoretical data is queried from select databases accessible through the Optimade project’s API.
Material discovery is a cornerstone of modern science, driving advancements in diverse disciplines from biomedical technology to climate solutions. Predicting synthesizability, a critical factor in realizing novel materials, remains a complex challenge due to the limitations of traditional heuristics and thermodynamic proxies. While stability metrics such as formation energy offer partial insights, they fail to account for kinetic factors and technological constraints that influence synthesis outcomes. These challenges are further compounded by the scarcity of negative data, as failed synthesis attempts are often unpublished or context-specific.
We present SynCoTrain, a semi-supervised machine learning model designed to predict the synthesizability of materials. SynCoTrain employs a co-training framework leveraging two complementary graph convolutional neural networks: SchNet and ALIGNN. By iteratively exchanging predictions between classifiers, SynCoTrain mitigates model bias and enhances generalizability. Our approach uses Positive and Unlabeled (PU) Learning to address the absence of explicit negative data, iteratively refining predictions through collaborative learning.
The model demonstrates robust performance, achieving high recall on internal and leave-out test sets. By focusing on oxide crystals, a well-characterized material family with extensive experimental data, we establish SynCoTrain as a reliable tool for predicting synthesizability while balancing dataset variability and computational efficiency. This work highlights the potential of co-training to advance high-throughput materials discovery and generative research, offering a scalable solution to the challenge of synthesizability prediction.