Filtern
Dokumenttyp
- Zeitschriftenartikel (2)
- Vortrag (2)
- Posterpräsentation (1)
Sprache
- Englisch (5)
Schlagworte
- Bonding analysis (3)
- Machine learning (3)
- Materials Descriptors (3)
- Bonding Analysis (2)
- Anharmonicity (1)
- Bonding Heterogeneity (1)
- Elastic Properties (1)
- Force constants (1)
- High-throughput (1)
- Large Language Models (1)
Organisationseinheit der BAM
- 6 Materialchemie (5)
- 6.6 Digitale Materialchemie (5)
- VP Vizepräsident (1)
- VP.1 eScience (1)
Eingeladener Vortrag (wissenschaftliche Konferenzen)
- nein (2)
Most machine learning models for materials science rely on descriptors based on materials compositions and structures, even though the chemical bond has been proven to be a valuable concept for predicting materials properties. Over the years, various theoretical frameworks have been developed to characterize bonding in solid‐state materials. However, integrating bonding information from these frameworks into machine learning pipelines at scale has been limited by the lack of a systematically generated and validated database. Recent advances in high‐throughput bonding analysis workflows have addressed this issue, and our previously computed Quantum‐Chemical Bonding Database for Solid‐State Materials was extended to include approximately 13,000 materials. This database is then used to derive a new set of quantum‐chemical bonding descriptors. A systematic assessment is performed using statistical significance tests to evaluate how the inclusion of these descriptors influences the performance of machine‐learning models that otherwise rely solely on structure‐ and composition‐derived features. Models are built to predict elastic, vibrational, and thermodynamic properties typically associated with chemical bonding in materials. The results demonstrate that incorporating quantum‐chemical bonding descriptors not only improves predictive performance but also helps identify intuitive expressions for properties such as the projected force constant and lattice thermal conductivity via symbolic regression.
Examining the bonding between their constituent atoms in crystalline materials has played a vital role in understanding material properties. For instance, low thermal conductivity in materials is typically attributed to its anharmonicity, which has been reported to arise from strong antibonding interactions and local environment distortions. Employing an automated for bonding analysis that we developed, we have generated for ~13000 crystalline compounds such bonding analysis data. To create new descriptors from these data automatically, we extended our package LobsterPy. The curated descriptors span different types, including statistical representations of bonding characteristics for traditional ML algorithms (e.g., random forests), textual descriptions for large language models (LLMs), and structure graphs for graph neural networks (GNNs). These descriptors are then tested by employing them in several state-of-the-art ML algorithms and architectures to predict the mechanical, vibrational, and thermal properties of crystalline materials. Through this work, we are not only able to demonstrate how one can enhance the model’s predictive accuracy by incorporating quantum chemical bonding-based descriptors alongside typical composition and structure-based descriptors, but it also aids in uncovering relationships between bonding and materials properties on a larger scale, which was not possible before.
Large Language Models (LLMs) are reshaping many aspects of materials science and chemistry research, enabling advances in molecular property prediction, materials design, scientific automation, knowledge extraction, and more. Recent developments demonstrate that the latest class of models are able to integrate structured and unstructured data, assist in hypothesis generation, and streamline research workflows. To explore the frontier of LLM capabilities across the research lifecycle, we review applications of LLMs through 34 total projects developed during the second annual Large Language Model Hackathon for Applications in Materials Science and Chemistry, a global hybrid event. These projects spanned seven key research areas: (1) molecular and material property prediction, (2) molecular and material design, (3) automation and novel interfaces, (4) scientific communication and education, (5) research data management and automation, (6) hypothesis generation and evaluation, and (7) knowledge extraction and reasoning from the scientific literature. Collectively, these applications illustrate how LLMs serve as versatile predictive models, platforms for rapid prototyping of domain-specific tools, and much more. In particular, improvements in both open source and proprietary LLM performance through the addition of reasoning, additional training data, and new techniques have expanded effectiveness, particularly in low-data environments and interdisciplinary research. As LLMs continue to improve, their integration into scientific workflows presents both new opportunities and new challenges, requiring ongoing exploration, continued refinement, and further research to address reliability, interpretability, and reproducibility.
The concept of the chemical bond has long served chemists in rationalizing material properties,[1]reaction pathways,[2] or crystal structure stability.[3,4] Despite several theoretical frameworks being developed over the years to characterize bonding in solid-state materials,[5–7] a comprehensive assessment of the impact of incorporating quantum chemical bonding descriptors into machine learning studies of material properties has remained elusive, partly due to the lack of data. To overcome this issue, a quantum-chemical bonding analysis workflow[8,9] was developed, enabling the high-throughput computation of orbital-based bonding descriptors derived from ab initio calculations. By utilizing this workflow, we have constructed a database of bonding descriptors for approximately 13,000 structures from the Materials Project. A total of 1,500 entries from this dataset have already been published as part of our initial database validation publication.[10,11] The LobsterPy[12] package developed alongside enabled the generation of summaries for the most important bonds in materials and provided tools to transform the raw bonding data from the database into machine learning-ready descriptors. The curated descriptors span different types, including statistical representations of bonding characteristics for traditional ML algorithms (e.g., random forests), textual descriptions for large language models (LLMs), and structure graphs for graph neural networks (GNNs). Here, we present the results from employing the statistical bonding descriptors in machine learning to predict the mechanical, vibrational, and thermal properties of crystalline materials. Through this work, we demonstrate that incorporating quantum chemical bonding-based descriptors alongside traditional composition and structure-based ones enhances the model performance. Using SISSO,[13] a symbolic regression method, we also demonstrate that one can discover simple, intuitive relationships between bonding and material properties on a larger scale, which was previously not possible.
Materials at both ends of the thermal conductivity spectrum are desirable for various technological applications. Despite substantial progress in modeling thermal transport within materials, identifying materials with the desired thermal conductivity remains a considerable challenge. This difficulty is partly attributable to the computationally intensive nature of such calculations and to the complexities of modeling many body interactions in solids.[1,2] Recognizing that chemical bonding within a material plays a crucial role in phonon dynamics, several studies have incorporated bonding analysis to investigate the origins of low lattice thermal conductivity. These studies have identified bonding-related features, including bonding heterogeneity, as among the important factors that induce low lattice thermal conductivity.[3–6] In this work, the investigation aims to determine whether we can find such an observation on a larger scale using machine learning techniques. To achieve this, a database of bonding analysis data obtained using the LOBSTER[7–10] program was first generated for approximately 13,000 materials[11,12] sourced from the Materials Project.[13] This data was subsequently transformed into machine-learning-ready descriptors that can numerically quantify the material's bonding heterogeneity. These descriptors were evaluated within machine learning algorithms (e.g., random forests) to assess how their inclusion, alongside traditional structure and composition-based descriptors, influences model predictive performance. The primary target property in these models is the total lattice thermal conductivity, including three-phonon interactions.[14] ML models, on average, showed a significant reduction in prediction errors, and feature importance analyses indicated that bonding heterogeneity descriptors exert a considerable influence. Finally, using SISSO,[15,16] a symbolic regression technique, a new descriptor was identified, revealing that increased bonding heterogeneity in a material correlates with a decrease in total lattice thermal conductivity.