TY - JOUR U1 - Wissenschaftlicher Artikel A1 - Budnikov, Mikhail A1 - Bykova, Anna A1 - Yamshchikov, Ivan T1 - Generalization potential of large language models JF - Neural Computing and Applications N2 - The rise of deep learning techniques and especially the advent of large language models (LLMs) intensified the discussions around possibilities that artificial intelligence with higher generalization capability entails. The range of opinions on the capabilities of LLMs is extremely broad: from equating language models with stochastic parrots to stating that they are already conscious. This paper represents an attempt to review LLM landscape in the context of their generalization capacity as an information theoretic property of those complex systems. We discuss the suggested theoretical explanations for generalization in LLMs and highlight possible mechanisms responsible for these generalization properties. Through an examination of existing literature and theoretical frameworks, we endeavor to provide insights into the mechanisms driving the generalization capacity of LLMs, thus contributing to a deeper understanding of their capabilities and limitations in natural language processing tasks. AB - The rise of deep learning techniques and especially the advent of large language models (LLMs) intensified the discussions around possibilities that artificial intelligence with higher generalization capability entails. The range of opinions on the capabilities of LLMs is extremely broad: from equating language models with stochastic parrots to stating that they are already conscious. This paper represents an attempt to review LLM landscape in the context of their generalization capacity as an information theoretic property of those complex systems. We discuss the suggested theoretical explanations for generalization in LLMs and highlight possible mechanisms responsible for these generalization properties. Through an examination of existing literature and theoretical frameworks, we endeavor to provide insights into the mechanisms driving the generalization capacity of LLMs, thus contributing to a deeper understanding of their capabilities and limitations in natural language processing tasks. Y1 - 2024 SN - 0941-0643 SS - 0941-0643 UN - https://nbn-resolving.org/urn:nbn:de:bvb:863-opus-57807 U6 - https://doi.org/10.1007/s00521-024-10827-6 DO - https://doi.org/10.1007/s00521-024-10827-6 VL - 37 IS - 4 SP - 1973 EP - 1997 S1 - 25 PB - Springer ER -