- Aims Systematic visualization of predefined anatomical sites is essential for quality assurance and training in esophagogastroduodenoscopy (EGD), but objective, automated assessment of examination completeness in Europe is lacking. We aimed to develop and preliminarily evaluate a deep learning algorithm that classifies still images from the upper gastrointestinal tract (UGIT) into 28 standardizedAims Systematic visualization of predefined anatomical sites is essential for quality assurance and training in esophagogastroduodenoscopy (EGD), but objective, automated assessment of examination completeness in Europe is lacking. We aimed to develop and preliminarily evaluate a deep learning algorithm that classifies still images from the upper gastrointestinal tract (UGIT) into 28 standardized anatomical categories, including 22 gastric segments according to the Systemic Stomach Screening (SSS) protocol by Kenshi Yao. Primary outcomes were macro F1-score and overall accuracy for this 28-class classification task.
Methods In this retrospective diagnostic machine learning study, we collected endoscopic still images from two sources: (1) the publicly available GastroHUN dataset (Colombia; Bravo et al., Sci Data 2025) and (2) an internal dataset from a European tertiary referral centre (Augsburg, Germany). All images were obtained during routine EGD using Olympus endoscopes and anonymized prior to analysis. A total of 1,553 patients were included: 385 from GastroHUN (5,202 images) and 1,114 from UKA (2,793 images), yielding 8,095 images. Each image was labelled at image level into one of 28 anatomical classes: oropharynx, esophagus, esophagogastric junction, duodenal bulb, second part of the duodenum, papilla and 22 gastric segments based on the SSS classification using a multi-annotator process. We trained a ConvNeXt-Tiny convolutional neural network for image-level classification using patient-wise 5-fold cross-validation to prevent information leakage between training and test sets. Performance was evaluated as mean macro F1-score and overall accuracy across folds.
Results Across 5-fold patient-wise cross-validation, the model achieved a mean macro F1-score of 0.79±0.01 and an overall accuracy of 0.79±0.01 for 28-class anatomical site classification. The network differentiated major UGIT regions (esophagus, stomach, duodenum) and also assigned images to fine-grained gastric segments according to the SSS protocol. Qualitatively, performance was higher in broader anatomical regions and lower in some of the most detailed gastric segments (e.g., lower vs. upper corpus posterior wall), consistent with the increased difficulty of very fine-grained localization.
Conclusions This preliminary study demonstrates the feasibility of automated recognition of UGIT anatomical sites using a ConvNeXt-Tiny CNN trained on a combination of public and tertiary hospital data. The model achieved moderate performance across 28 predefined anatomical categories from the oropharynx to the second part of the duodenum, including 22 SSS gastric segments, indicating that fine-grained anatomical localization of the entire upper GI tract is achievable from still images alone. To our knowledge, this is among the first studies to leverage a public endoscopy dataset together with institutional data for upper GI anatomical site classification.…

