Refine
Document Type
Is part of the Bibliography
- no (22)
Keywords
- Künstliche Intelligenz (4)
- Artificial Intelligence (3)
- Barrett-Ösophagus (1)
- Endoscopy (1)
- Forschung (1)
- Forschungsbericht (1)
- Medical Image Computing (1)
- Migration (1)
- Multikulturelle Gesellschaft (1)
- Smart Endoscopy (1)
Institute
- Research Center of Biomedical Engineering - RCBE (21)
- Research Center of Health Sciences and Technology - RCHST (21)
- Fakultät Informatik und Mathematik (20)
- Labor Regensburg Medical Image Computing (ReMIC) (20)
- Research Center for Artificial Intelligence - RCAI (19)
- Fakultät Sozial- und Gesundheitswissenschaften (1)
- Hochschulleitung/Hochschulverwaltung (1)
- Research Center of Energy and Resources - RCER (1)
- Zentrum für Forschung und Transfer (ZFT ab 2024; vorher: IAFW) (1)
Begutachtungsstatus
- peer-reviewed (20)
Aims
Evaluation of the add-on effect an artificial intelligence (AI) based clinical decision support system has on the performance of endoscopists with different degrees of expertise in the field of Barrett's esophagus (BE) and Barrett's esophagus-related neoplasia (BERN).
Methods
The support system is based on a multi-task deep learning model trained to solve a segmentation and several classification tasks. The training approach represents an extension of the ECMT semi-supervised learning algorithm. The complete system evaluates a decision tree between estimated motion, classification, segmentation, and temporal constraints, to decide when and how the prediction is highlighted to the observer. In our current study, ninety-six video cases of patients with BE and BERN were prospectively collected and assessed by Barrett's specialists and non-specialists. All video cases were evaluated twice – with and without AI assistance. The order of appearance, either with or without AI support, was assigned randomly. Participants were asked to detect and characterize regions of dysplasia or early neoplasia within the video sequences.
Results
Standalone sensitivity, specificity, and accuracy of the AI system were 92.16%, 68.89%, and 81.25%, respectively. Mean sensitivity, specificity, and accuracy of expert endoscopists without AI support were 83,33%, 58,20%, and 71,48 %, respectively. Gastroenterologists without Barrett's expertise but with AI support had a comparable performance with a mean sensitivity, specificity, and accuracy of 76,63%, 65,35%, and 71,36%, respectively.
Conclusions
Non-Barrett's experts with AI support had a similar performance as experts in a video-based study.
Einleitung
Third space Endoskopieprozeduren wie die endoskopische Submukosadissektion (ESD) und die perorale endoskopische Myotomie (POEM) sind technisch anspruchsvoll und gehen mit untersucherabhängigen Komplikationen wie Blutungen und Perforationen einher. Grund hierfür ist die unabsichtliche Durchschneidung von submukosalen Blutgefäßen ohne präemptive Koagulation.
Ziele
Die Forschungsfrage, ob ein KI-Algorithmus die intraprozedurale Gefäßerkennung bei ESD und POEM unterstützen und damit Komplikationen wie Blutungen verhindern könnte, erscheint in Anbetracht des erfolgreichen Einsatzes von KI bei der Erkennung von Kolonpolypen interessant.
Methoden
Auf 5470 Einzelbildern von 59 third space Endoscopievideos wurden submukosale Blutgefäße annotiert. Zusammen mit weiteren 179.681 nicht-annotierten Bildern wurde ein DeepLabv3+neuronales Netzwerk mit dem ECMT-Verfahren für semi-supervised learning trainiert, um Blutgefäße in Echtzeit erkennen zu können. Für die Evaluation wurde ein Videotest mit 101 Videoclips aus 15 vom Trainingsdatensatz separaten Prozeduren mit 200 vordefinierten Gefäßen erstellt. Die Gefäßdetektionsrate, -zeit und -dauer, definiert als der Prozentsatz an Einzelbildern eines Videos bezogen auf den Goldstandard, auf denen ein definiertes Gefäß erkannt wurde, wurden erhoben. Acht erfahrene Endoskopiker wurden mithilfe dieses Videotests im Hinblick auf Gefäßdetektion getestet, wobei eine Hälfte der Videos nativ, die andere Hälfte nach Markierung durch den KI-Algorithmus angesehen wurde.
Ergebnisse
Der mittlere Dice Score des Algorithmus für Blutgefäße war 68%. Die mittlere Gefäßdetektionsrate im Videotest lag bei 94% (96% für ESD; 74% für POEM). Die mediane Gefäßdetektionszeit des Algorithmus lag bei 0,32 Sekunden (0,3 Sekunden für ESD; 0,62 Sekunden für POEM). Die mittlere Gefäßdetektionsdauer lag bei 59,1% (60,6% für ESD; 44,8% für POEM) des Goldstandards. Alle Endoskopiker hatten mit KI-Unterstützung eine höhere Gefäßdetektionsrate als ohne KI. Die mittlere Gefäßdetektionsrate ohne KI lag bei 56,4%, mit KI bei 71,2% (p<0.001).
Schlussfolgerung
KI-Unterstützung war mit einer statistisch signifikant höheren Gefäßdetektionsrate vergesellschaftet. Die mediane Gefäßdetektionszeit von deutlich unter einer Sekunde sowie eine Gefäßdetektionsdauer von größer 50% des Goldstandards wurden für den klinischen Einsatz als ausreichend erachtet. In prospektiven Anwendungsstudien sollte der KI-Algorithmus auf klinische Relevanz getestet werden.
Einleitung
Die Differenzierung zwischen nicht dysplastischem Barrett-Ösophagus (NDBE) und mit Barrett-Ösophagus assoziierten Neoplasien (BERN) während der endoskopischen Inspektion erfordert viel Expertise. Die frühe Diagnosestellung ist wichtig für die weitere Prognose des Barrett-Karzinoms. In Deutschland werden Patient:innen mit einem Barrett-Ösophagus (BE) in der Regel im niedergelassenen Sektor überwacht.
Ziele
Ziel ist es, den Einfluss von einem auf Künstlicher Intelligenz (KI) basierenden klinischen Entscheidungsunterstützungssystems (CDSS) auf die Performance von niedergelassenen Gastroenterolog:innen (NG) bei der Evaluation von Barrett-Ösophagus (BE) zu untersuchen.
Methodik
Es erfolgte die prospektive Sammlung von 96 unveränderten hochauflösenden Videos mit Fällen von Patient:innen mit histologisch bestätigtem NDBE und BERN. Alle eingeschlossenen Fälle enthielten mindestens zwei der folgenden Darstellungsmethoden: HD-Weißlichtendoskopie, Narrow Band Imaging oder Texture and Color Enhancement Imaging. Sechs NG von sechs unterschiedlichen Praxen wurden als Proband:innen eingeschlossen. Es erfolgte eine permutierte Block-Randomisierung der Videofälle in entweder Gruppe A oder Gruppe B. Gruppe A implizierte eine Evaluation des Falls durch Proband:innen zunächst ohne KI und anschließend mit KI als CDSS. In Gruppe B erfolgte die Evaluation in umgekehrter Reihenfolge. Anschließend erfolgte eine zufällige Wiedergabe der so entstandenen Subgruppen im Rahmen des Tests.
Ergebnis
In diesem Test konnte ein von uns entwickeltes KI-System (Barrett-Ampel) eine Sensitivität von 92,2%, eine Spezifität von 68,9% und eine Accuracy von 81,3% erreichen. Mit der Hilfe von KI verbesserte sich die Sensitivität der NG von 64,1% auf 71,2% (p<0,001) und die Accuracy von 66,3% auf 70,8% (p=0,006) signifikant. Eine signifikante Verbesserung dieser Parameter zeigte sich ebenfalls, wenn die Proband:innen die Fälle zunächst ohne KI evaluierten (Gruppe A). Wurde der Fall jedoch als Erstes mit der Hilfe von KI evaluiert (Gruppe B), blieb die Performance nahezu konstant.
Schlussfolgerung
Es konnte ein performantes KI-System zur Evaluation von BE entwickelt werden. NG verbessern sich bei der Evaluation von BE durch den Einsatz von KI.
Objective
Despite high stand-alone performance, studies demonstrate that artificial intelligence (AI)-supported endoscopic diagnostics often fall short in clinical applications due to human-AI interaction factors. This video-based trial on Barrett's esophagus aimed to investigate how examiner behavior, their levels of confidence, and system usability influence the diagnostic outcomes of AI-assisted endoscopy.
Methods
The present analysis employed data from a multicenter randomized controlled tandem video trial involving 22 endoscopists with varying degrees of expertise. Participants were tasked with evaluating a set of 96 endoscopic videos of Barrett's esophagus in two distinct rounds, with and without AI assistance. Diagnostic confidence levels were recorded, and decision changes were categorized according to the AI prediction. Additional surveys assessed user experience and system usability ratings.
Results
AI assistance significantly increased examiner confidence levels (p < 0.001) and accuracy. Withdrawing AI assistance decreased confidence (p < 0.001), but not accuracy. Experts consistently reported higher confidence than non-experts (p < 0.001), regardless of performance. Despite improved confidence, correct AI guidance was disregarded in 16% of all cases, and 9% of initially correct diagnoses were changed to incorrect ones. Overreliance on AI, algorithm aversion, and uncertainty in AI predictions were identified as key factors influencing outcomes. The System Usability Scale questionnaire scores indicated good to excellent usability, with non-experts scoring 73.5 and experts 85.6.
Conclusions
Our findings highlight the pivotal function of examiner behavior in AI-assisted endoscopy. To fully realize the benefits of AI, implementing explainable AI, improving user interfaces, and providing targeted training are essential. Addressing these factors could enhance diagnostic accuracy and confidence in clinical practice.
BABS-Mi
(2025)
Aims
Artificial Intelligence (AI) systems in gastrointestinal endoscopy are narrow because they are trained to solve only one specific task. Unlike Narrow-AI, general AI systems may be able to solve multiple and unrelated tasks. We aimed to understand whether an AI system trained to detect, characterize, and segment early Barrett’s neoplasia (Barrett’s AI) is only capable of detecting this pathology or can also detect and segment other diseases like early squamous cell cancer (SCC).
Methods
120 white light (WL) and narrow-band endoscopic images (NBI) from 60 patients (1 WL and 1 NBI image per patient) were extracted from the endoscopic database of the University Hospital Augsburg. Images were annotated by three expert endoscopists with extensive experience in the diagnosis and endoscopic resection of early esophageal neoplasias. An AI system based on DeepLabV3+architecture dedicated to early Barrett’s neoplasia was tested on these images. The AI system was neither trained with SCC images nor had it seen the test images prior to evaluation. The overlap between the three expert annotations („expert-agreement“) was the ground truth for evaluating AI performance.
Results
Barrett’s AI detected early SCC with a mean intersection over reference (IoR) of 92% when at least 1 pixel of the AI prediction overlapped with the expert-agreement. When the threshold was increased to 5%, 10%, and 20% overlap with the expert-agreement, the IoR was 88%, 85% and 82%, respectively. The mean Intersection Over Union (IoU) – a metric according to segmentation quality between the AI prediction and the expert-agreement – was 0.45. The mean expert IoU as a measure of agreement between the three experts was 0.60.
Conclusions
In the context of this pilot study, the predictions of SCC by a Barrett’s dedicated AI showed some overlap to the expert-agreement. Therefore, features learned from Barrett’s cancer-related training might be helpful also for SCC prediction. Our results allow different possible explanations. On the one hand, some Barrett’s cancer features generalize toward the related task of assessing early SCC. On the other hand, the Barrett’s AI is less specific to Barrett’s cancer than a general predictor of pathological tissue. However, we expect to enhance the detection quality significantly by extending the training to SCC-specific data. The insight of this study opens the way towards a transfer learning approach for more efficient training of AI to solve tasks in other domains.
Die zuverlässige endoskopische Einschätzung der Infiltrationstiefe bei Neoplasien des Barrett-Ösophagus (BO) ist entscheidend für die Therapieentscheidung zwischen organerhaltender endoskopischer Resektion und chirurgischer Resektion. Ziel dieser Studie ist die Entwicklung eines KIVision- Modells, das endoskopisch kurativ resezierbare Läsionen (LGD; HGD; T1a m1 –T1b sm1, L0, V0 und G1-2; EREAC) von nicht kurativ endoskopisch resezierbaren Läsionen (≥ T1b sm2 und/oder L1, V1, G3; NEREAC) anhand endoskopischer Bildgebung unterscheidet. In einer Pilotstudie mit 116 Patient:innen wurde bei der Unterscheidung von T1a vs. T1b mit 230 Bildern ein F1-Score, eine Sensitivität und Spezifität von 74%, 77% und 64% erreicht. [1]
Material und Methodik Es handelt sich um eine ambidirektionale multizentrische Studie mit insgesamt n=578 Patient:innen und n=651 Läsionen im Zeitraum Januar 2014 bis Juni 2025. Eingeschlossen wurden Patient:innen mit gesicherter BO-assoziierter Neoplasie anhand des histologischen Goldstandards durch a) endoskopische Submukosadissektion (ESD) oder b) chirurgische Resektion. Insgesamt wurden 526 Läsionen im Arm EREAC und 125 Läsionen der Gruppe im Arm NEREAC akquiriert. Es werden Bild- und Videodaten aus Olympus und Fuji-Endoskopen verwendet. Aus der Videodokumentation wurden Frames extrahiert und anhand eines vordefinierten Schemas annotiert (u.a. Lasionssichtbarkeit, virtuelle (NBI, TXI) und farbbasierte Chromoendoskopie, Zoommodus). Es erfolgte eine mehrstufige Annotation durch einen medizinischen Doktoranden (J.B.), mit 1:1 Supervision durch einen endoskopisch tätigen Facharzt im Bereich Gastroenterologie (D.R.), sowie durch einen Barrett-Experten (A.E). Aufgrund einer Klassenimbalance zugunsten der EREAC-Gruppe erfolgt zunachst eine Zwischenanalyse aus 185 Patient:innen mit 215 Läsionen mit a 500 Frames. Mittels ConvNeXt Tiny Architektur wurde eine 5-fache Cross-Validierung mit Split auf Patientenebene durchgeführt. Primare Leistungsmetriken sind F1-Score, Sensitivität und Spezifität.
Ergebnisse Im aktuellen Test erreichte das KI-Modell für die Klassifikation (EREAC vs. NEREAC) einen mittleren F1-Score von 78.3% (resektabel: 79.8%, nicht 2 resektabel: 76.8%). Die mittlere Sensitivität und Spezifität lag bei je 78.3%. Die klassenspezifische Sensitivität und Spezifität zur Erkennung von NEREAC lag bei 72.0% und 84.6%.
Zusammenfassung Ein KI-Assistenzmodell kann die endoskopische Einschatzung der Resektabilität von Barrett-Neoplasien unterstutzen, indem es die Infiltrationstiefe aus endoskopischer Bildgebung pradiziert. Klinisch besonders relevant ist die hohe Spezifität in der Erkennung nicht-resektabler Läsionen.
Nächste Schritte sind das weitere Training mit höherer Datenmenge, die adressierte Aufarbeitung der Klassenimbalance, externe Validierung sowie die Echtzeit-Implementation.
Forschung 2019
(2019)
The endoscopic features associated with eosinophilic esophagitis (EoE) may be missed during routine endoscopy. We aimed to develop and evaluate an Artificial Intelligence (AI) algorithm for detecting and quantifying the endoscopic features of EoE in white light images, supplemented by the EoE Endoscopic Reference Score (EREFS). An AI algorithm (AI-EoE) was constructed and trained to differentiate between EoE and normal esophagus using endoscopic white light images extracted from the database of the University Hospital Augsburg. In addition to binary classification, a second algorithm was trained with specific auxiliary branches for each EREFS feature (AI-EoE-EREFS). The AI algorithms were evaluated on an external data set from the University of North Carolina, Chapel Hill (UNC), and compared with the performance of human endoscopists with varying levels of experience. The overall sensitivity, specificity, and accuracy of AI-EoE were 0.93 for all measures, while the AUC was 0.986. With additional auxiliary branches for the EREFS categories, the AI algorithm (AI-EoEEREFS) performance improved to 0.96, 0.94, 0.95, and 0.992 for sensitivity, specificity, accuracy, and AUC, respectively. AI-EoE and AI-EoE-EREFS performed significantly better than endoscopy beginners and senior fellows on the same set of images. An AI algorithm can be trained to detect and quantify endoscopic features of EoE with excellent performance scores. The addition of the EREFS criteria improved the performance of the AI algorithm, which performed significantly better than endoscopists with a lower or medium experience level.
Background
This study evaluated the effect of an artificial intelligence (AI)-based clinical decision support system on the performance and diagnostic confidence of endoscopists in their assessment of Barrett’s esophagus (BE).
Methods
96 standardized endoscopy videos were assessed by 22 endoscopists with varying degrees of BE experience from 12 centers. Assessment was randomized into two video sets: group A (review first without AI and second with AI) and group B (review first with AI and second without AI). Endoscopists were required to evaluate each video for the presence of Barrett’s esophagus-related neoplasia (BERN) and then decide on a spot for a targeted biopsy. After the second assessment, they were allowed to change their clinical decision and confidence level.
Results
AI had a stand-alone sensitivity, specificity, and accuracy of 92.2%, 68.9%, and 81.3%, respectively. Without AI, BE experts had an overall sensitivity, specificity, and accuracy of 83.3%, 58.1%, and 71.5%, respectively. With AI, BE nonexperts showed a significant improvement in sensitivity and specificity when videos were assessed a second time with AI (sensitivity 69.8% [95%CI 65.2%–74.2%] to 78.0% [95%CI 74.0%–82.0%]; specificity 67.3% [95%CI 62.5%–72.2%] to 72.7% [95%CI 68.2%–77.3%]). In addition, the diagnostic confidence of BE nonexperts improved significantly with AI.
Conclusion
BE nonexperts benefitted significantly from additional AI. BE experts and nonexperts remained significantly below the stand-alone performance of AI, suggesting that there may be other factors influencing endoscopists’ decisions to follow or discard AI advice.