Refine
Document Type
Is part of the Bibliography
- no (23)
Keywords
- Artificial Intelligence (7)
- Künstliche Intelligenz (7)
- Barrett-Ösophagus (3)
- Diagnose (2)
- Endoscopy (2)
- Medical Image Computing (2)
- Speiseröhrenkrankheit (2)
- Third-Space Endoscopy (2)
- Adenokarzinom (1)
- Barrett's Esophagus (1)
Institute
Begutachtungsstatus
- peer-reviewed (23)
The early diagnosis of cancer in Barrett’s esophagus is crucial for improving the prognosis. However, identifying Barrett’s esophagus-related neoplasia (BERN) is challenging, even for experts [1]. Four-quadrant biopsies may improve the detection of neoplasia, but they can be associated with sampling errors. The application of artificial intelligence (AI) to the assessment of Barrett’s esophagus could improve the diagnosis of BERN, and this has been demonstrated in both preclinical and clinical studies [2] [3].
In this video demonstration, we show the accurate detection and delineation of BERN in two patients ([Video 1]). In part 1, the AI system detects a mucosal cancer about 20 mm in size and accurately delineates the lesion in both white-light and narrow-band imaging. In part 2, a small island of BERN with high-grade dysplasia is detected and delineated in white-light, narrow-band, and texture and color enhancement imaging. The video shows the results using a transparent overlay of the mucosal cancer in real time as well as a full segmentation preview. Additionally, the optical flow allows for the assessment of endoscope movement, something which is inversely related to the reliability of the AI prediction. We demonstrate that multimodal imaging can be applied to the AI-assisted detection and segmentation of even small focal lesions in real time.
Background
This study evaluated the effect of an artificial intelligence (AI)-based clinical decision support system on the performance and diagnostic confidence of endoscopists in their assessment of Barrett’s esophagus (BE).
Methods
96 standardized endoscopy videos were assessed by 22 endoscopists with varying degrees of BE experience from 12 centers. Assessment was randomized into two video sets: group A (review first without AI and second with AI) and group B (review first with AI and second without AI). Endoscopists were required to evaluate each video for the presence of Barrett’s esophagus-related neoplasia (BERN) and then decide on a spot for a targeted biopsy. After the second assessment, they were allowed to change their clinical decision and confidence level.
Results
AI had a stand-alone sensitivity, specificity, and accuracy of 92.2%, 68.9%, and 81.3%, respectively. Without AI, BE experts had an overall sensitivity, specificity, and accuracy of 83.3%, 58.1%, and 71.5%, respectively. With AI, BE nonexperts showed a significant improvement in sensitivity and specificity when videos were assessed a second time with AI (sensitivity 69.8% [95%CI 65.2%–74.2%] to 78.0% [95%CI 74.0%–82.0%]; specificity 67.3% [95%CI 62.5%–72.2%] to 72.7% [95%CI 68.2%–77.3%]). In addition, the diagnostic confidence of BE nonexperts improved significantly with AI.
Conclusion
BE nonexperts benefitted significantly from additional AI. BE experts and nonexperts remained significantly below the stand-alone performance of AI, suggesting that there may be other factors influencing endoscopists’ decisions to follow or discard AI advice.
Aims
Human-computer interactions (HCI) may have a relevant impact on the performance of Artificial Intelligence (AI). Studies show that although endoscopists assessing Barrett’s esophagus (BE) with AI improve their performance significantly, they do not achieve the level of the stand-alone performance of AI. One aspect of HCI is the impact of AI on the degree of certainty and confidence displayed by the endoscopist. Indirectly, diagnostic confidence when using AI may be linked to trust and acceptance of AI. In a BE video study, we aimed to understand the impact of AI on the diagnostic confidence of endoscopists and the possible correlation with diagnostic performance.
Methods
22 endoscopists from 12 centers with varying levels of BE experience reviewed ninety-six standardized endoscopy videos. Endoscopists were categorized into experts and non-experts and randomly assigned to assess the videos with and without AI. Participants were randomized in two arms: Arm A assessed videos first without AI and then with AI, while Arm B assessed videos in the opposite order. Evaluators were tasked with identifying BE-related neoplasia and rating their confidence with and without AI on a scale from 0 to 9.
Results
The utilization of AI in Arm A (without AI first, with AI second) significantly elevated confidence levels for experts and non-experts (7.1 to 8.0 and 6.1 to 6.6, respectively). Only non-experts benefitted from AI with a significant increase in accuracy (68.6% to 75.5%). Interestingly, while the confidence levels of experts without AI were higher than those of non-experts with AI, there was no significant difference in accuracy between these two groups (71.3% vs. 75.5%). In Arm B (with AI first, without AI second), experts and non-experts experienced a significant reduction in confidence (7.6 to 7.1 and 6.4 to 6.2, respectively), while maintaining consistent accuracy levels (71.8% to 71.8% and 67.5% to 67.1%, respectively).
Conclusions
AI significantly enhanced confidence levels for both expert and non-expert endoscopists. Endoscopists felt significantly more uncertain in their assessments without AI. Furthermore, experts with or without AI consistently displayed higher confidence levels than non-experts with AI, irrespective of comparable outcomes. These findings underscore the possible role of AI in improving diagnostic confidence during endoscopic assessment.
Ziele:
Das Ziel der Studie war es, den Einfluss von KI auf die diagnostische Sicherheit (Konfidenzniveau) von Endoskopikern anhand von BÖ-Videos zu untersuchen und mögliche Korrelationen mit der Untersuchungsqualität zu erforschen.
Methodik:
22 Endoskopiker aus zwölf Zentren mit unterschiedlicher Barrett-Erfahrung untersuchten 96 standardisierte Endoskopievideos. Die Untersucher wurden in Experten und Nicht-Experten eingeteilt und nach dem Zufallsprinzip für die Bewertung der Videos mit oder ohne KI eingeteilt. Die Teilnehmer wurden in zwei Gruppen aufgeteilt: Arm A bewertete zunächst Videos ohne KI und dann mit KI, während Arm B die umgekehrte Reihenfolge einhielt. Die Untersucher hatten die Aufgabe, BÖ-assoziierte Neoplasien zu erkennen und ihr Konfidenzniveau sowohl mit als auch ohne KI auf einer Skala von 0 bis 9 anzugeben.
Ergebnis:
In Arm A erhöhte der Einsatz von KI das Konfidenzniveau bei beiden signifikant (p<0.001). Bemerkenswert ist, dass jedoch nur Nicht-Experten durch die KI eine signifikante Verbesserung der Sensitivität und Spezifität (p<0.001 bzw. p<0.05) erfuhren. Während Experten ohne KI im Vergleich zu Nicht-Experten mit KI ein höheres Konfidenzniveau aufwiesen, gab es keinen signifikanten Unterschied in der Genauigkeit. In Arm B zeigten beide Gruppen eine signifikante Abnahme des Konfidenzniveaus (p<0.001) bei gleichbleibender Genauigkeit. Darüber hinaus wurden in 9% der Entscheidungen trotz korrekter KI eine falsche Wahl getroffen.
Schlussfolgerung:
Der Einsatz künstlicher Intelligenz steigerte das Konfidenzniveau sowohl bei Experten als auch bei Nicht-Experten signifikant – ein Effekt, der im Studienmodell reversibel war. Darüber hinaus wiesen Experten mit oder ohne KI durchweg höhere Konfidenzniveaus auf als Nicht-Experten mit KI, trotz vergleichbarer Ergebnisse. Zudem konnte beobachtet werden, dass die Untersucher in 9% der Fälle die KI zuungunsten des Patienten ignorierten.
Objective
Despite high stand-alone performance, studies demonstrate that artificial intelligence (AI)-supported endoscopic diagnostics often fall short in clinical applications due to human-AI interaction factors. This video-based trial on Barrett's esophagus aimed to investigate how examiner behavior, their levels of confidence, and system usability influence the diagnostic outcomes of AI-assisted endoscopy.
Methods
The present analysis employed data from a multicenter randomized controlled tandem video trial involving 22 endoscopists with varying degrees of expertise. Participants were tasked with evaluating a set of 96 endoscopic videos of Barrett's esophagus in two distinct rounds, with and without AI assistance. Diagnostic confidence levels were recorded, and decision changes were categorized according to the AI prediction. Additional surveys assessed user experience and system usability ratings.
Results
AI assistance significantly increased examiner confidence levels (p < 0.001) and accuracy. Withdrawing AI assistance decreased confidence (p < 0.001), but not accuracy. Experts consistently reported higher confidence than non-experts (p < 0.001), regardless of performance. Despite improved confidence, correct AI guidance was disregarded in 16% of all cases, and 9% of initially correct diagnoses were changed to incorrect ones. Overreliance on AI, algorithm aversion, and uncertainty in AI predictions were identified as key factors influencing outcomes. The System Usability Scale questionnaire scores indicated good to excellent usability, with non-experts scoring 73.5 and experts 85.6.
Conclusions
Our findings highlight the pivotal function of examiner behavior in AI-assisted endoscopy. To fully realize the benefits of AI, implementing explainable AI, improving user interfaces, and providing targeted training are essential. Addressing these factors could enhance diagnostic accuracy and confidence in clinical practice.
Einleitung
Third-Space Interventionen wie die endoskopische Submukosadissektion (ESD) und die perorale endoskopische Myotomie (POEM) sind technisch anspruchsvoll und mit einem erhöhten Risiko für intraprozedurale Komplikationen wie Blutung oder Perforation assoziiert. Moderne Computerprogramme zur Unterstützung bei diagnostischen Entscheidungen werden unter Einsatz von künstlicher Intelligenz (KI) in der Endoskopie bereits erfolgreich eingesetzt. Ziel der vorliegenden Arbeit war es, relevante anatomische Strukturen mithilfe eines Deep-Learning Algorithmus zu detektieren und segmentieren, um die Sicherheit und Anwendbarkeit von ESD und POEM zu erhöhen.
Methoden
Zwölf Videoaufnahmen in voller Länge von Third-Space Endoskopien wurden aus der Datenbank des Universitätsklinikums Augsburg extrahiert. 1686 Einzelbilder wurden für die Kategorien Submukosa, Blutgefäß, Dissektionsmesser und endoskopisches Instrument annotiert und segmentiert. Mit diesem Datensatz wurde ein DeepLabv3+neuronales Netzwerk auf der Basis eines ResNet mit 101 Schichten trainiert und intern anhand der Parameter Intersection over Union (IoU), Dice Score und Pixel Accuracy validiert. Die Fähigkeit des Algorithmus zur Gefäßdetektion wurde anhand von 24 Videoclips mit einer Spieldauer von 7 bis 46 Sekunden mit 33 vordefinierten Gefäßen evaluiert. Anhand dieses Tests wurde auch die Gefäßdetektionsrate eines Experten in der Third-Space Endoskopie ermittelt.
Ergebnisse
Der Algorithmus zeigte eine Gefäßdetektionsrate von 93,94% mit einer mittleren Rate an falsch positiven Signalen von 1,87 pro Minute. Die Gefäßdetektionsrate des Experten lag bei 90,1% ohne falsch positive Ergebnisse. In der internen Validierung an Einzelbildern wurde eine IoU von 63,47%, ein mittlerer Dice Score von 76,18% und eine Pixel Accuracy von 86,61% ermittelt.
Zusammenfassung
Dies ist der erste KI-Algorithmus, der für den Einsatz in der therapeutischen Endoskopie entwickelt wurde. Präliminäre Ergebnisse deuten auf eine mit Experten vergleichbare Detektion von Gefäßen während der Untersuchung hin. Weitere Untersuchungen sind nötig, um die Leistung des Algorithmus im Vergleich zum Experten genauer zu eruieren sowie einen möglichen klinischen Nutzen zu ermitteln.
Einleitung
Die sichere Detektion und Charakterisierung von Barrett-Ösophagus assoziierten Neoplasien (BERN) stellt selbst für erfahrene Endoskopiker eine Herausforderung dar.
Ziel
Ziel dieser Studie ist es, den Add-on Effekt eines künstlichen Intelligenz (KI) Systems (Barrett-Ampel) als Entscheidungsunterstüzungssystem für Endoskopiker ohne Expertise bei der Untersuchung von BERN zu evaluieren.
Material und Methodik
Zwölf Videos in „Weißlicht“ (WL), „narrow-band imaging“ (NBI) und „texture and color enhanced imaging“ (TXI) von histologisch bestätigten Barrett-Metaplasien oder BERN wurden von Experten und Untersuchern ohne Barrett-Expertise evaluiert. Die Probanden wurden dazu aufgefordert in den Videos auftauchende BERN zu identifizieren und gegebenenfalls die optimale Biopsiestelle zu markieren. Unser KI-System wurde demselben Test unterzogen, wobei dieses BERN in Echtzeit segmentierte und farblich von umliegendem Epithel differenzierte. Anschließend wurden den Probanden die Videos mit zusätzlicher KI-Unterstützung gezeigt. Basierend auf dieser neuen Information, wurden die Probanden zu einer Reevaluation ihrer initialen Beurteilung aufgefordert.
Ergebnisse
Die „Barrett-Ampel“ identifizierte unabhängig von den verwendeten Darstellungsmodi (WL, NBI, TXI) alle BERN. Zwei entzündlich veränderte Läsionen wurden fehlinterpretiert (Genauigkeit=75%). Während Experten vergleichbare Ergebnisse erzielten (Genauigkeit=70,8%), hatten Endoskopiker ohne Expertise bei der Beurteilung von Barrett-Metaplasien eine Genauigkeit von lediglich 58,3%. Wurden die nicht-Experten allerdings von unserem KI-System unterstützt, erreichten diese eine Genauigkeit von 75%.
Zusammenfassung
Unser KI-System hat das Potential als Entscheidungsunterstützungssystem bei der Differenzierung zwischen Barrett-Metaplasie und BERN zu fungieren und so Endoskopiker ohne entsprechende Expertise zu assistieren. Eine Limitation dieser Studie ist die niedrige Anzahl an eingeschlossenen Videos. Um die Ergebnisse dieser Studie zu bestätigen, müssen randomisierte kontrollierte klinische Studien durchgeführt werden.
Background and aims
Celiac disease with its endoscopic manifestation of villous atrophy is underdiagnosed worldwide. The application of artificial intelligence (AI) for the macroscopic detection of villous atrophy at routine esophagogastroduodenoscopy may improve diagnostic performance.
Methods
A dataset of 858 endoscopic images of 182 patients with villous atrophy and 846 images from 323 patients with normal duodenal mucosa was collected and used to train a ResNet 18 deep learning model to detect villous atrophy. An external data set was used to test the algorithm, in addition to six fellows and four board certified gastroenterologists. Fellows could consult the AI algorithm’s result during the test. From their consultation distribution, a stratification of test images into “easy” and “difficult” was performed and used for classified performance measurement.
Results
External validation of the AI algorithm yielded values of 90 %, 76 %, and 84 % for sensitivity, specificity, and accuracy, respectively. Fellows scored values of 63 %, 72 % and 67 %, while the corresponding values in experts were 72 %, 69 % and 71 %, respectively. AI consultation significantly improved all trainee performance statistics. While fellows and experts showed significantly lower performance for “difficult” images, the performance of the AI algorithm was stable.
Conclusion
In this study, an AI algorithm outperformed endoscopy fellows and experts in the detection of villous atrophy on endoscopic still images. AI decision support significantly improved the performance of non-expert endoscopists. The stable performance on “difficult” images suggests a further positive add-on effect in challenging cases.
Aims
VA is an endoscopic finding of celiac disease (CD), which can easily be missed if pretest probability is low. In this study, we aimed to develop an artificial intelligence (AI) algorithm for the detection of villous atrophy on endoscopic images.
Methods
858 images from 182 patients with VA and 846 images from 323 patients with normal duodenal mucosa were used for training and internal validation of an AI algorithm (ResNet18). A separate dataset was used for external validation, as well as determination of detection performance of experts, trainees and trainees with AI support. According to the AI consultation distribution, images were stratified into “easy” and “difficult”.
Results
Internal validation showed 82%, 85% and 84% for sensitivity, specificity and accuracy. External validation showed 90%, 76% and 84%. The algorithm was significantly more sensitive and accurate than trainees, trainees with AI support and experts in endoscopy. AI support in trainees was associated with significantly improved performance. While all endoscopists showed significantly lower detection for “difficult” images, AI performance remained stable.
Conclusions
The algorithm outperformed trainees and experts in sensitivity and accuracy for VA detection. The significant improvement with AI support suggests a potential clinical benefit. Stable performance of the algorithm in “easy” and “difficult” test images may indicate an advantage in macroscopically challenging cases.
Aims
Evaluation of the add-on effect an artificial intelligence (AI) based clinical decision support system has on the performance of endoscopists with different degrees of expertise in the field of Barrett's esophagus (BE) and Barrett's esophagus-related neoplasia (BERN).
Methods
The support system is based on a multi-task deep learning model trained to solve a segmentation and several classification tasks. The training approach represents an extension of the ECMT semi-supervised learning algorithm. The complete system evaluates a decision tree between estimated motion, classification, segmentation, and temporal constraints, to decide when and how the prediction is highlighted to the observer. In our current study, ninety-six video cases of patients with BE and BERN were prospectively collected and assessed by Barrett's specialists and non-specialists. All video cases were evaluated twice – with and without AI assistance. The order of appearance, either with or without AI support, was assigned randomly. Participants were asked to detect and characterize regions of dysplasia or early neoplasia within the video sequences.
Results
Standalone sensitivity, specificity, and accuracy of the AI system were 92.16%, 68.89%, and 81.25%, respectively. Mean sensitivity, specificity, and accuracy of expert endoscopists without AI support were 83,33%, 58,20%, and 71,48 %, respectively. Gastroenterologists without Barrett's expertise but with AI support had a comparable performance with a mean sensitivity, specificity, and accuracy of 76,63%, 65,35%, and 71,36%, respectively.
Conclusions
Non-Barrett's experts with AI support had a similar performance as experts in a video-based study.