Research Center for Artificial Intelligence - RCAI
Refine
Year of publication
Document Type
- conference proceeding (article) (343)
- Article (295)
- Part of a Book (101)
- conference proceeding (presentation, abstract) (94)
- Preprint (62)
- conference talk (59)
- Report (20)
- Research Data (16)
- Working Paper (15)
- Other (9)
Language
- English (726)
- German (307)
- Multiple languages (2)
- Polish (1)
Is part of the Bibliography
- no (1036)
Keywords
- Künstliche Intelligenz (29)
- Artificial Intelligence (18)
- Deep Learning (18)
- Simulation (16)
- Digitalisierung (13)
- Akzeptanz (11)
- Diagnose (9)
- Photoacoustic spectroscopy (9)
- Chief Information Officer (8)
- Machine learning (8)
Institute
- Research Center for Artificial Intelligence - RCAI (1036)
- Fakultät Informatik und Mathematik (547)
- Research Center of Health Sciences and Technology - RCHST (450)
- Research Center of Biomedical Engineering - RCBE (233)
- Fakultät Elektro- und Informationstechnik (174)
- Fakultät Maschinenbau (170)
- Institut für Sozialforschung und Technikfolgenabschätzung (IST) (145)
- Fakultät Sozial- und Gesundheitswissenschaften (136)
- Laboratory for Safe and Secure Systems (LAS3) (121)
- Labor für Technikfolgenabschätzung und Angewandte Ethik (LaTe) (109)
Begutachtungsstatus
- peer-reviewed (485)
- begutachtet (40)
Process mining often yields highly complex “Spaghetti Models”, making them difficult to interpret and impeding informed decision-making. Therefore, researchers have explored clustering of event logs to simplify process models and reduce their complexity. However, the unsupervised nature of clustering can introduce an interpretation gap, necessitating manual effort to identify differences and similarities across the resulting process models. To address these issues, we propose an explainable clustering approach that identifies key subprocesses and applies eXplainable Artificial Intelligence (XAI) techniques to clarify the rationale behind model partitioning. Moreover, we integrate a Large Language Model (LLM) into the process discovery procedure to generate natural language descriptions and compare the discovered process models, enhancing user understanding, engagement, and making complex technical details more accessible. A case study demonstrates that our method operates effectively across various LLMs, preserving vital contextual information while simplifying the process discovery workflow. Our findings reveal that larger models generally ensure completeness, whereas smaller ones offer more efficiency at the expense of explanation quality, highlighting the importance of a balanced LLM choice for practical applications.
Explanatory Interactive Machine Learning queries user feedback regarding the prediction and the explanation of novel instances. CAIPI, a state-of-the-art algorithm, captures the user feedback and iteratively biases a data set toward a correct decision-making mechanism using counterexamples. The counterexample generation procedure relies on hand-crafted data augmentation and might produce implausible instances. We propose Bayesian CAIPI that embeds a Variational Autoencoder into CAIPI’s classification cycle and samples counterexamples from the likelihood distribution. Using the MNIST data set, where we distinguish ones from sevens, we show that Bayesian CAIPI matches the predictive accuracy of both, traditional CAIPI and default deep learning. Moreover, it outperforms both in terms of explanation quality.
Smart sensor systems are a key factor to ensure sustainable compute by enabling machine learning algorithms to be executed at the data source. This is particularly helpful when working with moving parts or in remote areas, where no tethered deployment is possible. However, including computations directly at the measurement device places an increased load on the power budget. Therefore, we introduce the Hierarchical Machine Learning framework “HiMLEdge” which enables highly specialized models that are tuned using an energy-aware multi-criteria optimization. We evaluate our framework with prognostic health management in a three-part feasibility study: First, we apply an exhaustive search to find hierarchical taxonomies, which we benchmark against hand-tuned flat classifiers. This test shows a decrease in power consumption of up to 47.63% for the hierarchical approach. Second, the search strategy is improved with Reinforcement Learning. As a novel contribution, we include real measurements in the reward function, instead of using a surrogate metric. This inclusion leads to a different optimal policy in comparison to the literature, which shows the error that may be introduced by an approximation. Third, we conduct tests on the system level, including communication and system-off power draw. In this scenario, the optimized hierarchical model can perform four times as many readings per hour as a flat classifier while achieving the same five years of battery life with similar accuracy. In turn, this also means that the battery life can be increased by the same amount if the readings per hour are kept constant.
The rise of machine-learning applications in domains with critical end-user impact has led to a growing concern about the fairness of learned models, with the goal of avoiding biases that negatively impact specific demographic groups. Most existing bias-mitigation strategies adapt the importance of data instances during pre-processing. Since fairness is a contextual concept, we advocate for an interactive machine-learning approach that enables users to provide iterative feedback for model adaptation. Specifically, we propose to adapt the explanatory interactive machine-learning approach Caipi for fair machine learning. FairCaipi incorporates human feedback in the loop on predictions and explanations to improve the fairness of the model. Experimental results demonstrate that FairCaipi outperforms a state-of-the-art pre-processing bias mitigation strategy in terms of the fairness and the predictive performance of the resulting machine-learning model. We show that FairCaipi can both uncover and reduce bias in machine-learning models and allows us to detect human bias.
Pflegekultur Im Wandel: Digitale Technologien Machtstrukturen und Migration Im Gesundheitssektor
(2025)
Die fortschreitende Digitalisierung bringt die komplexen Wechselwirkungen zwischen Pflege, Kultur und Gesellschaft durcheinander. Die Beiträger*innen zeigen, dass in der Pflege ein Kulturwandel in vollem Gange ist, der die Entwicklung einer kultursensiblen Pflege erfordert. Neben den Auswirkungen der neuen Technologien beleuchten sie auch die Themen Migration und darauf aufbauende Machtstrukturen und bieten so eine umfassende, interdisziplinäre Einführung in die modernen Herausforderungen der Pflege, die nicht nur für Wissenschaftler*innen interessant ist, sondern auch praxisbezogene Fragen aufgreift.
BACKGROUND
Ischemic stroke (IS) and retinal ischemia (IR) share similar vascular risk factors, but differ in their risk for subsequent or recurrent stroke and therapeutic options. This study characterizes the cardiovascular risk profiles and magnitude of atherosclerosis of the carotid artery of patients with central retinal artery occlusion (CRAO) in relation to the presence of the retrobulbar "spot sign" on orbital color-coded sonography (OCCS).
METHODS
We performed a retrospective analysis on the detailed cardiovascular risk factors and neuroimaging data in patients with IR presenting between 2009 and 2023. Based on OCCS findings, CRAO were further divided into hyperechoic ("spot sign positive", ssCRAO) or hypoechoic CRAO (heCRAO). Statistical analyses were performed with Mann-Whitney-U and χ [2] testing. P-values were considered significant if < 0.05.
RESULTS
Overall, 112 patients were identified (heCRAO: n = 32; ssCRAO: n = 80). ssCRAO patients were significantly older (median 74 years vs. 66.5 years, Mann-Whitney-U: p-value < 0.001). Overall, 15/103 (14.6%) patients had concurrent acute ischemic stroke- 9 in the ipsilateral internal carotid territory, 2 in other territories and 4 disseminated. Further significant differences were found regarding the echogenicity of atherosclerosis (AS) in the two subgroups with (mainly) echorich AS being more common in the ssCRAO group (p-value < 0.001, n = 108) and the distribution of high-grade vs. low-grade stenoses of the ipsi- and contralateral carotid artery (p-value < 0.05, n = 99). 20 out of 112 patients had atrial fibrillation (aFib) with 17 of these being on ongoing oral anticoagulation.
CONCLUSION
According to this study, atherosclerosis may be one of the most important risk factors for IR while a specific embolic source could not be demonstrated (i.e. acute plaque rupture). By contrast, current oral anticoagulation for aFib in CRAO patients was high, thus only an incidental finding and may be an incidental finding due to its prevalence in the elderly. Furthermore, we were able to distinguish two subgroups of IR that differ in risk factors and most likely also in etiology, therapy and prognosis. The study underlines the importance of OCCS to detect "spot signs" in IR with indications for both, acute thrombolysis and secondary prevention.
Aims Precise intraprocedural phase recognition during complex endoscopic procedures like endoscopic submucosal dissection (ESD) could facilitate automatic objective reporting, specific target-oriented training interventions and measurable quality control. Artificial intelligence algorithms have already been applied to phase recognition in laparoscopic operations and peroral endoscopic myotomy successfully. The aim of this study was to develop and validate an algorithm for automated intraprocedural phase recognition during ESD on a single frame basis.
Methods The ESD procedure was divided into 5 macro phases: diagnostics, marking, needle injection, dissection and bleeding. Macro-phases were subdivided into micro-phases, which included scope manipulation, electric current application and injection, leading to a total of 11 phases. A training dataset was compiled from 92 full-length ESD videos (7.930.412 frames) and each frame allocated to one phase. All procedures were performed with the same endoscope system (GIF EZ1500, Olympus, Tokyo, Japan) A video swin transformer was trained in the recognition of the procedural phases. Temporal information was incorporated by uniform frame sampling from past and future of the analyzed frame. The algorithm was validated internally on a validation set (16 ESD procedures, 1.580.394 frames). External validation was performed on 8 ESD procedures (759.346 frames) from a live pig study using a different endoscopy system (GIF H190, Olympus, Tokyo, Japan). In a further evaluation, 2 videos from the external validation set were incorporated into the training data.
Results The overall internal validation yielded an accuracy of 86% and an F1 score of 86%. The external validation showed values of 69% and 70% for the same parameters. The evaluation for macro phases showed overall accuracies of 89% and 80% and F1 scores of 90% and 80% for internal and external test sets, respectively. Incorporation of 2 videos from the test set into the training set led to the following results: The accuracy and F1-score for all phase validation were 78% and 78%. For macro-phase evaluation, these values were 87% and 87%.
Aims Systematic visualization of predefined anatomical sites is essential for quality assurance and training in esophagogastroduodenoscopy (EGD), but objective, automated assessment of examination completeness in Europe is lacking. We aimed to develop and preliminarily evaluate a deep learning algorithm that classifies still images from the upper gastrointestinal tract (UGIT) into 28 standardized anatomical categories, including 22 gastric segments according to the Systemic Stomach Screening (SSS) protocol by Kenshi Yao. Primary outcomes were macro F1-score and overall accuracy for this 28-class classification task.
Methods In this retrospective diagnostic machine learning study, we collected endoscopic still images from two sources: (1) the publicly available GastroHUN dataset (Colombia; Bravo et al., Sci Data 2025) and (2) an internal dataset from a European tertiary referral centre (Augsburg, Germany). All images were obtained during routine EGD using Olympus endoscopes and anonymized prior to analysis. A total of 1,553 patients were included: 385 from GastroHUN (5,202 images) and 1,114 from UKA (2,793 images), yielding 8,095 images. Each image was labelled at image level into one of 28 anatomical classes: oropharynx, esophagus, esophagogastric junction, duodenal bulb, second part of the duodenum, papilla and 22 gastric segments based on the SSS classification using a multi-annotator process. We trained a ConvNeXt-Tiny convolutional neural network for image-level classification using patient-wise 5-fold cross-validation to prevent information leakage between training and test sets. Performance was evaluated as mean macro F1-score and overall accuracy across folds.
Results Across 5-fold patient-wise cross-validation, the model achieved a mean macro F1-score of 0.79±0.01 and an overall accuracy of 0.79±0.01 for 28-class anatomical site classification. The network differentiated major UGIT regions (esophagus, stomach, duodenum) and also assigned images to fine-grained gastric segments according to the SSS protocol. Qualitatively, performance was higher in broader anatomical regions and lower in some of the most detailed gastric segments (e.g., lower vs. upper corpus posterior wall), consistent with the increased difficulty of very fine-grained localization.
Conclusions This preliminary study demonstrates the feasibility of automated recognition of UGIT anatomical sites using a ConvNeXt-Tiny CNN trained on a combination of public and tertiary hospital data. The model achieved moderate performance across 28 predefined anatomical categories from the oropharynx to the second part of the duodenum, including 22 SSS gastric segments, indicating that fine-grained anatomical localization of the entire upper GI tract is achievable from still images alone. To our knowledge, this is among the first studies to leverage a public endoscopy dataset together with institutional data for upper GI anatomical site classification.
Aims Endoscopic detection of gastric cancer is operator-dependent, and almost all existing artificial intelligence (AI) systems have been developed in high-incidence East Asian settings. We aimed to develop and preliminarily evaluate an AI model for detection and segmentation of gastric cancer in a Western population, where this pathology is low-incidence, using endoscopic submucosal dissection (ESD) histopathology as reference standard. The model was fine-tuned from a previously established semi-supervised network for Barrett’s neoplasia (Meinikheim et al., Endoscopy 2024). Transfer learning is appropriate because both esophageal and gastric cancers share key endoscopic characteristics and multimodal imaging features across WLI, NBI, and TXI. The feature representations learned from large-scale Barrett’s datasets therefore provide a strong initialization. Primary outcomes were Dice similarity coefficient for tumor segmentation and image-level detection sensitivity.
Methods In this retrospective single-centre study at a tertiary-care hospital in Germany (University Hospital Augsburg), we included 84 patients with histologically confirmed gastric adenocarcinoma (T1a or higher) treated by ESD. From these patients, 827 endoscopic images (Olympus systems; WLI, NBI, TXI, with or without indigo chromoendoscopy) showing visible gastric cancer were extracted. Tumor extent was delineated on still images by experts informed by the corresponding ESD specimens and full pathology reports. All data were de-identified before analysis.
A convolutional neural network segmentation model was initialized with weights from the Barrett AI system, which had been pre-trained on 55,273 endoscopic images from 557 patients with Barrett’s esophagus and related neoplasia, and then supervisedly fine-tuned for gastric cancer segmentation. Data were split patient-wise into five train/validation folds (80/20 per fold) to avoid information leakage between patients, and a separate model was trained for each fold. For segmentation, performance was assessed using the Dice coefficient. For image-level detection, a tumor was considered detected in a given image if the entire predicted lesion region overlapped the ground-truth mask with Dice≥75%; detection performance was summarized as sensitivity at this threshold and as the area under a recall–overlap curve obtained by varying the Dice threshold from 0 to 1.
Results Across the five patient-wise folds, tumor Dice averaged 82.93% with a standard deviation of 1.39% [80.91–84.37%]. Image-level detection sensitivity at Dice≥75% averaged 94.86% with a standard deviation of 3.12% [90.23–97.59%]. The area under the recall–overlap curve averaged 0.912 with a standard deviation of 0.009. These results indicate a consistent and robust performance across all validation images.