CARISSMA Institute of Automated Driving (C-IAD)
Smartphone applications routinely collect and share personal data, yet users often struggle to understand these practices, particularly when third-party data sharing is involved. Existing mechanisms such as privacy policies and notices provide limited support for user understanding. To address this gap, we investigated how explainability concepts can enhance contextual privacy policies for mobile apps. We designed two interface prototypes integrating explanation strategies: contrastive explanations, which clarify data-sharing boundaries, and example-based explanations, which illustrate counterfactual scenarios. In an exploratory between-subjects user study ( N = 30), we evaluated their impact on comprehensibility, simplicity, cognitive load, and experience. Statistical analysis revealed no significant differences between the two types. Hence, the findings should be read as descriptive trends rather than confirmatory effects. These trends suggested example-based explanations better supported users mental models, while contrastive explanations better supported decision-making. Our findings contribute design recommendations on applying explanation strategies in privacy interfaces, offering guidance for developers and researchers seeking to improve user understanding and trust in digital systems.
Making AI explainable requires more than algorithmic transparency: it demands understanding who needs explanations and why. In our sixth CHI workshop on Human-Centered XAI (HCXAI), we shift focus to agentic AI systems. LLM-based agents foundationally challenge existing explainability paradigms. Unlike traditional AI that produces single outputs, agents plan multi-step strategies, invoke tools with real-world consequences, and coordinate with other systems; yet current XAI approaches fail to address these complexities. Users need to understand not just what an agent might do, but the cascade of actions it could trigger, the risks involved, and why responses take time. Even our expanded HCXAI frameworks struggle with these new demands. Through our workshop series, we have built a strong community making important conceptual, methodological, and technical impact. This year, we re-examine what human-centered explainable AI means in the agentic era, bringing together researchers and practitioners to shape explainability for both users and developers of these systems.
Uncertainty-Aware Diffusion Model for Multimodal Highway Trajectory Prediction via DDIM Sampling
(2026)
Objectives
Driving under the influence of alcohol (DUI) remains a major contributor to fatal traffic crashes worldwide. With increasing regulatory pressure, such as requirements by Euro NCAP for in-vehicle impairment detection, there is a growing need for reliable, real-time monitoring solutions. While traditional DUI detection approaches focus on driving behavior or eye movement analysis, this study explores the potential of thermal imaging as a noninvasive alternative for detecting alcohol impairment.
Methods
We conducted a large-scale experimental study with 120 participants in a high-fidelity driving simulator, capturing thermal facial data under both non-impaired and alcohol-impaired conditions. A novel temperature extraction method was developed based on facial landmarks, incorporating multiple frames to reduce noise and improve temporal stability. Ambient cabin temperature was also recorded to normalize facial temperature readings and control for environmental influences. Several machine learning classifiers, including Logistic Regression, Random Forest, Support Vector Machine (SVM), and Gradient-Boosting Models, were trained using five facial temperature features (cheek, temple, ear, forehead, and nasal tip) and evaluated via five-fold subject-wise cross-validation.
Results
Significant temperature changes were observed in specific facial regions (particularly the cheek, ear, temple, and nasal tip) under alcohol influence. Among the evaluated models, Logistic Regression achieved the highest average classification accuracy (62%), while SVM demonstrated the most stable performance across folds. The model showed a slight conservative bias toward predicting the baseline (non-impaired) class, thereby reducing the risk of false positive classifications. Environmental conditions, including cabin temperature, were verified to be stable across both driving sessions, ensuring the validity of the physiological measurements.
Conclusions
This study demonstrates the feasibility of using thermal imaging for in-vehicle DUI detection under realistic conditions. Our contributions include the development of a robust facial temperature processing pipeline, the creation of a unique dataset collected under ecologically valid conditions, and a comprehensive comparison of seven state-of-the-art classification models. Thermal imaging represents a promising complementary modality for future driver monitoring systems focused on safety and impairment detection.
Artificial intelligence (AI)-driven clinical decision support systems (CDSS) hold promise to improve diagnostic accuracy and efficiency in computational pathology. However, collaboration between human experts and AI may give rise to cognitive biases, such as automation and anchoring bias, wherein users may be inclined to blindly adopt system recommendations or be disproportionately influenced by the presence of AI predictions, even when they are inaccurate. These biases may be exacerbated under time pressure, pervasive in routine pathology diagnostics, or shaped by individual user characteristics. To investigate these effects, we conducted a web-based experiment in which trained pathology experts (n = 28) estimated tumor cell percentages twice: once independently and once with the aid of an AI. A subset of the estimates in each condition was performed under time constraints. Our findings indicate that AI integration generally enhances diagnostic performance. However, it also introduced a 7% automation bias rate, quantified as the number of accepted negative consultations, where a previously correct independent assessment gets overturned by inaccurate AI guidance. While time pressure did not increase the frequency of automation bias occurrence, it appeared to intensify its severity, as evidenced by a performance decline linked to increased automation reliance under cognitive load. A linear mixed-effects model (LMM) analysis, simulating weighted averaging, revealed a statistically significant positive coefficient for AI advice, indicating a moderate degree of anchoring on system output. This effect was further intensified under time pressure, suggesting that anchoring bias may become more pronounced when cognitive resources are limited. A secondary LMM evaluation assessing automation reliance, used as a proxy for both automation and anchoring bias, demonstrated that professional experience and self-efficacy were associated with reduced dependence on system support, whereas higher confidence during AI-assisted decision-making was linked to increased automation reliance. Together, these findings underscore the dual nature of AI integration in clinical workflows, offering performance benefits while also introducing risks of cognitive bias–driven diagnostic errors. As an initial investigation focused on a single medical specialty and diagnostic task, this study aims to lay the groundwork for future research to explore these phenomena across diverse clinical contexts, ultimately supporting the establishment of appropriate reliance on automated systems and the safe, effective integration of human–AI collaboration in medical decision-making.
Accelerating the Approval of Automated Driving Vehicles through standardized XiL test environments
(2026)
Support methodology to the informational and conceptual design of small recreational powerboats
(2024)
This article proposes a methodology for the design of motor recreational boats which are a crucial part of the Brazilian nautical industry. Despite the high demand for these boats, design methodologies for small leisure vessels are not well explored in the literature, leading to not mapped aspects as design information treatment, requirements engineering, and trade-off conflicts. The proposed methodology aims to systematize the obtaining of design specifications, generation and evaluation of solution concepts, employing de- sign support tools. The methodology was evaluated by applying it on a small recreational motorboat, and the results were positively evaluated by nautical industry and academic specialists, leading to acceptation of the methodology and some suggestions for improvement.
Enhancing UX in Automated Vehicles through Biophilic Interfaces: Insights from Prospective End Users
(2025)
Sensor degradation poses a significant challenge in autonomous driving. During heavy rainfall, interference from raindrops can adversely affect the quality of LiDAR point clouds, resulting in, for instance, inaccurate point measurements. This, in turn, can potentially lead to safety concerns if autonomous driving systems are not weather-aware, i.e., if they are unable to discern such changes. In this study, we release a new, large-scale, multi-modal emulated rain dataset, REHEARSE-3D, to promote research advancements in 3D point cloud de-raining. Distinct from the most relevant competitors, our dataset is unique in several respects. First, it is the largest point-wise annotated dataset (9.2 billion annotated points), and second, it is the only one with high-resolution LiDAR data (LiDAR-256) enriched with 4D RADAR point clouds logged in both daytime and nighttime conditions in a controlled weather environment. Furthermore, REHEARSE-3D involves rain-characteristic information, which is of significant value not only for sensor noise modeling but also for analyzing the impact of weather at the point level. Leveraging REHEARSE-3D, we benchmark raindrop detection and removal in fused LiDAR and 4D RADAR point clouds. Our comprehensive study further evaluates the performance of various statistical and deep learning models, where SalsaNext and 3D-OutDet achieve above 94% IoU for raindrop detection.
For expert users to accept Generative AI (GenAI) as a true collaborative partner, it must move beyond simple task-awareness to an understanding of their workflow’s underlying structural rules. This paper introduces a paradigm for AI collaborators that moves beyond simple task awareness to an understanding of the semantic and hierarchical relationships within a component-based system. We investigate this concept within the context of the design-to-code workflow, where inefficiencies arise from the modification of components within design systems. Through two empirical studies with designers and developers, we found that GenAI output was often rejected because it violated the component hierarchy. Designers required granular and visual control for refinements, whereas developers valued automated setup but required transparent validation of the generated code’s logic. Based on these findings, we contribute design guidelines for achieving Component-Structure Awareness (CSA), with two core principles: the Atomic Recommender, which provides assistance that respects the component hierarchy, and Communication Archetypes, which allow GenAI to adapt its interaction style to the user’s role and the atomic nature of their task. This work provides a new, higher-level concept for designing the next generation of truly collaborative GenAI agents.
As automated vehicle technology advances, explainable AI has emerged as a critical tool to enable users to understand and predict the behavior of AI systems, particularly in safety-critical applications such as automated driving. However, increased transparency in AI explanations may inadvertently contribute to an “illusion of control”, a cognitive bias in which drivers overestimate their influence or understanding of the AI’s actions. We aim to better understand how the level of detail in AI explanations affects users of automated vehicles. In a virtual reality study, N = 44 participants experienced different explanation levels (low, medium, high) in an automated ride (SAE L4) compared to a baseline condition with no explanations. The results show a significant improvement in participants’ user experience, acceptance, and explanation satisfaction, with more detailed explanations. Our findings also indicate that as AI explanations become more detailed, users’ perceived level of control increases significantly, although this perception does not correlate with actual control capabilities. At the same time, it decreased their desire to take control, indicating users’ susceptibility to the ’illusion of control’ bias in the context of automated driving. Overall, this suggests that the design of explanation interfaces should strive for a balanced level of detail that promotes AI transparency without causing cognitive overload. At the same time, explainable AI can be utilized to decrease users’ desire to intervene in the AI’s actions.
Enhancing Pedestrian Realism in Adverse-Weather Driving Simulations Using Motion Capture Data
(2025)
Inclusive Vehicle Dashboard Design: Supporting Neuro diverse ADHD Drivers Through Visual Simplicity
(2025)
In urban traffic, while the fraction of collisions involving Vulnerable Road Users (VRU) is low, their importance is high due to the higher injury risk for VRU. Their infrequent occurrence on average (compared with far more common individual perceptual and behavioral errors by both drivers and VRUs) reflects an underlying fault tolerance in traffic processes. However, the degree of fault tolerance varies among traffic situations. The underlying perceptual and cognitive processes involved are complex and can require a high level of attention and concentration, particularly in situations with intersecting trajectories. These processes can occasionally fail, leading to collision risk. The situation of right-turning motorists (in right-hand-drive countries) encountering cyclists moving straight on a bike lane (with right of way) has a particularly low error tolerance, since motorists must actively scan for cyclists approaching from behind. In order to develop, test and assess solutions that mitigate collision risk in this situation, the behavior-related causation mechanisms need investigation. This is the focus of this article. We conducted a trial on our closed test track with n = 35 subjects. The experiment was designed as a within-subject design with three independent factors: maneuver, target velocity, and cognitive load in an n-back task. The trial included observations of participants' gaze control. A primary research focus was the quality and efficiency of the safeguarding gaze behavior of participants in order to draw conclusions on the causation mechanisms of collisions in this situation. For this purpose we define metrics in order to quantify the quality and efficiency of a specific gaze behavior. Furthermore, we studied the effect of factors cognitive load and target velocity on safety and secondary (n-back) task performance. Remarkably, only four out of 35 participants reached a collision risk of 0% relating to the defined quality metric. Furthermore, we identified four distinct gaze strategy groups through hierarchical clustering, where one group performed particularly few glances overall. This group showed significant differences with respect to the defined quality metric whereas the other groups showed only slight differences to each other. The results have implications on subsequent crash causation model development.
To develop truly human-centered automated systems, it is essential to acknowledge that human reasoning is prone to systematic deviations from rational judgment, known as Cognitive Biases. The present study investigated such flawed reasoning in the context of automated driving. In a multi-step study with N = 34 participants, the occurrence of four Cognitive Biases was examined: Truthiness Effect, Automation Bias, Action Bias, and Illusory Control. Additionally, the study explored how the Explainability of the automation’s behavior and the driver’s Mental Model influenced the manifestation of these biases. The findings indicate a notable susceptibility to the Truthiness Effect and Illusory Control, although all biases appeared highly dependent on the specific driving context. Moreover, Explainability strongly impacted the perceived credibility of information and participants’ agreement with the system’s behavior. Given the exploratory nature of the study, this work aims to initiate a discussion on how Cognitive Biases shape human reasoning and decision-making in interactions with automated vehicles. Based on the results, several directions for future research are proposed: (1) investigation of additional cognitive biases, (2) analysis of biases across different levels of automation, (3) exploration of mitigation strategies versus deliberate use of biases, (4) examination of dynamic and context-dependent manifestations, and (5) validation in high-fidelity simulations or real-world settings.
This paper investigates the effects of motion mismatches on simulator sickness and subjective ratings of the motion. In an open-loop driving simulator experiment, participants were driven through a recorded urban drive twelve times, in which mismatches were induced by manipulating the following three aspects in motion cueing: (i) mismatches in specific vehicle axes, (ii) mismatch types (scaling, missing, and false cues), and (iii) inconsistent scaling between different motion axes. Subjects (N=52) reported simulator sickness post-hoc (after each drive), as well as continuously during each drive, a first in simulator sickness research. Furthermore, subjective post-hoc motion incongruence ratings on the quality of the motion were extracted. Results show that longitudinal motion mismatches lead to the most simulator sickness and the highest ratings, followed by mismatches in lateral motion, then yaw rate. False cues induce the most sickness, followed by missing and then scaled motion. Inconsistent scaling between the axes has no significant effect. The continuous sickness ratings support that the occurrence and severity of simulator sickness are indeed related to mismatches in simulator motion of specific maneuvers. This paper contributes to an improved understanding of the relationship between simulator motion and sickness, allowing for more targeted motion cueing strategies to prevent and reduce sickness in driving simulators. These strategies may include the appropriate selection of the simulator, the motion cueing, and the sample of participants, following the presented results.
Driving automation aims to enhance comfort, safety, and traffic flow by removing the human driver from the control loop. However, the human experience of commuting involves more than just reaching a destination or assuming the role of a driver. Factors like personal driving style and courtesy towards fellow road users are integral to the driving experience but often overlooked in the development of driving algorithms for automated vehicles. In this study, we explored the needs of passengers in highly automated vehicles. A qualitative use case analysis was conducted (N=16). In a second study, N=15 participants experienced the resulting use cases in an automated vehicle. In these scenarios, they were able to interact with the automation through a cooperation HMI. Results indicate that most participants expressed a desire for cooperative driving, albeit varying with the driving situation. Moreover, allowing cooperation improves passengers’ overall experience by satisfying psychological needs for autonomy, security, competence, and relatedness.