Refine
Year of publication
- 2026 (1)
Document Type
- Master's Thesis (1)
Language
- English (1)
Has Fulltext
- yes (1)
Is part of the Bibliography
- no (1)
Institute
Generative AI tutoring tools predominantly refine the wording, structure, and feasibility of a student’s initial problem framing. Whether AI assistance that instead poses counter-proposals – alternative framings, stakeholder perspectives, and questioned assumptions – changes the originality of the resulting framing is an open empirical question for data science education. This thesis investigates that Question in an exploratory study with thirty postgraduate students at Universität Koblenz. Each participant interacted with a bespoke web platform that walked them through three fixed-order stages on a sharedWi-Fi-dormitory dataset description: an Editorstyle AI stage, a Challenger-style AI stage, and a final unassisted synthesis. The final synthesis topic was rated on four dimensions (originality, feasibility, reasoning quality, clarity) by a human expert rater and an LLM second rater (Anthropic Claude) blinded to group assignment.
Because the planned counterbalanced within-subjects crossover could not be implemented, participants were sorted post hoc into an Editor-influenced and a Challengerinfluenced group based on triangulated self-reported impact and preference. Topics from the Challenger-influenced group were rated substantially higher on originality (Cohen’s d = 3.67, 95% CI [2.50, 4.84], p < .001, ICC = 0.859); topics from the Editor-influenced group were rated higher on feasibility (d = −2.04, 95% CI [−2.92,−1.16]) and clarity (d = −1.07, 95% CI [−1.84,−0.30]); a secondary unexpected association favoured the Challenger-influenced group on reasoning Quality (d = 1.03, 95% CI [0.27, 1.79]). The latter three dimensions had inter-rater reliability below the pre-specified 0.70 threshold and their effect-size magnitudes are interpreted as exploratory estimates. Perception data and open-ended responses suggested participants treated the two AI styles as functional complements assigned to distinct task-contexts, and 94% indicated interest in an integrated switcher mode. Because the analytic groups were defined by self-report and because the stage order was fixed, these results should be read as associations rather than causal effects of prompt style. The thesis contributes (a) empirical evidence of large betweengroup differences in framing originality consistent with a counter-proposal mechanism; (b) a methodological case study of a hybrid human-LLM rating protocol with dimension-specific reliability documentation; and (c) design directions for educational AI tools that surface, rather than choose for the learner, the kind of cognitive assistance offered.