
AI-generated learning stories were clear and well paced, but their quizzes did not belong in the plot
Finn Rogosch, Andreas Schrader
EDULEARN26 Proceedings
Resumen de 500 palabras

Rogosch and Schrader examine a useful intermediate question about generative AI in educational games. Before measuring whether an AI-generated story improves learning, can learners follow it, accept its length and experience the educational content as part of the story? Their EDULEARN26 paper evaluates short, choice-driven interactive-fiction episodes for higher education. The work deliberately studies perceived quality rather than knowledge gains, treating clarity, coherence and engagement as prerequisites for a later learning-outcome trial.
The authors used SINE, a domain-agnostic pipeline that combines an open-weight language model with deterministic validation and repair. They fixed the pipeline around Qwen3 14B, then created 20 media-technology content seeds covering sampling, quantization and compression. Three generations per seed produced a controlled pool; after automated playability and validation filters, 48 scenarios remained. The validator checked reachability and content fidelity, but it could not decide whether a technical question felt causally necessary inside the plot. Each participant received one English-language episode intended to take five to ten minutes. The prompt configuration and content base were held constant so variation came mainly from narrative generation.
Twenty-two adults with a STEM higher-education connection completed the online study: eight students, eleven university staff members and three recent graduates. They played 19 distinct scenario files, then answered a short German-language questionnaire. Ten positive Likert items measured narrative clarity, story-content coherence, engagement and length acceptance. Gameplay telemetry recorded duration and quiz responses, while an open prompt gathered comments. The analysis was descriptive because the sample was small.
Clarity and length were the strongest results. Narrative clarity averaged 4.11 on the five-point scale and length acceptance 4.14; median playtime was 5.7 minutes. Engagement averaged 3.08, with its confidence interval spanning the neutral midpoint. Story-content coherence was the bottleneck at 2.92, with most of its confidence interval below neutral. First-try quiz accuracy averaged 0.71, and correlations between perceived quality and gameplay measures were small to moderate and not statistically significant. Low coherence therefore was not simply a reaction to getting answers wrong.
Ten participants left comments. Six questioned the artificial in-story motivation for quiz prompts: characters appeared to demand technical knowledge without a believable narrative reason. Others noted abrupt changes of location or character, obvious distractors, repeated questions and missing story consequences for wrong answers. Positive remarks appeared alongside these criticisms, suggesting that participants accepted the interactive-fiction format while rejecting the seam between the plot and the quiz.
The evidence is intentionally limited. This was a convenience sample from one institution, one STEM sub-domain, one pipeline-model configuration and one exposure. Adapted scales were not fully validated, qualitative coding used one unblinded rater and no learning outcome was measured. For AIEDHK, the design lesson is nevertheless concrete: technical playability and verbatim quiz fidelity are weak proxies for educational integration. A better pilot should make questions causally necessary to the story, give choices meaningful consequences, use semantic rather than verbatim content checks, review complete paths with educators and learners, and only then test independent knowledge and transfer.


