← Back to Research News
A diverse group of educators reviews an instructional video storyboard, narration timeline, and multimedia-learning checks at a university media lab
Conference PaperConference paper202623 Aug 2026· 2 min

Dual gatekeeping improved ratings of AI-generated instructional video without testing student learning

Yearim Kim, Njun Baek, Nojun Kwak

CHI 2026 Workshop on Understanding and Engaging Critical Resistance to AI in Education

500-word summary

A diverse group of educators reviews an instructional video storyboard, narration timeline, and multimedia-learning checks at a university media lab

Listen to the paper summary

Audio summary

0:00/0:00

Yearim Kim, Njun Baek and Nojun Kwak study a common risk in AI-generated educational video: a polished output can still sequence ideas poorly, include distracting material or misalign narration and visuals. Their PedaCo system treats resistance to an AI draft as a designed part of authoring rather than a failure. It combines educator review before rendering with automated checks after video synthesis.

The first gate operates at the script stage. Educators select criteria from Mayer's Cognitive Theory of Multimedia Learning, including coherence, signaling, redundancy, contiguity, segmenting and pre-training. An AI reviewer flags possible violations by principle, and the educator can reject, revise or regenerate the script. The system presents its comments as possible problems rather than final judgments. Reviewing text upstream is cheaper than repairing a pedagogical error after narration and visuals have been rendered.

The second gate evaluates the finished video on five dimensions: coherence, redundancy, temporal contiguity, modality and image quality. The educator sees the scores and decides whether to accept the video or return to the script. The design intentionally reserves context-sensitive decisions, such as tone and learner appropriateness, for people while using automation for structural features that are easier to measure.

The authors evaluated the workflow in two ways. In a within-subject study, 23 educators briefed on the multimedia-learning principles compared reviewed videos with videos generated without those guidelines across three topics representing causal, abstract and procedural demands. Mean ratings increased from 3.07 to 3.86 on a five-point scale. The paper reports statistically significant improvements across the principles, including large gains in prerequisite sequencing, removal of irrelevant material and overall instructional validity. Participants also rated production efficiency 4.26 out of five.

A separate automated comparison used 14 videos covering seven science and philosophy topics, with two conditions per topic. Coherence rose from 0.646 to 0.729 and temporal contiguity from 0.273 to 0.294; both differences were statistically significant. Modality, redundancy and image quality did not differ significantly. The authors keep those measures as possible regression checks rather than claiming universal improvement.

The evidence is promising but early. The paper is a four-page workshop report, the educator sample is small, and participants were already briefed on the framework used for judging. The study evaluates ratings and proxy metrics, not comprehension, retention, transfer or accessibility for learners. It also does not establish whether repeated review remains manageable during ordinary teaching or whether automated flags work equally well across languages, ages and subjects.

For AIEDHK, the useful contribution is the placement of two explicit acceptance gates. A school or university could require a teacher-reviewed script before generation, then inspect timing, coherence, captions, accessibility and source accuracy on the final media. A stronger trial should compare student learning and teacher workload across workflows, pre-register outcomes, include bilingual materials and report disagreements between reviewers and metrics. PedaCo supports a practical principle: generated teaching media should remain "not yet" until both professional judgment and inspectable checks support its use.

Related papers

A university student explains a geometry construction to a lecturer while a classmate follows and a laptop displays a related digital diagram
Industry7 Sept 2026
Industry 112

Commentary: Astra's AGI claim puts evidence of human learning at the centre of education

AIED.HK Editorial

AI Product News Commentary

OpenAI launched GPT-6 Astra on 3 September 2026 amid claims about the arrival of AGI. This commentary treats that label as a claim, not an established consensus. For education, the immediate challenge is to distinguish what an AI can produce from what a learner can explain, question and transfer independently—and to use stronger agents to support that learning.

product newscommentaryGPT-6 Astra
Read 500-word summary →
Three education and software colleagues review illustrated lesson cards, an annotated chart and a digital prototype in a bright university design studio
Industry7 Sept 2026
Industry 113

Commentary: Fable 5.1 brings longer AI workflows to AIED—and makes educational validation more important

AIED.HK Editorial

AI Product News Commentary

Anthropic released Claude Fable 5.1 on 1 September 2026 with stronger long-running coding and knowledge-work capabilities and cheaper cache reads. For AIED, the opportunity is a faster cycle from teaching idea to reviewable prototype and research analysis. The test is whether teams can turn that speed into better pedagogy and credible evidence, while accounting for total cost, data conditions and human review.

product newscommentaryClaude Fable 5.1
Read 500-word summary →
A diverse group of university students and a lecturer examine clustered dialogue cards, a ten-trait matrix and an exam-progress chart in a bright learning analytics studio
Tool / Dataset2026
Tool / Dataset 94

Principal Trait Analysis linked AI-tutor dialogue patterns to outcomes, but not yet to transferable skills

Hunter McNichols, Kai Du, Andrew Lan

arXiv preprint

McNichols, Du and Lan introduce Principal Trait Analysis, an LLM-assisted pipeline that turns human-AI conversation traces into interpretable behavioral traits. On 1,540 university AI-tutor sessions and 2,774 professional coding-agent sessions, selected traits added explanatory or predictive signal beyond prior performance. Conceptual questioning aligned positively with some exam outcomes, but cross-semester inconsistency, contradictory coefficients and mostly flat temporal patterns mean the traits cannot yet be treated as transferable AI collaboration skills.

Principal Trait AnalysisAI tutoring dialoguehuman-AI collaboration
Read 500-word summary →