← Back to Research News
A diverse TESOL trainee team builds a meaningful multimodal teaching website while comparing human and AI contributions with their instructor
Journal PaperPeer-reviewed study20268 Jul 2026· 3 min

Hong Kong TESOL trainees used generative AI selectively across 10 multimodal website projects

Benjamin Luke Moorhouse, Christoph A. Hafner, Tsz Ying Ho

TESOL Quarterly

500-word summary

A diverse TESOL trainee team builds a meaningful multimodal teaching website while comparing human and AI contributions with their instructor

Listen to the paper summary

Audio summary

0:00/0:00

Moorhouse, Hafner and Ho examine how prospective language teachers used generative AI while creating digital multimodal projects in a Hong Kong MA TESL course. The 13-week course asked groups to build a website as a capstone task, with AI use optional rather than required. Students received a three-hour workshop, submitted projects in week 12 and commented on peers' work in week 13. The study asks not merely which tools appeared, but how participants decided when AI supported or weakened their purposes as multimodal composers and future teachers.

The course enrolled 137 students; 22 consented to the research and represented 10 group projects. Nineteen participants were women and three were men, all aged 22 to 28 and from mainland China. Most reported little or no teaching experience and expected to teach in mainland China, Hong Kong or Macao. Evidence comprised 10 individual or group stimulated-recall interviews lasting about 50 to 65 minutes and the 10 completed websites. Interviews were conducted in Mandarin by the third author, who was not the course teacher, and machine-assisted transcription or translation received human oversight.

All 10 groups used generative AI selectively. Nine used large language models, two used image generators and two used video generators; every group also used multimodal or website-building tools. Across creation activities, eight groups generated images, seven summarized material, six produced data visualizations, six analyzed data and four generated text. Eight groups sought ideas or advice from AI and seven sought feedback. These counts show a broad repertoire rather than uniform dependence: groups combined tools differently and sometimes rejected an AI contribution when it conflicted with the intended message or their desired level of authorship.

Participants described decisions involving efficiency, perceived capability, authenticity, meaning-making and degree of personal involvement. They reported benefits such as confidence, creativity, digital literacy and critical reflection, alongside concerns about reliability and contextual fit. Those are valuable accounts of experience, not objective gains. The study has no comparison group, pre-post assessment, blind scoring or measure of later classroom transfer. Only 22 of 137 students volunteered, all came from mainland China, and the single postgraduate course cannot represent all Hong Kong trainees, practicing teachers or language-learning contexts.

For Hong Kong teacher education, the study offers a concrete curriculum pattern. Programmes can teach multimodal design and AI evaluation together, require students to record which suggestions they accepted or rejected, and assess whether words, images, charts and navigation serve a coherent teaching purpose. A reflection can distinguish assistance with routine production from decisions that require disciplinary, cultural and pedagogical judgment. Bilingual reviewers can check English, Chinese and local classroom context, while protected student or placement data should remain outside unapproved tools.

The defensible finding is that these participants exercised selective agency across 10 authentic projects; the paper does not prove that generative AI improved teacher competence or learner outcomes. Future work could compare scaffolded and unscaffolded cohorts, score products blind to condition, examine delayed independent composing and follow graduates into classroom practice. Meanwhile, the study gives Hong Kong educators a useful assessment target: not the number of AI features used, but the quality of the human reasoning that determines when a generated contribution belongs in a purposeful educational design.

Related papers

A programming lecturer and two diverse university students inspect compiled code, an inheritance diagram, and a grading rubric in a computer laboratory
Journal Paper2026
Journal Paper 102

Five AI systems outscored the average OOP cohort but still failed compilation and advanced concepts

Marina Lepp, Joosep Kaimre

arXiv preprint

Lepp and Kaimre evaluated ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, and Microsoft 365 Copilot on authentic introductory OOP tests and examinations using student grading criteria. Systems exceeded the historical average and often solved long tasks, yet some code did not compile and interfaces, abstract classes, inheritance, and image-based questions remained difficult. The results challenge take-home assessment validity without proving student learning.

programming assessmentobject-oriented programminggenerative AI
Read 500-word summary →
A diverse group of university students explores a branching media-technology learning story while an instructor traces where quiz choices connect to the narrative
Conference Paper2026
Conference Paper 90

AI-generated learning stories were clear and well paced, but their quizzes did not belong in the plot

Finn Rogosch, Andreas Schrader

EDULEARN26 Proceedings

Rogosch and Schrader tested AI-generated interactive-fiction episodes with 22 STEM higher-education participants. The five-to-ten-minute stories were rated clear and appropriately long, but story-content coherence averaged below the neutral midpoint and engagement sat near it. Participants most often questioned why characters suddenly demanded technical answers, showing that a playable educational story can still fail to integrate its learning task.

interactive fictioneducational gamesgenerative AI
Read 500-word summary →
Four diverse adults analyze a business problem with a laptop, charts and an unassisted written follow-up in a workforce-learning laboratory
Journal Paper2026
Journal Paper 54

Generative AI closed three quarters of an education-based performance gap during assisted work, but effort shaped what carried forward

Guillermo Cruces, Diego Fernández Meijide, Sebastian Galiani, Ramiro H. Gálvez, María Lombardi

arXiv working paper

In a preregistered randomized online experiment with 1,174 Argentine adults, GPT-4.1 assistance raised workplace-style problem-solving performance for both education groups and reduced the baseline gap from 0.548 to 0.139 standard deviations. Lower-education participants retained a modest gain after AI was removed, but stronger follow-up performance appeared when intensive assistance was paired with sustained human effort.

generative AIrandomized experimenteducation inequality
Read 500-word summary →