Serious games in pharmacy education stop at engagement and short-term knowledge, not clinical competency

An umbrella review finds serious games in pharmacy education boost engagement and short-term knowledge, but not clinical competency or patient outcomes.

Direct answer

A 2026 scoping umbrella review of ten secondary evidence sources concludes that serious games in pharmacy education are consistently linked to learner engagement and short-term knowledge gains, but evidence for workplace behavior and patient outcomes remains limited [1]. The review maps this pattern onto Miller's pyramid and Kirkpatrick's model, showing that most assessments cluster at the Knows and Knows How tiers rather than the Shows How or Does levels [1]. Critically, seven of ten included reviews pooled digital and physical interventions without reporting modality-specific outcomes, which weakens any claim that digital game affordances independently drive learning [1]. A 2025 randomized comparative study in neurologic emergency medicine found that a serious game produced comparable multiple-choice test results to a traditional clinical case seminar, reinforcing that knowledge gains are achievable but not uniquely superior [2]. Together, these findings position serious games as supplementary formative tools rather than validated instruments for high-stakes competency certification [1].

8sources cited

This article was generated with WisPaper-powered search and paper analysis.

What earlier evidence established about games and simulation in health education

Before the anchor review, the field had already demonstrated that serious games could produce measurable knowledge gains in medical education. A 2025 randomized comparative intervention study by Heidrich et al. compared a serious game on epileptic seizures and ischemic strokes against a traditional clinical case seminar and a no-intervention control in 205 final-year medical students, finding no significant difference in multiple-choice test results between the serious game and seminar groups at either immediate or three-week follow-up [2]. This established that well-designed serious games can match conventional teaching for factual knowledge retention, but it did not test whether that knowledge translates to workplace behavior or patient outcomes [2]. Separately, a 2023 systematic review of high-fidelity assessments in pharmacy education examined student perceptions of summative simulation-based assessments, providing a precursor frontier that highlighted the need for stronger validity evidence and competency alignment in pharmacy simulation [3]. These earlier works collectively set the baseline: games and simulations can teach, but the field had not systematically asked whether they change clinical practice.

The evaluation frameworks themselves were also well established. Miller's pyramid distinguishes four tiers of clinical competence—Knows, Knows How, Shows How, and Does—while Kirkpatrick's model separates reaction, learning, behavior, and results [1]. A 2023 systematic review applying Kirkpatrick's model to interprofessional simulation activities involving pharmacy students found that six included studies failed to justify the simulation's learning outcomes to competencies, indicating that alignment between educational interventions and clinical competency frameworks was already a recognized weakness [4]. This validation evidence confirms that the problem the anchor review identifies is not new, but the anchor review is the first to map the entire serious games pharmacy literature against these frameworks at the review level [1].

The anchor review's contribution: mapping ten reviews onto competency frameworks

The anchor scoping umbrella review searched Web of Science, PubMed, and Scopus through December 31, 2025, for publications from January 1, 2015, to December 31, 2025, screening 914 unique records and ultimately including ten secondary evidence sources—eight from database searching and two from backward citation tracking [1]. The included reviews were published between 2017 and 2023, with five appearing in 2023 alone, and their first authors were based in six countries including the United States (n=4) and Australia (n=2) [1]. Populations were predominantly pharmacy trainees (n=9), with some reviews extending to practicing pharmacists and pharmacy support personnel (n=4) and one including patients receiving medication management interventions [1]. Interventions spanned digital game-based learning, gamified instructional modules, educational escape rooms, and computer-based simulation systems, with application settings primarily in higher education teaching environments (n=9) [1].

Using Miller's pyramid, Kirkpatrick's model, and the Graafland framework for serious game validity, the review found that engagement and short-term knowledge were the outcomes most frequently reported across the included reviews [1]. Evidence for workplace behavior and patient outcomes remained limited, clustering at the lower tiers of both frameworks [1]. AMSTAR 2 appraisal rated nine of ten included reviews as critically low and one as low, driven mainly by missing or incomplete items such as prospective protocol registration, explicit exclusion lists, and integration of risk-of-bias assessments into interpretation [1]. The authors interpret these ratings as indicators of reporting and appraisal transparency rather than as a definitive rejection of the underlying educational studies, noting that several included reviews were scoping or structured narrative reviews designed to map the field rather than estimate pooled effects [1].

Why mixed-modality reporting undermines attribution to digital games

A central finding of the anchor review is that seven of ten included reviews pooled digital and physical interventions without reporting modality-specific outcomes, limiting the ability to attribute observed benefits to digital-native affordances such as real-time feedback, adaptive difficulty, or interaction logging [1]. For example, reviews by Sera and Wheeler, Hope et al., Oestreich and Guy, and Kanaan et al. synthesized outcomes across screen-based and paper-and-pencil games without separating digital from physical effects [1]. Similarly, Hintze et al. aggregated digital platforms with physical lockbox clues, and Garnier et al. combined physical practice hoods with 3D interactive programs without reporting digital features separately from the broader simulation [1]. Only three reviews—Abraham et al., Silva et al., and Gharib et al.—focused explicitly on pure digital environments, but even these showed similar limitations in assessment validity and competency translation, relying largely on self-report scales and cognitive testing rather than stealth assessment or high-order clinical behavioral data [1].

This aggregation problem is not merely methodological housekeeping. Because digital and physical modalities differ in their affordances, pooling them can obscure whether observed benefits relate to digital feedback, repeated practice, social collaboration, novelty, or other instructional features [1]. The anchor review treats media confounding as a limitation in review-level interpretation rather than as proof that digital tools lack independent value, but it concludes that current evidence cannot isolate the incremental contribution of digital interventions from surrounding course designs or instructional contexts [1]. This limitation is compounded by the broader problem that control variables are rarely considered as sources of heterogeneity in systematic reviews; a 2025 umbrella review of 50 systematic reviews of observational studies found that differences in confounder sets were mentioned as a potential source of heterogeneity in only 5 of 50 reviews (10%), with control for mediators and colliders mentioned in zero reviews [7]. While that review focused on observational epidemiology rather than educational interventions, it illustrates a generalizable weakness in how systematic reviews handle contextual variables that could similarly confound serious game outcomes [7].

The assessment validity gap: from self-report to stealth assessment

The anchor review found that many included reviews reported reliance on post-hoc self-report satisfaction surveys and cognitive pre/post-test analyses, measures that capture subjective experiences and short-term knowledge gains but have limited ability to capture decision-making processes or behavioral change [1]. Viewed through Miller's pyramid, this evaluative pattern clusters evidence mainly at the Knows and Knows How tiers, with fewer data addressing clinically meaningful Shows How or Does outcomes [1]. The review notes that stealth assessment—real-time capture of decision pathways, response times, and resource-use patterns through back-end interaction logs during naturalistic interaction with a digital system—remains underexamined in the current secondary evidence base [1]. Most studies still treated digital tools primarily as content-delivery platforms rather than as environments that could support more embedded assessment [1].

This assessment gap is not unique to pharmacy education. A 2023 systematic review of high-fidelity assessments in pharmacy education found that six studies failed to justify the simulation's learning outcomes to competencies, and the review called for stronger validity evidence in summative simulation-based assessments [3]. The anchor review extends this concern by showing that even when digital games are used, the assessment paradigms remain largely unchanged from traditional paper-and-pencil or self-report formats [1]. Under the Graafland framework, stronger claims about educational effectiveness require evidence relevant to construct and concurrent validity; if assessment modalities remain confined to tests that cannot capture authentic clinical behaviors, it becomes difficult to substantiate a link between in-game performance and professional competence [1]. The review therefore positions digital serious games as better suited for formative assessment and contextualized training than for high-stakes competency certification, and any future high-stakes use would require stronger construct-validity, concurrent-validity, and objective performance evidence [1].

Boundaries of the conclusion and what remains uncertain

The anchor review's conclusions are explicitly bounded by its scoping umbrella design and the limitations of the included reviews. The authors state that the synthesis is best read as a cautious map of review-level evidence and reporting gaps, not as a complete primary-study traceback or reappraisal [1]. The search was restricted to three databases—Web of Science, PubMed, and Scopus—and the authors acknowledge that education-, psychology-, nursing-, and engineering-focused databases may contain additional records not captured by this strategy [1]. The AMSTAR 2 ratings, while indicating widespread reporting limitations, should not be read as a judgment that the underlying educational studies are invalid, because several included reviews were mapping-oriented rather than effectiveness-oriented [1]. The review cannot prove that serious games are ineffective for clinical competency translation; it can only show that the current review-level evidence does not support strong claims about such translation [1].

The broader evidence base on digital interventions in health professions education also suggests that context and implementation factors matter greatly. A 2026 hybrid umbrella review of nurse-led tele-palliative care found that telehealth interventions improved symptom management and family support across 28 studies, but the authors cautioned that findings should be interpreted cautiously due to heterogeneity in study designs, reliance on secondary evidence, and variable methodological quality [5]. Similarly, a 2024 umbrella review of the Mediterranean Diet for cardiovascular prevention found that better quality reviews are needed and that AMSTAR 2 ratings were frequently critically low, illustrating that reporting quality challenges are not unique to educational technology reviews [6]. A 2025 umbrella review of mental disorder symptoms among university students found that 65% of included meta-analyses were rated critically low on AMSTAR 2, further demonstrating that methodological transparency is a widespread issue in secondary evidence synthesis [8]. These parallel findings reinforce that the anchor review's cautious conclusions reflect a general state of the review literature rather than a specific failure of serious games research [1][5][6][8].

About These Sources

This research page is built on 8 peer-reviewed studies — published from 2023 to 2026, 6 from 2024 or later, collectively cited 189 times — selected as the most relevant from 13 studies that passed quality screening, drawn from 85 papers retrieved from a database of over 500 million.

Sources used in this answer

1

Intervention fidelity and competency translation of serious games in pharmacy education: a scoping umbrella review

The anchor scoping umbrella review of ten secondary evidence sources found that serious games in pharmacy education most frequently report engagement and short-term knowledge outcomes, with limited evidence for workplace behavior and patient outcomes, and that seven of ten reviews pooled digital and physical interventions without modality-specific reporting, limiting attribution to digital affordances [1].

2

Education Research: teaching neurologic emergencies through serious games: a randomized comparative intervention study

A 2025 randomized comparative intervention study in neurologic emergency medicine found no significant difference in multiple-choice test results between a serious game group and a traditional clinical case seminar group among 205 final-year medical students, establishing that serious games can match conventional teaching for factual knowledge but do not demonstrate superior knowledge gains [2].

3

Shifting to authentic assessments? a systematic review of student perceptions of high-fidelity assessments in pharmacy

A 2023 systematic review of high-fidelity assessments in pharmacy education examined student perceptions of summative simulation-based assessments and found that six included studies failed to justify simulation learning outcomes to competencies, highlighting the need for stronger validity evidence in pharmacy simulation [3].

4

The application of Kirkpatrick's evaluation model in the assessment of interprofessional simulation activities involving pharmacy students: a systematic review

A 2023 systematic review applying Kirkpatrick's evaluation model to interprofessional simulation activities involving pharmacy students found that six studies failed to justify the simulation's learning outcomes to competencies, validating concerns about competency alignment in pharmacy simulation education [5].

5

Nurse-led tele-palliative care for symptom management and family support: A hybrid umbrella review of reviews and primary studies

A 2026 hybrid umbrella review of nurse-led tele-palliative care found that telehealth interventions improved symptom management and family support across 28 studies, but cautioned that findings should be interpreted cautiously due to heterogeneity in study designs, reliance on secondary evidence, and variable methodological quality [6].

6

The effectiveness of the Mediterranean Diet for primary and secondary prevention of cardiovascular disease: An umbrella review

A 2024 umbrella review of the Mediterranean Diet for cardiovascular prevention found that better quality reviews are needed and that AMSTAR 2 ratings were frequently critically low, illustrating that reporting quality challenges are not unique to educational technology reviews [7].

7

An umbrella review reveals that control variables are rarely considered as a source of heterogeneity in systematic reviews of observational studies.

A 2025 umbrella review of 50 systematic reviews of observational studies found that differences in confounder sets were mentioned as a potential source of heterogeneity in only 5 of 50 reviews (10%), with control for mediators and colliders mentioned in zero reviews, demonstrating that control variables are rarely considered as sources of heterogeneity in systematic reviews [8].

8

Prevalence of mental disorder symptoms among university students: An umbrella review.

A 2025 umbrella review of mental disorder symptoms among university students found that 65% of included meta-analyses were rated critically low on AMSTAR 2, further demonstrating that methodological transparency is a widespread issue in secondary evidence synthesis [10].