New meta-analysis of 51 studies reveals ChatGPT has a *large positive impact* on student learning performance and moderately improves both learning perception and higher-order thinking. → ChatGPT works best in *skills development courses* and when used for *problem-based learning* → Optimal usage period is 4-8 weeks; effectiveness decreases with shorter or longer use → As an intelligent tutor in STEM courses, ChatGPT is particularly effective at enhancing higher-order thinking → Effect size for learning performance is significantly larger than traditional AI assessment tools (g = 0.390) The research suggests specific implementation strategies: provide clear learning frameworks when developing higher-order thinking, encourage cross-grade usage, integrate into different course types (especially STEM), and use ChatGPT flexibly as tutor, partner, or learning tool. We're moving beyond questioning IF AI helps students learn to understanding HOW to optimize its implementation. Paper: https://lnkd.in/eMs2tgJg
Research Methods
Explore top LinkedIn content from expert professionals.
-
-
Ever since ChatGPT arrived, there has been a wave of excitement, skepticism, and curiosity about whether - and how - it actually helps students. Now, a systematic review and meta-analysis by Deng et al. in Computers & Education has pulled together findings from 69 experimental studies, shedding new light on what ChatGPT means for teaching and learning. What Did the Research Reveal? 1️⃣ Stronger Academic Performance Studies show that ChatGPT-assisted interventions often lead to higher grades and better written work - especially in language-rich subjects. One caveat? Many experiments did not make it clear whether students were allowed to use ChatGPT during exams, raising questions about genuine mastery versus AI-assisted output. 2️⃣ Positive Motivation - But Mostly for College Learners University students typically felt more engaged and motivated. In K-12 settings, however, the motivational boost was not as pronounced - suggesting a need for age-appropriate strategies and scaffolds. 3️⃣ Perceived Gains in Higher-Order Thinking Learners reported enhanced creativity and critical thinking. The big “but”: most studies relied on self-reports, so future work needs objective assessments (e.g., problem-solving tasks or performance-based measures) to confirm actual skill growth. 4️⃣ Reduced Mental Effort, Uncertain Self-Efficacy ChatGPT may lighten cognitive load - learners felt tasks were less “taxing.” At the same time, studies showed a mixed or non-significant effect on self-efficacy, implying we need a deeper look at whether students gain real confidence or just convenience. What This Means for Educators & Academics? 1️⃣ Design Rich Assessments: To spot genuine skill gains, use project-based tasks that demand application and originality. 2️⃣ Spell Out Tech Policies: Clearly specify whether and how learners can use ChatGPT - especially for graded work. 3️⃣ Look for Long-Haul Impact: Do not just check excitement levels right after introducing ChatGPT; measure whether those positive vibes (and scores) persist weeks or months down the road. 4️⃣ Mind the Methods: If you are studying ChatGPT’s educational impact, conduct power analyses (to ensure you have enough participants) and randomize group assignments to get the most reliable data. This meta-analysis provides early - but promising - evidence that ChatGPT can enrich students’ learning experiences. The next step? Refining the methods, tracking long-term outcomes, and ensuring actual learning gains are assessed - not just AI’s ability to produce polished outputs. Reference: Deng, R., Jiang, M., Yu, X., Lu, Y., & Liu, S. (2024). Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies. Computers & Education, 105224. https://lnkd.in/eXe8agAT
-
Could willow be the beef farmer's new best friend? Quite possibly, according to a scientific publication from our Northern Irish Lighthouse Farms. In a full field trial at Brook Hall Estate & Gardens, led by the Institute for Global Food Security and QUB School of Biological Sciences, "a 27% reduction in methane production was observed for cattle grazing a willow silvopasture, compared to cattle grazing perennial ryegrass only". The paper concludes: "These findings suggest that integrating willow fodder into beef production systems can help meet emissions reduction targets" (link in comments). Is this the silver bullet for climate-smart ruminant farming that will put the debate to bed for good? Probably not, but it is yet another positive addition to the toolbox of systemic solutions that our Lighthouse Farm community #ArcZero is pioneering across Northern Ireland (link in comments). Watch out for news from Prof. John Gilliland OBE DSc, Rachel Creamer and Deborah De Groot on this week's Soil Health Benchmarks meeting with two generations of ArcZero farmers!
-
The marketing curse 😂 The fix: Dara Denney's 5 step framework that brings data and creatives team together, battle tested with over $100M of ad spend. Dara's take: Creative freedom is a myth. “You need to attack the sources of ambiguity within the creative process. This is the secret to building high performing creative teams" 1. Remove ambiguity with SOPs "The most ambiguous parts of the creating process have the biggest impact on performance" Think of all the ambiguity that exists in your creative production workflow: Research: Who is conducting competitor research? Where is the team documenting customer reviews, and how are you using the performance data you’ve collected? Roadmap: Is everyone clear about the goals and tasks in your creative production pipeline? Or does every new request feel chaotic? Performance: Does your designer know why the last ad bombed? Is data on performance understood or locked in some spreadsheet? To remove ambiguity, Dara suggests formalizing the creative project lifecycle stages research, execution, review, client submission, and launch—for streamlined creation. She calls these stages Standard Operating Procedures (SOPs). 2. Hire a dedicated Creative Strategist Creative strategists remove ambiguity from the creative process by doing the hard work of understanding customer psychology, the competitor landscape, deep s of performance data, and uncovering the strategic problems that ads need to solve. Without a creative strategist, your growth and creative teams become disconnected. For in house teams, this leads to internal politics, mistrust between teams, and low output. 3. Make data accessible AND exciting Not sure which metrics to narrow down on? Focus on your primary KPIs, such as spend, purchases, and cost per lead. These metrics will give you a good understanding of your campaign's performance. Additionally, look at storytelling KPIs, like drop off rates, average video watch time, hook and hold rates, and CTRs. Use a visual analytics platforms to make the data accessible and interesting for your creatives (that's what Motion (Creative Analytics) does btw) 4. Roll out a sprint structure Here's a simple structure you can start with: - Monthly roadmaps, metric checkpoints, bi-weekly retros - Keep the process on track with daily stand-ups Regularly analyze ad formats and metrics as a team during your live sessions and set up a Slack channel for sharing high and low performing ads where you can chat async on what you're seeing 5. Build a data driven creative culture You need to embed Creative Strategy into your org culture. Start all brainstorms with a data download. Ex: share CI research, customer insights, past performance but make sure you start from data or bring it into how you operate. To keep momentum up, create a "win" Slack channel to celebrate learnings and top performing ads and conduct monthly retros to keep the team aligned and engaged with data.
-
UPDATE May 2026: This paper was retracted on April 22, 2026 due to discrepancies in the meta-analysis. See my new post for context. Important new evidence on ChatGPT in education: Wang & Fan's (2025) meta-analysis of 51 studies shows we're at an inflection point. The technology demonstrably improves learning outcomes, but success depends entirely on implementation. The research reveals optimal conditions: sustained use (4-8 weeks), problem-based contexts, and structured support for critical thinking development. Effect sizes tell the story; large gains for learning performance (g=0.867), moderate for critical thinking (g=0.457). Quick fixes don't work. Thoughtful integration does. Particularly compelling: ChatGPT excels in skills development courses and STEM subjects when used as an intelligent tutor over time. The key? Providing scaffolds like Bloom's taxonomy for higher-order thinking tasks. As educators, we have emerging empirical guidance for AI adoption. Not whether to use these tools, but how to use them effectively - maintaining rigor while enhancing accessibility and engagement. The future of education isn't human or AI. It's human with AI, thoughtfully applied.
-
𝗗𝗼 𝘆𝗼𝘂 𝗵𝗮𝘃𝗲 𝗲𝗻𝗼𝘂𝗴𝗵 𝗼𝗳 𝗮 𝗴𝗿𝗼𝘄𝘁𝗵 𝗺𝗶𝗻𝗱𝘀𝗲𝘁 𝘁𝗼 𝗮𝗰𝗰𝗲𝗽𝘁 𝘁𝗵𝗮𝘁 𝗺𝗮𝘆𝗯𝗲 𝗶𝘁 𝗶𝘀𝗻'𝘁 𝗮 𝗿𝗲𝗮𝗹 𝘁𝗵𝗶𝗻𝗴? The text below is from a Twitter thread by Brooke Macnamara reporting the results of her #systematicreview and #metaanalysis of #growthmindset interventions on students’ academic achievement. It's long but the nuance is important. First, some findings from the systematic review: • studies authored by researchers with financial incentives to report positive effects were > 2.5x as likely to report positive effects • > 90% of studies had confounds in their study design • some studies found null results but interpreted them as significant anyway (inc. highly-cited studies) • some studies didn’t adjust for clustering, leading them to erroneously report significant effects (inc. a highly-cited study) • 97% of samples were not preregistered. In fact, there were more studies that described themselves as preregistered that were not preregistered than there were actual preregistered studies. • many studies never tested whether students’ mindsets were affected by the intervention • of the studies that tested whether students’ mindsets were affected by the intervention, many found no evidence of a mindset change Meta-analysis 1: Across all studies, we found a small effect. No theoretically-meaningful moderators were significant. We tested for publication bias using multiple approaches (Egger’s, Duval & Tweedie’s, PET-PEESE). All suggested publication bias. When correcting for publication bias, the overall effect was non-significant. Meta-analysis 2: We next tested if any effects from growth mindset interventions were from the assumed cause—change in students’ mindsets from the intervention. Here, we included all studies that demonstrated the intervention changed treatment students’ mindsets. < 25% of studies demonstrated the intervention changed treatment students’ mindsets. For these studies, the overall effect was non-significant. Meta-analysis 3: We focused on the highest quality evidence. We aimed to only include interventions that changed students’ mindsets & met 100% best practices—e.g., no confounds, full blinding, active control group, no authors with financial COIs. No study met these criteria. We had to considerably lower the threshold for what was considered the highest-quality studies in the growth mindset intervention literature. Among the highest-quality studies available, the effect on academic achievement was not significant. We then conducted over 200 meta-analytic models examining adherence to every combination of best practice criteria. As the number of best practices adhered to increased, the number of significant models decreased. Preprint here: https://lnkd.in/egQctJPt
-
Thematic analysis isn’t a single method, it’s a family of approaches that help us make sense of qualitative data in UX research. Each has its own strengths, structure, and purpose depending on your project goals, timeline, and team setup. Here’s a quick overview of the major types and when they’re most useful: Reflexive Thematic Analysis is ideal when you’re exploring meaning, emotion, or experience. Themes are developed through active interpretation. This method is flexible, deep, and works well for solo researchers or small teams working in open-ended problem spaces. Codebook Thematic Analysis brings more structure. It’s great for teams that need consistent coding. You create a shared codebook, define how to use each code, and track agreement across analysts. Perfect for comparative studies or when stakeholders want clear reliability. Template Analysis blends theory and flexibility. You start with a set of expected themes (based on theory or past research) but revise as you go. It’s well suited for projects where you’re building on existing models or frameworks but still want to stay open to new patterns. Framework Analysis is matrix-based and often used in applied contexts like policy or evaluation. It allows you to systematically compare users and themes in a clear format, which is useful for large, structured datasets or when presenting results to non-research stakeholders. Interpretative Phenomenological Analysis (IPA) is a deep-dive method focused on how individuals make sense of their experiences. You go case by case, then identify patterns across them. It’s especially powerful when you’re designing for sensitive topics or specific identities. AI-Augmented Thematic Analysis uses tools like large language models (LLMs) to support early-stage coding or theme exploration. It can speed up work on large datasets like surveys or reviews, but always benefits from human oversight to ensure context and nuance aren’t lost. Participatory Thematic Analysis involves users in the analysis process. This might include co-coding sessions, theme sorting, or discussions about meaning. It’s especially valuable in community-based or inclusive design work where power-sharing and voice matter. SAMMSA is a structured, step-by-step approach combining summary, micro and meso themes, synthesis, and cross-case analysis. It’s great for training researchers or working with complex qualitative data that needs to be carefully layered and integrated. Textual-Visual Thematic Analysis (TVTA) is designed to analyze visuals and text together-helpful when working with screenshots, photo diaries, or social media posts. It lets you integrate what people say with what they show. Topic Modeling + Thematic Analysis combines computational methods (like LDA) with traditional qualitative coding. It’s useful for very large datasets, letting you identify topic clusters before interpreting them in depth.
-
Very happy to share our latest working paper investigating the impacts of farmer field schools, farmer field schools plus cash, and farmer field schools plus input transfers in southern Malawi. And more is coming! We have funding to go back and interview these same farmers 2 years after the project ended. I'll keep you all posted. Abstract: This study examines the effectiveness of integrated agricultural extension and material support programmes on agricultural productivity and food security among smallholder farmers in Malawi. It evaluates three interventions: Farmer Field Schools (FFS) alone, FFS combined with input transfers, and FFS combined with cash transfers. Using a clustered randomized controlled design and a non-experimental comparison group, the study finds that FFS participation significantly improves food security and agricultural productivity by promoting the adoption of practices such as reduced planting spacing and timely fertilizer application. Both input and cash transfers further amplify these gains, primarily through increased fertilizer use, with no significant differences observed between the two modalities. https://lnkd.in/d49fv5ar
-
🚨 New publication in Motivation Science! 🚨 Why do some teachers, parents, coaches, and leaders adopt highly controlling approaches, while others are far more supportive of people’s motivation? I’m excited to share our new systematic review and meta-analysis—the first of its kind—examining the antecedents of interpersonal communication styles within a Self-Determination Theory (SDT) framework. 🔍 While previous SDT reviews have focused on the consequences of interpersonal styles (need-supportive vs. need-thwarting), none have systematically reviewed what drives these styles in the first place. 📊 Our synthesis draws on 122 studies across education, parenting, sport, healthcare, and more. We: Identified 59 candidate antecedents of interpersonal styles. Grouped them into 13 general factors and 3 higher-order themes: 1️⃣ Socio-contextual factors 2️⃣ Motivators’ personal factors 3️⃣ Motivators’ perceptions of motivatees’ motivation & behaviour Integrated these findings into a model extending the “classic SDT sequence.” 💡 This classification system can help: Pinpoint moderators of intervention effectiveness Identify new targets for interventions to foster more need-supportive styles 📄 Read the full paper here: https://lnkd.in/dhAeCjBf #SelfDeterminationTheory #Motivation #MetaAnalysis #SystematicReview #Psychology #Education #Coaching #Parenting #Leadership Thank you to the action editor Marylène Gagné and two anonymous reviews for a constructive round of reviews. Andreas Ivarsson Dennis Bengtsson Chris Lonsdale Jennie Hancox Eleanor Quested
-
🚨 Our tutoring meta-analysis is now online at RER. Beth Schueler, Grace Falken & I analyzed 263 RCTs. What we learned has important implications for the future of tutoring and the use of meta-analyses to inform policy. Open access: https://lnkd.in/eacY9xT4 In 2020 the research and policy community, myself included, saw tutoring as a promising, evidence-based approach to accelerate learning in the wake of COVID-19 school closures. We were making inferences about the efficacy of tutoring at a scope and scale previously untested. Several prior meta-analyses had found pooled effect sizes of eye-popping 0.3-0.4 SDs. Pooling across multiple studies affords stronger external validity across variable contexts. But we made inferences about LARGE tutoring programs across MANY grades & subjects. Focusing on RCTs strengthens the internal validity of meta-analytic estimates, but does not automatically produce estimates that generalize to broader efforts to scale tutoring across grades and subjects. It turns out that most RCTs evaluate relatively small programs. And most RCTs focus on literacy in elementary schools where there are larger gains to be had by helping students learn the alphabet, phonics, and core preliteracy skills. We sought to refine our meta-analytic sample to have greater external validity for the large tutoring programs that districts and states have attempted to build post-COVID. Effects with stronger target-equivalence are 0.16 to 0.22 SD, relative to an overall pooled effect of 0.40 SD. What drives these declines with scale? Exploratory analyses point to several factors: 1) Real changes to program design features (⬆️ student-tutor ratios; ⬇️ intended dosage) 2) Expanding targeted programs to serve students w/ fewer needs 3) Declining implementation quality (holding program design features fixed). Difficulties cutting through red tape. Delivering low levels of intended dosage. More limited ability to be selective in tutoring hiring. What about programs that were larger but did not stray from best practices in an effort to serve more students on a fixed budget? The evidence is promising! Maintaining "high-dosage/high-impact" features partially inoculates programs from this attenuation at scale. There has been an explosion of high-quality research on tutoring in recent years affirming our findings. Taking tutoring to scale presents both great potential and considerable challenges. I remain optimistic, but also more humble in what we should expect from tutoring. My thinking about what it means to scale successfully has also evolved. I think tutoring might be the type of program that scales best horizontally - more small programs across districts - rather than vertically, where each district attempts to provide tutoring to more of its students. Many thanks to our amazing team of RAs and hats off to the scholars and organizations that have advanced this national effort to scale tutoring.