Moorwell
Pages
Look
Open Moorwell

Our clinical document, audited row by row — including the rows that did not survive.

We went back to the original research behind what this app does and read it. The register is public, and it includes the rows that did not survive. The 28 below are the ones that decided something you can see in the app — a technique it offers, a rule it follows, or something it deliberately will not do. Three of them did not hold up and one could not be checked at all. Those four are in the list with the rest.

The rest are working notes: background reading, the limits that sit on top of the rows below, and ideas the project looked at and chose not to build.

What the audit found
  • Holds 16
  • Holds, with a qualification 8
  • Did not hold 3
  • Unverifiable 1

Each bar is a share of the 28 rows on this page. Choose one to read only those rows.

g = 0.22

95% CI 0.10 to 0.33

Averaged across 18 trials of 22 mental-health apps: how much they moved depressive symptoms, compared with people who were doing something else rather than nothing at all.

The ceiling everything else fits inside

Apps like this one help, and they help by a small amount. That number is the size of it. It is a limit rather than a promise, and nothing claimed on this page is allowed to sit outside it.

None of the studies below is a trial of Moorwell. No outcome study of this app has been run. Every row is evidence about a technique the app draws on, which is a different and weaker thing.

The same number, drawn

g = 0.22 is the measurement. This is what it looks like. Take one person who used an app like this and one person who did something else instead, and compare how their depressive symptoms moved. Do that a hundred times.

50 of them would go the app user's way if these apps did nothing whatsoever. That is a coin toss, and it is the number to compare against — not zero.

56 go their way in fact. The 6 extra squares are the entire effect that 18 trials of 22 apps could find. Small, real, and measured.

Pooling the studies differently moves the figure between 52.8 and 59.2 — the same 95% interval as the one printed beside g = 0.22 above, said the other way. Converted with the probability of superiority, Φ(g ÷ √2), which assumes both groups are normally distributed with equal variance and treats Hedges' g as Cohen's d. It is the same measurement in different units, not a second finding: the conversion is monotone and cannot make a small effect large.

Against people doing nothing rather than something else, the same meta-analysis finds a larger effect. This page shows the smaller one on purpose: an app should be judged against the alternative a person actually has, not against an empty chair.

What d, g and SMD mean

Every number on this page is one of four things, and all four are the same kind of measurement: how far apart two groups ended up.

d
Cohen’s d

The gap between the people who got the thing and the people who did not, measured in standard deviations. 0.5 means the average person who did it ended up better off than about 69 in 100 of those who did not.

g
Hedges’ g

The same idea as d, corrected for small studies, which nudges the number slightly down. Treat g and d as the same scale when you compare them.

SMD
Standardised mean difference

The umbrella term for both. A minus sign only means the thing being measured went down — with sleep problems or low mood, down is the good direction.

d+
Pooled d

A d averaged across many studies at once, so it carries more weight than a single trial’s number.


What counts as big

0.2 small0.5 moderate0.8 large

These labels are conventions, not laws. In mental health a small effect delivered to a lot of people at no cost can matter more than a large one that needs a therapist and a waiting list — which is the argument for an app existing at all.

What the techniques measure, against the honest ceiling

Each bar is the headline effect from the paper named beside it, read from that paper's own abstract. The band behind them is where unguided self-help apps land as whole products — and it is a range rather than a line, because the same apps measure 0.56 against an inactive control and 0.22 against an active one. A strong technique does not lift an app above that band.

  • Implementation intentions — if-then plans d+ = 0.99 · Toli 2016

    k = 28, N = 1,636, clinical and analogue samples, one outlying effect excluded. The outcome is goal attainment, not symptoms — the one bar here measuring something different from the rest.

  • Behavioural activation — against an inactive control g = 0.83 · Stein 2021

    28 studies. The design credits this to “Cuijpers 2023”; it is Stein, Carl, Cuijpers, Karyotaki & Smits 2021, and the abstract records that study quality was generally low with evidence of publication bias.

  • Behavioural activation — against an active control g = 0.15, n.s. · Stein 2021

    Same meta-analysis, same 28 studies, the other comparator. The difference was not significant. This bar is on the chart because leaving it off is how 0.83 gets quoted.

  • Self-compassion — on self-criticism g = 0.51 [0.33–0.69] · Wakelin 2022

    19 papers, 1,350 participants. The paper’s own moderator analysis found larger reductions against passive controls than against active ones.

  • Best possible self — on positive affect d+ = 0.511 · Carrillo 2019

    29 studies, N = 2,909. This is the largest of four outcomes in that paper: wellbeing is 0.325 and optimism 0.334. A reader who assumed “wellbeing” would be reading a number well over half again too big.

  • Sleep support, ages 10–24 SMD = 0.32 [0.11–0.53] · Salamanca-Sanabria 2025

    18 RCTs, N = 13,296, mean age 19, 55.5% unguided — the closest population and delivery format to this app of anything in the register. Sleep quality; the same paper's insomnia-severity figure is an order of magnitude apart and is not usable.

  • Memory specificity — on symptoms d = 0.29 · Barry 2019

    k = 13, against control at post-intervention. The design plots 0.47, which is that paper’s within-group pre-post change and not comparable with the rest of this chart. The abstract adds that the benefit over control was mostly lost by follow-up.

  • Gratitude — on wellbeing g = 0.22 [0.11–0.33] · Kirca 2023

    25 RCTs, 6,745 participants, against neutral comparison groups. The smallest bar here, and the one whose confidence interval most nearly reaches the floor of the band.

The shaded band is where unguided self-help apps land as whole products — 0.22 against an active control and 0.56 against an inactive one (Firth et al. 2017, 18 RCTs, 3,414 participants). A technique above the band is not this app above the band. Every citation, comparator and interval here is written into the project’s research registry, with the paper each one came from.

Every technique in the app, and the paper behind it

The three findings that did not hold up, and the one we could not check at all, are here at the same size as everything else. A list that shows only its wins is an advert.

Verdict
Narrow the register
Mechanism
Population

Showing all 28 rows.

  • Details for What an app of this kind can move at all

    What an app of this kind can move at all

    Firth et al. 2017

    Survives
    Study
    Meta-analysis of RCTs · 18 RCTs, 22 apps, N = 3,414
    Population
    Adults with depressive symptoms, mixed settings
    Replication
    Direction replicated by a later 176-RCT meta-analysis

    This is the governing ceiling for the whole project, and every claim on this site has to fit inside it. Against an active comparator — people who were doing something else, not nothing — the pooled effect is g = 0.22 (0.10–0.33).

    PMID 28941113 10.1002/wps.20472

  • Details for The same question, asked again across 176 trials

    The same question, asked again across 176 trials

    Linardon et al. 2024

    Survives
    Study
    Meta-analysis · 176 RCTs; N = 33,567 depression, 22,394 anxiety
    Population
    Adults, mixed settings
    Replication
    Consistent in direction and size with the row above

    Apps containing CBT do better than apps that do not, which is why the content bank stays anchored in cognitive and behavioural material rather than in general wellness.

    PMID 38214614 10.1002/wps.21183

  • Details for Guidance changes the result more than the content does

    Guidance changes the result more than the content does

    Moshe et al. 2021

    Survives
    Study
    Meta-analysis · 83 studies, N = 15,530
    Population
    Adults with depression
    Replication
    —

    Guided programmes outperform unguided ones, and trials run in real services underperform trials run in ideal conditions. Moorwell is unguided and free, so the bottom of that range is the honest expectation for it.

    PMID 34898233 10.1037/bul0000334

  • Details for Nerves reframed as your body getting ready

    Nerves reframed as your body getting ready

    Jamieson et al. 2016, 2022

    Qualified
    Study
    Two randomised classroom trials, both from the group that built the manipulation · 93 and 339
    Population
    Community-college students sitting real exams — the best population match in the whole register
    Replication
    A 2026 replication across twelve courses at seven institutions, against a placebo control, did not reproduce it. That study is the next row, and it is the reason this one reads Qualified rather than Survives

    The part that holds is the control arm, and it decides a design question: “ignore the stress” was the placebo condition. That is why the app never tells anyone to calm down before an evaluation — a rule that rests on what the trials compared against, not on the size of what they found.

    PMID 34292050 10.1177/1948550616644656 10.1037/xge0000893

  • Details for …and the replication that did not find it

    …and the replication that did not find it

    Thormodsæter et al. 2026

    Qualified
    Study
    Replication study in real courses · 12 courses across 7 institutions, intervention against an active placebo
    Population
    Undergraduates in real courses
    Replication
    This is the failure, and it is cited beside the finding it weakens

    A 2026 replication did not reproduce the effect. The project's own audit records that the claim needs a dated correction beside it naming this study. It is on this page for the same reason it is in the register: a citation that only survives when you leave out the replication is not a citation.

    PMID 41730015 10.1187/cbe.25-04-0055

  • Details for Writing your worries down before a test

    Writing your worries down before a test

    O'Meara & Lovett 2026

    Did not hold
    Study
    Meta-analysis · 21 studies, N = 1,457, 30 effect sizes
    Population
    Students taking real tests — exact population, exact moment
    Replication
    The pooled result is the correction

    Single-session expressive writing before an exam is one of the most repeated pieces of study advice there is. Pooled, it does not hold. The app does not do it, and this row is why.

    PMID 40968670

  • Details for Lines that make a claim about the reader

    Lines that make a claim about the reader

    Wood, Perunovic & Lee 2009; Flynn et al. 2020

    Qualified
    Study
    Lab experiments, plus a replication report titled “On the failure to replicate…” · Small samples
    Population
    Undergraduates
    Replication
    Mixed — the failure is reported and cited

    One study found that “you are enough” lands worst on the person who least believes it. Two later attempts did not find the same thing. Mixed evidence about exactly the people most likely to read the line is a reason to place it carefully rather than to use it or drop it, so the app holds those lines back on a bad day and never writes one about you.

    10.1111/j.1467-9280.2009.02370.x 10.1016/j.jcbs.2020.03.003

  • Details for Speaking to yourself the way you would to a friend

    Speaking to yourself the way you would to a friend

    Wakelin et al. 2022

    Survives
    Study
    Meta-analysis · 29 studies reviewed; the pooled effect is k = 28, N = 1,636
    Population
    People with clinical diagnoses or elevated symptom scores — the best population match for any pooled effect in the register
    Replication
    None located

    This is the strongest row the register holds for a technique the app actually ships.

    PMID 33749936 10.1002/cpp.2586

  • Details for Self-criticism predicting depression on its own

    Self-criticism predicting depression on its own

    Gittins & Hunt 2020

    Did not hold
    Study
    Prospective, three waves, cross-lagged · 243 adolescents, mean age 12.08
    Population
    Early adolescents
    Replication
    The clearest direct test located, and it does not support the claim as stated

    The app still answers self-criticism, because the row above supports the technique. What died is the stronger claim that self-criticism predicts later depression independently of current distress. Both halves are on this page.

    PMID 33339021 10.1371/journal.pone.0244182

  • Details for Putting the feeling into words

    Putting the feeling into words

    Lieberman et al. 2007

    Qualified
    Study
    Single fMRI study · Small imaging sample
    Population
    Volunteers viewing standardised images
    Replication
    No meta-analysis located — a search of the literature returns zero title-level meta-analyses of affect labelling

    This is a single imaging study carrying a design rule, and the register says so. What survives it is narrow and is what the app does: the naming has to be yours, so the check-in asks what is going on and never offers a feeling to pick from a menu.

    PMID 17576282 10.1111/j.1467-9280.2007.01916.x

  • Details for Finding a more exact word for it

    Finding a more exact word for it

    Nook et al. 2018

    Survives
    Study
    Developmental study · Large developmental sample
    Population
    Across adolescence into adulthood
    Replication
    —

    Emotion differentiation follows a curve with its low point in adolescence — which is the age the app was built for, and a reason the vocabulary it offers matters.

    PMID 29878880 10.1177/0956797618773357

  • Details for Absolute words as a signal

    Absolute words as a signal

    Al-Mosaiwi & Johnstone 2018

    Survives
    Study
    Three text-analysis studies · 63 forums, 6,400+ members
    Population
    Self-selected internet forum members
    Replication
    None located

    Absolutist words track severity better than negative-emotion words do. The population is people who chose to post on a forum, which is not a general sample, and the register records that rather than smoothing it.

    PMID 30886766 10.1177/2167702617747074

  • Details for A plan with a when and a where in it

    A plan with a when and a where in it

    Toli, Webb & Hardy 2016

    Survives
    Study
    Meta-analysis · Clinical and analogue mental-health samples
    Population
    People with mental health problems
    Replication
    Supported by a later synthesis of 642 tests

    “I will do more work” is a wish. “After my last lecture, at the desk by the window” is a plan. The app stores the when and the will as two separate fields, because the shape is the part doing the work.

    PMID 25965276 10.1111/bjc.12086 10.1080/10463283.2024.2334563

  • Details for Doing one small thing, first

    Doing one small thing, first

    So et al. 2025

    Qualified
    Study
    Single RCT · N = 67, ages 20–30, 8 weeks
    Population
    Young adults 20–30
    Replication
    None located

    A single small trial of a behavioural-activation app. The population is close to the app's own, and one trial is one trial — the register grades it qualified for that reason, and we have kept that grade.

    PMID 41301298

  • Details for Slower breathing, at about six breaths a minute

    Slower breathing, at about six breaths a minute

    Fincham et al. 2023

    Survives
    Study
    Meta-analysis · k = 12, N = 785
    Population
    General adults
    Replication
    Mechanism replicated, including a published null on the in-to-out ratio

    The null is the interesting half: a 2024 study found that lengthening the exhale specifically made no difference to heart-rate variability during slow-paced breathing. The app's exercise still runs longer out than in, and we do not claim the ratio is what does the work.

    PMID 36624160 PMID 38507210

  • Details for Letting the muscles go

    Letting the muscles go

    Donato et al. 2026

    Qualified
    Study
    Meta-analysis · 31 RCTs, 2,277 adults
    Population
    Medically ill and inpatient adults
    Replication
    Contradicted as a component of CBT by a component analysis; loses to active comparators

    The population does not transfer — the register marks this one no on that column outright, and the app treats relaxation accordingly rather than as a headline technique.

    PMID 41633054

  • Details for Mindfulness, and why it is not the app's answer

    Mindfulness, and why it is not the app's answer

    Goyal et al. 2014, and three later syntheses

    Survives
    Study
    Meta-analyses · 47 trials / 3,515; 51 RCTs; 83 studies / 6,703; 43 studies / 1,427
    Population
    Mixed; one of the four is university students
    Replication
    —

    Meditation programmes are no better than any active treatment they were compared against, and adverse events are reported in a measurable minority. That is not a reason to think badly of mindfulness. It is a reason this app did not build its centre on it.

    PMID 24395196

  • Details for Sleep, and which parts of the therapy carry it

    Sleep, and which parts of the therapy carry it

    Furukawa et al. 2024; a student trial

    Survives
    Study
    Component network meta-analyses, plus a null RCT · 241 trials / 31,452; 80 / 15,351; 195 students
    Population
    Mean age 45.4 in the largest; the student trial is the null
    Replication
    Two independent network analyses converge on the same two components

    The two components that survive are the two the app implements: a wake time you keep and a bed kept for sleeping. The population gap is real and the project's clinical file states it — the big evidence is middle-aged, and the trial closest to a student sample is the null one.

    PMID 38231522 PMID 41345462

  • Details for Digital CBT reaching student anxiety

    Digital CBT reaching student anxiety

    Oliveira et al. 2023

    Survives
    Study
    Meta-analysis of RCTs · 15 studies, N = 1,619 university students
    Population
    University students — exact population
    Replication
    Largest synthesis located for this population

    This is the closest the register comes to evidence for this kind of programme in this kind of person. It is still a category result, not a result about Moorwell.

    PMID 36873307 10.1016/j.invent.2023.100609

  • Details for Worry and rumination are one process

    Worry and rumination are one process

    Stenzel et al. 2025

    Survives
    Study
    Meta-analysis · Multi-study synthesis
    Population
    Mixed
    Replication
    —

    Treatment does not care which register the thought is in. That is why the app aims at the loop rather than at whether someone calls it worrying or going over things.

    PMID 39916353 10.1017/S0033291725000017

  • Details for Postponing a worry, tested against an active control

    Postponing a worry, tested against an active control

    Versluis et al. 2016; a 2026 replication

    Survives
    Study
    Two RCTs against an active thought-log control · 51 and 117 adults with clinical GAD
    Population
    Adults with diagnosed GAD, not students
    Replication
    Replicated, and against an active comparator rather than a waiting list

    The register marks the population transfer no as evidenced: this is a clinical GAD sample. What does transfer is the delivery — a phone, a short window, repeated prompts.

    PMID 26511764 10.1111/bjhp.12170

  • Details for Media portrayal, and the rule it sets for the app

    Media portrayal, and the rule it sets for the app

    Domaradzki 2021

    Survives
    Study
    Literature review · 108 papers
    Population
    General media audiences; effects vary by age
    Replication
    Long-standing and repeatedly observed

    How a thing is described changes what happens next. This is the row underneath the app's safe-messaging rules, and it is the reason the crisis pathway is fixed wording rather than anything generated.

    PMID 33804527 10.3390/ijerph18052396

  • Details for Gratitude, across 28 countries

    Gratitude, across 28 countries

    Choi et al. 2025

    Survives
    Study
    Meta-analysis of RCTs · 25 RCTs, 6,745 participants in the earlier synthesis
    Population
    Mixed adults, many countries
    Replication
    Direction supported across countries, with significant between-country variation and no moderator that explains it

    It works, it is small, and nobody can yet say why it varies by country. For an app whose first population is Indian students, an unexplained cross-country variance is exactly the kind of thing worth printing.

    PMID 40627390 10.1073/pnas.2425193122

  • Details for How long people actually use an app like this

    How long people actually use an app like this

    Baumel et al. 2019

    Qualified
    Study
    Panel usage analysis of real installs · 93 apps, 10,000+ installs each
    Population
    Android app users in the wild
    Replication
    —

    Real-world engagement with mental-health apps is far below what trials report. Moorwell is designed around that rather than against it — nothing resets, nothing is lost by a gap, and the app never mentions that you were away.

    PMID 31573916 10.2196/14567

  • Details for Choice overload

    Choice overload

    Iyengar & Lepper 2000

    Did not hold
    Study
    Three studies, including the jam field study · Small
    Population
    Shoppers and students
    Replication
    Failed. A 2010 meta-analysis found the mean effect to be about zero

    The app still offers few options at a time, and the reason is not this study. It is that a long menu is more work to read at 2am. The behaviour survived; the citation under it did not, and swapping the reason is cheaper than pretending.

    10.1037/0022-3514.79.6.995 10.1086/651235

  • Details for Growth mindset, and the size it actually is

    Growth mindset, and the size it actually is

    Sisk et al. 2018; Macnamara & Burgoyne 2023

    Survives
    Study
    Two meta-analyses plus a later sceptical synthesis · Large, multi-study
    Population
    Students
    Replication
    The 2023 synthesis is the sceptical follow-up and it agrees with the 2018 one

    The claim that survives here is the cautious one: brief mindset interventions run to a very small effect, and the later literature moved further that way rather than back. This is a case study in a finding that reached the whole world before the pooling did, and the app builds none of it.

    10.1177/0956797617739704 10.1037/bul0000352

  • Details for Evidence about Indian students specifically

    Evidence about Indian students specifically

    Louis 2016; Kumar et al. 2023

    Unverifiable
    Study
    Cross-sectional descriptive studies · 500 in one; not stated in the other
    Population
    Indian students and professionals — the app's own population
    Replication
    Searches run and recorded; no outcome or intervention study located

    This is the most important row on the page. For the population this app was built for, the literature that exists is descriptive — it measures how common something is, not whether anything helps. The register grades that unverifiable rather than leaving the column blank, and the site is not going to imply otherwise.

    PMID 27833225 PMID 37034864

  • Details for The one Indian outcome trial the search did locate

    The one Indian outcome trial the search did locate

    Michelson et al. 2020; Gonsalves et al. 2022

    Qualified
    Study
    Outcome trial with 12-month follow-up, plus a pilot RCT of the app version · School students aged 12–20 in New Delhi
    Population
    Delivered in Hindi by a human lay counsellor, against a printed booklet
    Replication
    The app version is a pilot

    Structured problem-solving is the only candidate in the whole research pass with an outcome trial in this population. It was delivered by a person, and the project's own rule for that pass says no row in it may be described to a reviewer as evidence for an unguided app.

    PMID 32585185 PMID 36573376

How to read a verdict

Survives
We found the paper. It says what we said it says, and the study behind it is strong enough for the use we put it to.
Qualified
We found the paper, and our claim about it is too broad — the wrong group of people, a weaker kind of study, or one paper asked to carry more than it can.
The feature may still be right. The reason we gave for it is not good enough as written.
Did not hold
Someone ran the study again and did not get the same answer, or better evidence points the other way.
Unverifiable
We could not find the paper at all, or we could not trace the number back to the source it was credited to.
A claim that fails stays on the list, with the date and the correction next to it. Delete it, and the next person works it out from scratch and believes it again.

A psychologist has read the app’s own words

Everything above is this project auditing the research it leans on. Reading what the app actually says to you is a different job, and a psychologist has done it — the wording changed where she asked for it.

Read on