Moorwell
Pages
Look
Open Moorwell

You write a sentence. What comes back was chosen, not composed.

Moorwell does not write to you. It reads what you wrote, works out what it is allowed to offer, and picks one. Every sentence it can say already existed before you opened it, and a person could read all of them.

Five stages, and the order is the safety property

The stages run in this order every time, on every input. Nothing later can undo something earlier — a claim you could check by reading the code, not a diagram drawn after the fact.

  1. You say how the day went

    A word, a sentence, a paragraph. It is read and answered on the device. What you typed picks a reply and is then gone — not written to a database, not synced.

    Held back: nothing. The person is the only source of input — there is no sensor, no microphone and no background reading of anything.

  2. It is read for danger, before it is read for anything else

    The risk check runs first, on every input, always. It is not one option among several planners, it cannot be switched off, and it cannot be swapped out. It is a table of reviewed phrases matched as whole words, with negation handled — deliberately a list rather than a model, so a clinician can read it and say exactly what it will and will not catch.

    Held back: the phrases, and every rule about them. What is publishable is the shape and the honesty about its limit: the project writes this down as a floor, not a guarantee, and the always-visible route to help exists precisely because this layer will sometimes miss.

  3. The day is placed in a band, and a model may only raise it

    Not a diagnosis and not a score. A band is a coarse reading of how much room there is for anything demanding, and it decides what the libraries are willing to hand over. Where a model has a say at all, it is allowed to move the reading in exactly one direction: it may raise an assessed level of distress, and it may never lower one.

    Held back: the bands, their boundaries, and what moves a reading between them. The one-way rule is the part worth publishing, because it is the part that is checkable: it is written once so no call site has to remember it, restated at six more places in the code as defence in depth, and held by its own test. A model that could lower a reading turns a misread into a withheld crisis resource, and that is the one failure this app cannot absorb.

  4. Eligibility removes everything that may never appear

    Hard rules live here and nowhere else. If the answer to may this ever appear in this state? is no, the item is filtered out — never scored down. The gates sit inside the libraries themselves rather than on the screens that call them, so a screen cannot get it wrong by forgetting.

    Held back: which rule applies where. The reason for it is a bug this project shipped: a hard rule expressed as a score penalty is not a hard rule, because under any ranking penalised means unlikely rather than never — and a game once appeared on a good day for exactly that reason.

  5. Ranking chooses among what is left, and hands back an identifier

    Only preferences live here — what tends to suit you, what you said you care about, what fits the length of time you have. Ranking sorts inside the gate and can never open one: a better-matching line at the wrong band loses to a worse-matching line at the right one.

    Held back: the scoring. The output type is the most important fact on this page: the last stage returns an id, not a sentence, the identifier of an entry in the bank. That is what makes every word the app can say reviewable in advance.

Three guarantees you can verify yourself

Each is an absence in the build, which is a stronger promise than a policy.

  • No generated text

    Nothing is composed at runtime. Every string ships in the app and can be read before you install it.

  • No upload to understand you

    Assessment runs on the device. Aeroplane mode changes nothing about the reply.

  • No diagnosis

    A band is not a condition. Moorwell never names one, and never implies you have one.

The guard that makes review possible at all

A model in this system never produces the words you read. It is handed a slate of things it may choose from, and what it returns is a choice — validated against the catalogue before anything is drawn, and discarded whole if it does not match.

What a model is not allowed to return

"You've had a hard week. Try being kinder to yourself tonight."

Prose, composed on the spot, arriving at a person who may be having the worst evening of their year. Nobody reviewed it, because it did not exist until the moment it was shown.

What it returns instead

{ "pick": "val-x4" }

An identifier from a slate the engine built. The entry it names was written by a person, carries its own review status, and sits behind the same gates as everything else. If the identifier is not on the slate, the whole response is thrown away rather than partly applied.

As built today there is no wire field that could carry prose in the first place — the type has nowhere to put a sentence. The validator exists anyway, and the on-device and cloud paths go through the same one, so the remote path can never be checked less carefully than the local one.

The whole vocabulary of the app, counted

Everything it can say ships inside it, which is what makes reading all of it possible.

  • 2,227 written entries Across 25 libraries — assets/content/.
  • 0 lines generated at runtime Every sentence ships in the app — which is what makes reading all of it before you install it possible.
  • 1,848 of them reviewed The other 379 are not, and this site counts that as the weaker state rather than rounding it up.

Why rules, and not a model

A clinician can read every sentence this app is able to say.

That is a property no system that writes its own words can offer, at any budget, because the sentences do not exist until they are shown.

How do we know that Close

The packet a clinician would read is built and current: the 2,227 entries counted above, sorted into 232 groups, every one gated and traceable. What that count does not tell you is who did the checking.

Machine-audited
1,823 — a tool checking each entry against the clinical rationale it was written from
Read by a psychologist
The words Moorwell says have been read by a psychologist — for the wording, and for whether each one is safe to say.

The risk check is deliberately rule-based, and recall is chosen over precision.

A false positive shows somebody a route to help they did not need, which is mildly patronising. A false negative misses somebody who did. Those are not symmetrical, so the design is not symmetrical either.

How do we know that Close

The code says so in its own header, and it gives three reasons. The first is auditability: a neural decision boundary cannot be reviewed the way a table can, and this is the one part of the app where being wrong is not recoverable. The second is the asymmetry above. The third is the limit — keyword matching cannot catch every disclosure, so the project records it as a floor rather than a guarantee.

There is no model anywhere in the risk path. Its only imports are language-pack utilities; there is no embedder, no runtime, no inference of any kind. One diagram in the project's internal documentation draws an optional embedder into that box, and it is wrong — the path does not exist, and the code and the safety-invariant note both say so independently.

The loops, and what interrupts each one

Moorwell does not go after a diagnosis. It goes after the loops that keep a bad week going — the ones that quietly make tomorrow harder than it needed to be. Every part of the app names the loop it interrupts.

A loop, drawn

Withdrawal is the easiest one to see, so start there. Read it round once, then cut it.

The withdrawal loop

You do less

Fewer good things

Mood drops

Less feels possible

This one is withdrawal. Nothing in the ring points out of it: each step makes the next more likely, so a bad week keeps its own shape without anything new going wrong. That is what maintaining means.

Behavioural activation takes out the last step. You do one small thing without waiting to feel like it, and the doing supplies the good thing that went missing. Three steps are still there; the fourth, the one that started the ring again, is gone.

Rumination
Going over the same thing again and again. It feels like working the problem out; it is not, and the mood drops as it goes.
Affect labelling. Putting a feeling into words takes some of the heat out of it: in the imaging, naming it quietened the alarm part and brought the deliberate part in. The naming has to be yours, so Moorwell asks what is going on and never hands you a feeling to pick from a menu.
Withdrawal
You do less, so fewer good things happen, so your mood drops — the loop drawn above.
Behavioural activation. Action first, motivation second: waiting until you feel like it hands the decision to the part of you that is already flat. Moorwell suggests one small thing, and keeps two kinds apart — things you are good at, and things you enjoy. Under pressure people drop the second kind first.
Self-criticism
"I am not good enough." This is the part that turns ordinary pressure into something that hurts.
Self-compassion. Answering a bad moment the way you would answer a friend in it. Not praise, and not talking yourself up: what changes is the tone of the voice that shows up with the setback. Moorwell never tells you that you are enough. One study found a line like that landing worst on the person who least believes it, and two later attempts did not find the same thing. Where to put such a line is a decision Moorwell takes carefully rather than confidently.
Threat appraisal
Reading a racing heart before an exam as proof you are failing, rather than as your body getting ready.
Arousal reappraisal. The signal stays the same and the label on it changes: a fast heart before an exam is blood and oxygen arriving where they are wanted, which is what a body does before something that matters. Moorwell never tells you to calm down before an evaluation — that instruction was the placebo arm in the trials. The register grades this Qualified rather than Survives. Two field trials found the effect; a broader replication across 12 courses at 7 institutions, against an active placebo, found none — on test anxiety or on performance. The honest reading is that it holds in the labs that built it and has not travelled. Both are in the register.
Sleep disruption
Bad sleep and low mood each feed the other. Sleep is usually the one that will move first.
Sleep scheduling and stimulus control, two standard parts of CBT for insomnia: a wake time you keep, and a bed kept for sleeping. Moorwell asks for a wake time and a wind-down, and keeps a diary you fill in by hand. No sensor is allowed to write to it.
Emotion-driven procrastination
Putting the task off to put the feeling off. Then the guilt makes the next attempt harder.
Implementation intentions. A plan shaped as "when this happens, I will do that". The cue does the remembering, so starting stops being a decision you take again every time. Moorwell stores the when and the will as two fields, because the shape is the part doing the work.

Two techniques are not aimed at a loop at all — they are for the ten minutes you are in. Paced breathing slows you to about six breaths a minute. Grounding gives your eyes and your hands a job, so the loop has less room to run. The five-four-three-two-one form of grounding is taught nearly everywhere and has not been tested much on its own.

Three things it does not do

Each was measured rather than estimated, and each is a claim this site could have made without anybody catching it.

Often assumed
That it notices when you are ruminating and answers that specifically.
Measured
The signal exists and has never fired. Its two conditions cannot both be true at once, so no input can satisfy it.
Run against the project's own corpus of check-in notes.
Often assumed
That a reframe is matched to the mechanism behind your week.
Measured
The machinery exists, and the mechanism reading is absent for the large majority of check-ins — for most people, most of the time, what arrives was chosen well on other grounds.
Same corpus.
Often assumed
That it learns what helps you.
Measured
It learns what you are interested in. The preference it keeps is read by the parts that choose topics; nothing reads it to decide what worked.
That capability is on the Vision page, in the future tense, where it belongs.

What it does instead

Each choice below cost something that would demo well and bought something better in its place.

In Moorwell
What other apps ship here, and what this one does in its place.
Nothing to collect, no bar, nothing that can be broken
What almost every app like this ships. A promised reward makes people want the thing less for its own sake; a bar that hardly moves during a hard week is a report card; and a counter that resets meets coming back after a gap with a failure. Moorwell notices what you already did, afterwards, in a list that can only get longer, and never mentions the gap. The full set of these absences, and the tests that hold them, is on Safety.
It asks, and takes your answer
The clever feature everybody wants to build. Apps that time themselves to your mood are an early idea with thin support behind it. Moorwell works from what you typed and what you told it, such as the date of an exam. When the reading of what you wrote is poor it says so and asks you to put it another way, because guessing confidently is the failure worth avoiding.

One engine we wrote, and two small models that make it less literal

The rules are the product and run on their own. The two models sit on top of them to catch what a rule written in advance cannot: the phrasing nobody thought of, and the feeling underneath it.

The engine, and it is ours

A deterministic planner over a bank a person wrote — the five stages above. It is the product: the app is complete with both models below switched off, and it behaves the same way twice on the same input, which is what makes it reviewable at all.

A sentence-embedding model, for the phrasings we did not anticipate

Small, open-licensed, and from the retrieval family rather than the generative one. Its job is to notice that what you wrote and something in the bank are about the same thing even when they share no words. It changes which line is chosen — never what the line says.

A compact emotion classifier, which is allowed to say it does not know

It reads for feeling, and its promise is precision on the readings it accepts rather than accuracy on all of them — so where it is not confident it abstains, and the engine proceeds without it. It may raise how much difficulty the app assumes you are in. It may never lower it.

Both models run on your device, and neither produces a sentence you read. And neither is anywhere near the part of the app that reads for danger — that path is keyword matching over reviewable phrase lists, with no model of any kind in it, because it is the one place where being wrong is not recoverable.