
The people we call unsupported are rarely refusing help, they are stuck in the ‘corridor’: the stretch between noticing that something is wrong and being seen by someone qualified to do anything about it.
That ‘corridor’ especially when it comes to mental health issues can run for months.
An app cannot shorten the route to a qualified clinical assessment, but it can change what happens to a person while they wait, and whether they are still in as much difficulty when their name is finally called.
Many of the people in dire need of help need some sort of interim support, rather than merely being left waiting without protocols or strategies to adopt to lessen their suffering.
Fortune Business Insights values the global mental health apps market at USD 7.5 billion in 2025, projecting USD 25 billion by 2034 at a compound annual growth rate of over 14%. Two segment findings in that report matter more than the 14% headline number.
First, depression and anxiety management are the most keenly felt by the reviewed cohort, driven by conditions common enough that some useful app-based support could reach a meaningful segment of the patient population.
Second, homecare settings dominate end use, because people want support in the evening, at the weekend, and at the point of distress rather than at the point of appointment, which again means that they are looking for support to appear where they are habitually, in their own environments.
Growth is not the same as quality, however. ORCHA, which reviews health apps for NHS organisations across most English regions, scores products against several hundred criteria in three areas: clinical assurance, data privacy, and usability. An analysis of 436 mental health apps drawn from ORCHA’s review data, published in JMIR Human Factors in 2026, found that most fell below ORCHA’s quality threshold, and that many had been built without a qualified health professional involved, without literature behind the content, and without testing to show that the content held up.
Clinical assurance is reliably the weakest of the three scores. There are roughly 20,000 mental health apps in the app stores. Very few of them would survive an independent review.
But there is a problem, apps end up sitting so close to the advice boundary and the criticality of service that flags them as software as a medical device, that utilising AI enablement in this space (especially LLMs) creates a vulnerability around explainability and outcomes, that’s why apps working alongside clinicians is still the safest and most beneficial route, even as we run at pace into the AI age.

Self-guided apps held the largest market share in 2025. But that creates a problem.
The same report puts the hybrid model, combining self-guided tools with therapist or coach involvement, as the most proficient and safe support methodologies out there.
The direction of travel is towards products that connect to a clinician rather than substitute for one, and the recent partnership activity supports that reading.
In June 2026, Limbic partnered with TORTUS to bring regulated ambient voice and AI-scribe capability into Limbic Care, used across NHS Talking Therapies. That is AI reducing the administrative load on therapists so they can see more people, which is a supply-side intervention dressed as a product feature, but it’s a really, really useful one.
In February 2026, Wysa and Imperial College London announced a USD 7.2 million Wellcome-funded programme to adapt an AI-enabled digital mental health programme for adolescent girls in rural India.
Very different problem, but the same product logic: reach people for whom the alternative may well be nothing at all, but do it as an addendum to real-world support.
The TLDR is don’t create solutions where vulnerable people will be left to their own treatment path.
Any product operating in this space will eventually encounter a user in crisis. So, when a model flags risk, a clinician needs to see the reasoning behind the flag, not just the score. As a confidence value with no visible basis cannot be acted on safely, and it cannot be defended in an incident review.
The same applies to the underlying data.
Mental health information is among the most sensitive a person will ever generate. Processing on-device where you can, collecting the minimum you need, and being explicit about what leaves the phone are design decisions with clinical consequences, not compliance tickboxes.
Anything that triages, screens or escalates is likely to sit close to (or over) the medical device boundary. This means MHRA classification, DTAC and clinical safety governance under DCB0129 assessed properly by qualified clinicians and regulatory people before you build, not after your first escalation.
So any app in this space needs serious experience of regulated spaces and the intricacies of building and launching products, which will be viewed close to, if not over the boundary for software as a medical device.

Cross that boundary and a second problem arrives, one that sits underneath explainability rather than alongside it.
A medical device submission rests on a quality management system, and a quality management system rests on being able to show that the software does the same thing in the field that it did in verification. An unconstrained model response cannot give you that. The same input can return a materially different answer from the one your tester saw, and there is nowhere in a submission to put that.
Full visibility of the inputs behind a generated response does not fix it. You can show a reviewer every signal the model saw and still be unable to evidence that the answer will be the same next time. The MHRA has been open about the size of the problem: its AI Airlock sandbox was set up partly to work through hallucination and non-deterministic output in generative products, and has now run two phases with eleven innovators, TORTUS among them.
Mental health products have already failed this test in public. Tessa was a rule-based chatbot built with clinical input for the US National Eating Disorders Association, deliberately restricted to a fixed set of responses so that it could not go beyond what had been tested. A generative question-and-answer feature was added later, and in 2023 the bot started giving people with eating disorders advice on calorie deficits and weight loss. It was pulled within days. The original design held up. The generative layer added on top of it had never been through the same testing, and nobody had treated that addition as a change to the product’s clinical risk profile.
The practical consequence is a boundary drawn inside your own architecture. Generative components belong where a clinician reads the output before a patient
does: note-taking, summarisation, drafting, a triage prompt someone signs off. Anything a patient reads unsupervised needs to be deterministic, versioned and testable, whether that is a fixed content library, retrieval limited to approved material, or a rules engine doing the talking. Drawing that line early will mean you can still write a specification a tester can verify against, which is what makes the product submittable.
As much as AI optimists will tell you that ‘intelligent’ apps that ‘hyper-personalise’ responses are just an inevitability in the next 5 years, we need to be honest about where that contextualisation can create harm.
And in the case of mental health issues, harm is never far from any therapeutic course of action.
If a new service was being offered by a private hospital or as an add on to private medical insurance, let’s think about how we’d plan for best outcomes.
Having built mental health digital products in collaboration with our clients, we understand the difference between supporting content and support-adjacent therapeutic advice.
It isn’t trivial and needs people who deeply understand the space to help, from infrastructure to data handling, interface design to reporting mechanisms.

The evidence base for AI-enabled therapy is still thin on the ground, and the honest position is that it will probably stay thin for some time.
The evidence for AI reducing clinician admin, improving triage accuracy and holding onto people while they wait is stronger and growing.
So the question worth asking of any product in this market is not how sophisticated the model is. It is whether the thing gets a person to another person faster than they would have got there alone.
That’s the standard that needs to be achieved to reduce suffering.

Digital Health
Read more

