New research suggests AI chatbots can demonstrate some of the skills used in cognitive behavioural therapy, but their inconsistent performance raises questions about whether they can safely and effectively provide personalised mental health care.
Can AI deliver psychological therapy?
As demand for mental health care continues to grow, artificial intelligence is increasingly being explored as a way to expand access to psychological support. Large language models can produce human-like conversations, but there are still important questions about whether they can deliver therapy effectively, particularly when treatment needs to be adapted to an individual.
A new study published in Computers in Human Behavior: Artificial Humans examined whether an AI chatbot could conduct a full cognitive behavioural therapy (CBT) session and apply recognised therapeutic techniques.
Testing an AI therapy session
The researchers recruited 65 university students experiencing mild to moderate psychological distress, including presentation anxiety, occasional worrying and perfectionism. Each participant took part in a 30-minute session with a locally hosted AI chatbot specifically configured to deliver CBT.
The researchers then assessed the conversations using the Cognitive Therapy Scale, a standard measure of therapeutic competence. This looks at both general skills, such as empathy and collaboration, and more specific CBT skills, including helping people identify unhelpful beliefs and develop new ways of thinking.
The chatbot’s performance was also compared with a meta-analysis of 18 previous studies assessing human therapists using the same scale.
Chatbot performance varied considerably
The chatbot achieved the minimum threshold for adequate clinical competence in 30 of the 65 sessions. Overall, it scored slightly below the average for human practitioners. However, when compared with human therapists in the highest-quality studies, there was no statistically significant difference in scores.
One of the most notable findings was the variation between individual sessions. The chatbot performed well in areas such as expressing empathy, validating feelings and creating a collaborative atmosphere, but was less consistent when applying specific CBT techniques.
It struggled, for example, with identifying important beliefs, guiding users towards their own insights and adapting interventions to the individual.
“We were surprised by how much variation the LLM-chatbot showed in its skillfulness across CBT sessions,” study author Arthur Bran Herbener said.
“This is an important observation, as it suggests that we need research to ensure consistently competent care across individuals, and to understand when and why performance dips.”