Presentation
The Use of Artificial Intelligence Systems in Mental Health Treatment: A Thematic Review
SessionPoster Session 1
DescriptionArtificial intelligence (AI) systems have been rapidly advancing in their sophistication, particularly with large language models (LLMs). These models have received the colloquial term "chatbot" for their ability to replicate human conversation. However, in spite of its appearance to the user, the model has no real “comprehension” of the meaning behind the words it's presenting on the screen. Instead, it predicts the next most likely word or sentence based on a data set. This raises some questions regarding chatbots' effectiveness in various scenarios that may require a full understanding of what is being said in the conversation. One such scenario is in the field of mental health treatment, which has seen an increase of AI supplementation in recent years. In a 2024 survey conducted by Cross et al. on Australian mental health practitioners (n = 86) and members of the general Australian population above the age of 16 (n = 107), approximately 28% of the 107 community members reported using AI. Of those, 60% reported using AI for rapid mental health relief, and 47% reported using AI as a "personal therapist." With such a large amount of AI users utilizing it as a mental health aid, it is important that this phenomenon is studied further.
To explore the topic of AI in mental health care, an initial search was conducted by inputting general keywords into Embry-Riddle Aeronautical University's academic search tool that includes over 100 databases. The articles found from those keywords were then reverse searched to gather an overview of the topic. From the patterns found in the initial search, four keyword sets were developed to look deeper into the literature, and a secondary search was conducted. The combination of the initial and secondary search led to a total of 51 articles being reviewed, resulting in 5 themes and 8 general best practices being extracted from the literature.
The common themes found in the literature include interaction quality (27 articles), explainability (16 articles), trust (16 articles), ethical concerns (17 articles), and anthropomorphism (9 articles). Subthemes for interaction quality included engagement, workload, and the system's usability. For ethical concerns, subthemes included privacy and safety. These themes contribute to the understanding of the best practices, which are as follows:
The theme of interaction quality informed our first best practice. That is, developers should consider demographic differences in mental health care and design the system with those differences in mind. Training data sets have the potential to be biased, so it is important to be aware of how sampling bias may affect the effectiveness of its responses.
Continuing with the theme of interaction quality, our second best practice found that the design of chatbots should aim to help foster the relationship between the therapist and patient rather than replace it entirely. The importance of the clinician-patient relationship was highlighted throughout the literature, and the AI models studies were not shown to be as effective as human clinicians.
Based on the trust and explainability themes, the third best practice is to use a "Glass Box" model rather than a "Black Box" model for AI design. These terms relate to the level of explainability, and increased explainability was associated with higher levels of trust and better patient outcomes throughout the literature.
Interaction quality in tandem with ethical concerns informed the fourth best practice. Designers should develop the system to be encouraging towards patients. Users have experienced frustrations with the responses of AI chatbots, and patient encouragement is important to mental health outcomes.
Themes of explainability and interaction quality derived the fifth best practice: the iterative design process should consistently include clinician input. The incorporation of users and subject matter experts when available is standard best practice for user experience research as it allows developers to understand things they may have missed in the initial design.
Based on the themes of interaction quality and explainability, the sixth best practice is to provide some form of training, such as a tutorial, to users of the system. Some systems in the literature were AI-assisted peer support systems, but designers cannot expect users to automatically have sufficient knowledge of the system or how to use it to provide support to their peers.
Interaction quality and ethical concerns both played a role in the seventh best practice, which is to conduct response testing to ensure that the chatbot is listening and responding appropriately. Some research showed LLMs, especially those in early development, to provide risky responses to mental health prompts, which may cause harm to users.
The final best practice was determined through the anthropomorphism theme. While anthropomorphism may be helpful in establishing trust, system designers should try to avoid the "uncanny valley" effect. Due to this phenomenon, increasing anthropomorphism may also lead to an increase in psychological distress and may decrease trust in the system.
From the literature reviewed, future research directions may include longitudinal studies to analyze the long-term effects of AI usage in mental health settings, developing strategies for conveying AI accuracy, and further validation of different scales for AI usability such as the Artificial Social Agent Questionnaire (ASAQ) by Fitrianie et al. (2025) or the Chatbot Usability Scale (CUS) by Borsci et al. (2022).
References
Borsci, S., Malizia, A., Schmettow, M., van der Velde, F., Tariverdiyeva, G., Balaji, D., & Chamberlain, A. (2022). The Chatbot Usability Scale: The design and pilot of a usability scale for interaction with AI-based conversational agents. Personal and Ubiquitous Computing, 26(1), 95–119. https://doi.org/10.1007/s00779-021-01582-9
Cross, S., Bell, I., Nicholas, J., Valentine, L., Mangelsdorf, S., Baker, S., Titov, N., & Alvarez-Jimenez, M. (2024). Use of AI in mental health care: Community and mental health professionals survey. JMIR Mental Health, 11. https://doi.org/10.2196/60589
Fitrianie, S., Bruijnes, M., Abdulrahman, A., & Brinkman, W.-P. (2025). The Artificial Social Agent Questionnaire (ASAQ) — Development and evaluation of a validated instrument for capturing human interaction experiences with artificial social agents. International Journal of Human-Computer Studies, 199, 103482-. https://doi.org/10.1016/j.ijhcs.2025.103482
To explore the topic of AI in mental health care, an initial search was conducted by inputting general keywords into Embry-Riddle Aeronautical University's academic search tool that includes over 100 databases. The articles found from those keywords were then reverse searched to gather an overview of the topic. From the patterns found in the initial search, four keyword sets were developed to look deeper into the literature, and a secondary search was conducted. The combination of the initial and secondary search led to a total of 51 articles being reviewed, resulting in 5 themes and 8 general best practices being extracted from the literature.
The common themes found in the literature include interaction quality (27 articles), explainability (16 articles), trust (16 articles), ethical concerns (17 articles), and anthropomorphism (9 articles). Subthemes for interaction quality included engagement, workload, and the system's usability. For ethical concerns, subthemes included privacy and safety. These themes contribute to the understanding of the best practices, which are as follows:
The theme of interaction quality informed our first best practice. That is, developers should consider demographic differences in mental health care and design the system with those differences in mind. Training data sets have the potential to be biased, so it is important to be aware of how sampling bias may affect the effectiveness of its responses.
Continuing with the theme of interaction quality, our second best practice found that the design of chatbots should aim to help foster the relationship between the therapist and patient rather than replace it entirely. The importance of the clinician-patient relationship was highlighted throughout the literature, and the AI models studies were not shown to be as effective as human clinicians.
Based on the trust and explainability themes, the third best practice is to use a "Glass Box" model rather than a "Black Box" model for AI design. These terms relate to the level of explainability, and increased explainability was associated with higher levels of trust and better patient outcomes throughout the literature.
Interaction quality in tandem with ethical concerns informed the fourth best practice. Designers should develop the system to be encouraging towards patients. Users have experienced frustrations with the responses of AI chatbots, and patient encouragement is important to mental health outcomes.
Themes of explainability and interaction quality derived the fifth best practice: the iterative design process should consistently include clinician input. The incorporation of users and subject matter experts when available is standard best practice for user experience research as it allows developers to understand things they may have missed in the initial design.
Based on the themes of interaction quality and explainability, the sixth best practice is to provide some form of training, such as a tutorial, to users of the system. Some systems in the literature were AI-assisted peer support systems, but designers cannot expect users to automatically have sufficient knowledge of the system or how to use it to provide support to their peers.
Interaction quality and ethical concerns both played a role in the seventh best practice, which is to conduct response testing to ensure that the chatbot is listening and responding appropriately. Some research showed LLMs, especially those in early development, to provide risky responses to mental health prompts, which may cause harm to users.
The final best practice was determined through the anthropomorphism theme. While anthropomorphism may be helpful in establishing trust, system designers should try to avoid the "uncanny valley" effect. Due to this phenomenon, increasing anthropomorphism may also lead to an increase in psychological distress and may decrease trust in the system.
From the literature reviewed, future research directions may include longitudinal studies to analyze the long-term effects of AI usage in mental health settings, developing strategies for conveying AI accuracy, and further validation of different scales for AI usability such as the Artificial Social Agent Questionnaire (ASAQ) by Fitrianie et al. (2025) or the Chatbot Usability Scale (CUS) by Borsci et al. (2022).
References
Borsci, S., Malizia, A., Schmettow, M., van der Velde, F., Tariverdiyeva, G., Balaji, D., & Chamberlain, A. (2022). The Chatbot Usability Scale: The design and pilot of a usability scale for interaction with AI-based conversational agents. Personal and Ubiquitous Computing, 26(1), 95–119. https://doi.org/10.1007/s00779-021-01582-9
Cross, S., Bell, I., Nicholas, J., Valentine, L., Mangelsdorf, S., Baker, S., Titov, N., & Alvarez-Jimenez, M. (2024). Use of AI in mental health care: Community and mental health professionals survey. JMIR Mental Health, 11. https://doi.org/10.2196/60589
Fitrianie, S., Bruijnes, M., Abdulrahman, A., & Brinkman, W.-P. (2025). The Artificial Social Agent Questionnaire (ASAQ) — Development and evaluation of a validated instrument for capturing human interaction experiences with artificial social agents. International Journal of Human-Computer Studies, 199, 103482-. https://doi.org/10.1016/j.ijhcs.2025.103482
Event Type
Poster Presentation
TimeMonday, March 234:45pm - 6:15pm EDT
LocationRhinelander Gallery
Digital Health
