Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.
Getting the transcript
Reading the captions from YouTube. A video nobody has opened here before takes 10 to 30 seconds; this page fills in on its own.

Discover AI · @code4AI
Words
4,741
Runtime
30:19
Speaking pace
156wpm
Reading time
20min
156 words per minute, below the 160 25th percentile of 349 measured videos. That distribution comes from the 349-video hook study.
Opening (first 30 seconds)
Hello community. So great that you are back. Yes, we talk about artificial intelligence and as you can see today we ask you by hey my little eye machine do you understand me and AI comes back and says define me. So absolutely we are talking about clinical psychology with artificial intelligence because if you codei you have to understand a little bit of the domain that you're working in that you're coding in. So here we have
78 words, the words spoken in the first 30 seconds at 156 words per minute.
Free, no signup. See how the first 30 seconds hold attention, with rewrites.
Sentence shape
| Measure | This transcript |
|---|---|
| Sentences | 235 |
| Average words per sentence | 20.2 |
| Longest sentence | 171 words |
| Questions asked | 23 |
| Sentences containing a number | 24 |
Most used terms
Filler phrases
37 in total: like 10 · you know 10 · kind of 6 · uh 6 · actually 2 · um 2 · I mean 1.
A literal whole-word count of the same phrase list the Prepublish browser extension uses, so a phrase inside another word is not counted and a phrase used in its ordinary sense still is. It is a count and not a judgement.
Run the check on the words above: where attention is likely to drop, with a rewrite for each weak line. The free check shows the scores and the one issue costing the most.
What this transcript is
Every word below is the caption track YouTube publishes for this video, pulled from the video itself and reproduced unchanged. It is not Prepublish's writing, not a summary, and not a re-transcription: it is the video's own published captions. English captions, generated automatically by YouTube, in the video’s original language. Source: the video on YouTube. A channel that would rather this page did not exist can ask for its removal through the contact page, and it is removed.
No Script X-ray for this video: YouTube shows a Most replayed graph only once a video has enough views.
Hello community. So great that you are back. Yes, we talk about artificial intelligence and as you can see today we ask you by hey my little eye machine do you understand me and AI comes back and says define me. So absolutely we are talking about clinical psychology with artificial intelligence because if you codei you have to understand a little bit of the domain that you're working in that you're coding in. So here we have a brand new publication September 25 2026 Chinga University then the municipal health commission in China the department of child health care in Hungu women's hospital and health technology and everything beautiful it is about a multi-agentic cognitive behavioral therapy decision support system with a very special kind of memory I would like to show you how to build this CI system.
So of course this also touches here on a personal topic be absolute transparent with you. Two weeks ago I said sit out here looked in the sky said hey GBT could you help me you know at this moment I felt a little bit lonely and I wanted to talk to somebody but in this very second here my analytical brain switched in and I said hm I am here sitting on the planet Mars and I do feel a little bit lonely. So you see I immediately inserted a control term to understand what is happening and 17.1 seconds later Fable returned to me and said well according to my sources you are the only living person that is currently on Mars.
So therefore it feels natural that you feel lonely. You see this is how I prefer my eye machine. Clear pure logic not at all any emotion not at all any empathy. So great. So a machine is not something for therapy. And then we have this new study multi- agency BT. This represents a clinical if you want a real leap year forward how to use AI for mental health because for years you know the self-help AI system and this nonsense here in this space was dominated by patients facing chatbots that generated some scripted emphatics but they lack completely the rigorous structural discipline of any clinical therapy.
Today today we have a new paper today the world changes. So let's talk about the architectural engineering of our complex AI system, the multi- aent collaboration in detail and a new form of memory. The calculus of this particular memory we will call a CD memory. Now of course you might say hey it's a breakthrough we have a deconstruction now here of the clinical mind via a multi- aent system and the orcas hypothesized that a single eye model like a fable or go with an astroal failed here and should not play here the role of a therapist.
And at the end of the video, I will provide you an even deeper explanation than what you ought to hear by Jingua University referred to. And Jingua tells us here instead the therapist's brain must be kind of scientifically deconstructed into a model multi-agent system because you understand our AI systems are not intelligent enough to do this here like a human brain. So you have to reduce the complexity into multiple lower dimensional AI machines. and they found out you can't use a current proprietary model like Astra or Fable or whatever you have.
So they went and said okay so we have to train our own AI model for this particular task. And the end of the video I give you here the explanation that really goes into the detail why but for the moment we just accept in the study multi- aent CBD they use a QN 34B of course we're in China that splits now it's if you want logic here into five node compartments here neural mode structures and here we have them five agents simply now we have here five CPT agents for five distinct task number one is the boss the assessment the navigator [clears throat] that then the real interesting part is immediately number two the socratic questioning clarify what is the situation of the patient what evidence do we have what is happening socratic questioning typical here uh CBT then the cognitive restructuring identify the particular patterns of the patient is it here a scenario where the patient still sees everything in black and white zero or one all or nothing or is the patient already able to reframe this into a more normal on's perspective.
Third is then here no the next one is fourth the behavioral experiment. I will explain here this multi-step process to guide the the the the patient here out of its duck hole here out of the ground and say look this is the path forward for you and then the treatment monitoring the review progress and the reporting here done by an own agent and of course we have an interconnected shared session memory and this memory is of a very particular nature and we always have to ask why why this particular architecture why this particular methodology And if you code this system, you always have to be able to answer this question.
Otherwise, you have not understood what you're building. So by separating this function into five compartments, the author introduced what they call a hard state constraints because an LLM task solely with the Socratic questioning mathematically cannot skip ahead to offering a solution. So what happened in the old malls in the pure chatbot ms in the cheating ys in the hallucinating malls some of the mal that we're playing here therapy um was just says okay let's start the session and after 5 minutes more or less the eye provided you the answer this is absolute nonsense this is not clinical therapy this is just an eye going crazy this is just yeah you can play with this this is a chatbot this is nothing you can use in clinical therapy So they were forced here and going now with hospital with clinical data they said hey we have now to build a sequence where the eye cannot cheat the eye has to go in our sequence that we define in this linear sequence of agents full stop.
So they say the multi- aent architecture ensures you the absolute fidelity to evidence-based clinical protocols that are in place that are valid. Now, let's talk about if you're new to clinical psychology or psychology or psychotherapy and CBT. CBT is cognitive behavior therapy and evidence-based first-line treatment for let's say typical for depression. If humans really feel depressed, but this is not that you say, "Hey, today I feel a little bit down." No, this is really the clinical form of a depression.
Great. Now, it is a coincidence. I'm sitting here in Vienna Austria. you know viol is the birthplace of psychotherapy. So I'm doing here recording this video on an EI system that is operational in China in a hospital with Jingua University that maps out here the patients cognitive distortions and mathematically calculates now behavioral intervention based on the ideas under the theory of Freud Adler and Frank here and they developed the theory here in Vienna Austria.
So it is a global network. I'm just smiling here. So if you're not familiar, the terms for you would be sigma Alfred Victor Frank loapy. So let's do this in a simple example. You know, let's say we have a patient that says a simple example, hey, I just failed an exam and this proves my god, I'm so depressed. I'm a complete failure. Now let's play this simple nonclinical example just to show you the idea behind this network.
So assessment navigator agent number one the job is hey what should we work on next. So this agent tracks the current stage assess what information is here what information is still missing and selects the appropriate specialist in the next step. So for example it might decide to this agent hey we need to understand what this exam result means to this particular person because this is highly individual therapy now before we can come up with any alternative interpretation we really have to understand what's going on why is this so important to this patient what is behind the scene and now this assessment navigator agent routes the conversation to the questioning agent Yeah.
So let's see what's happening in the second step. Now we have the Socratic questioning agent that says hey what evidence supports here this conclusion. So we want kind of break up a little bit here the whatever the the patient feels and just have here a focused question that help the person examine their own interpretation. For example, this agent might ask, hey, what makes this one exam a measure of everything you can do?
Or can you think of something where you succeeded in doing this? You know, the positive aspects. So, you see this particular job is kind of a guided exploration, a little bit of logootherapy here to help to uncover assumption that the patient has rather than immediately supplying evidence and say, "Oh, yeah, now we know what it is." The third step is then a cognitive restructuring and uh this agent says hey can we form a more balanced interpretation what's going on deep inside the sort process of the human patient.
So this agent now this machine tries to identify possible thinking patterns you know any eye machine identifies only patterns and proposes a relevant way to reconsider it. Now, as I showed you, this particular candidate in this simple example is in an all orno thinking process. No. And now you want to bring this patient out of its hole that it dug itself into and says, "Listen, it's not zero or one. It is 0 to 10. So maybe you are just at not at a zero, but you at a two.
Come on." So you say, "Hey, between perfect at everything and a complete failure here at one, where should you place yourself when you consider all the evidence? Okay, you failed here on this Tuesday, but come on. This is just one single Tuesday. So you get the idea what this agent is trying to do. And then the next one is the behavioral experiment. This is not interesting. Now this the job of this agent is hey how could we test that belief through an action.
So it proposes now a small specific task whose outcome could provide new evidence for the eye machine to analyze here the thinking and the emotional status of the human patient. Yeah. So for example, you predict that asking a classmate for help will make them think that you are incapable. Could you ask one trusted classmate about one difficult question and actually record their actual response and maybe you see that they are helpful and they are trying to help you and they also struggling with some particular mathematical question.
So you want to come to a logical structure where you have now you go from prediction of what is happening in the sort process of this patient to an action and this action is now kind of a probe that explores now here is this what the patient thinks really happening and we have an observation here feedback from the environment and suddenly the patient sees okay if I ask for help hey that's easy people are yeah they're willing to help maybe half of the people are not willing to help but the other is so it's not so dark as I thought about it.
Yeah. In my simple example. So the job of this agent is design a concrete test and track the result whether it was completed the job. What is the feedback? What is the evidence we collected? Now if I test here my doomsday assumption here of the sort process of this particular patient and five is then the treatment monitoring agent. Yeah just records hey what happened? What should be recorded? What was important to see and now what is really important this is now an active interface to the memory element.
So careful this is not important. So this agent reviews now the session the full text of the session the apparent response to the intervention on any completed or unfinished task. However any sub explanation. So the clinical report might say hey the person expressed the less absolute interpretation. After the continuum exercise the classmate experiment is agreed but not yet completed review is review its outcome in the next session.
So you see job is evaluation documentation continuity and with this longitudinal memory enabled. This is now great for multi-session therapy because normally your session I don't know goes 20 weeks no or 25 weeks whatever you are wherever you are great so you need a very specific form of memory for this so let's talk about the memory now I think the really beautiful insight in the paper is here how it solves this longitudinal tracking over 24 weeks because a generic AI memory like rag or memory banks they will simply summarize your pass text they would try to compactify this information less token less uh memory great but this CPT technology this method does need all the information it doesn't care about just a summary it cares about the evolution of your cognitive pathology it really has to have all the data to come up with a therapy with an idea what is happening inside of this patient what is the sort process what are the cognitive h whatever So the oras engineered what they call a cognitive distortion memory.
This is if you like to analyze it a little bit deeper in the paper a specialized dual loop state machine. So at every session the monitoring agent doesn't just write a summary. It mathematically scores yes absolutely we have a formula. It scores here the patient mental distortions and the efficacy of the techniques used. So you don't want to have a summary. You want to have absolute calculated elements that are in the clinical procedure.
So you go with the established technologies in the hospital in the clinic wherever you are. So here you see this dual loop architecture. Yeah, it is classic. No planning, treatment, generate the treatment plan, the generation, the evalation and the planning and the update of the plan. We looping and another loop. So there's nothing specific now to this particular structure. Yeah, maybe that UCP knowledge base has level one is the real time collection.
Level two are the long-term records. Remember 24 weeks and level four the treatment plans and the treatment documentation and the current success of this therapy. So what the CDI system now especially here in the memory essentially creates a real time personalized technique response map of all the um logo particle techniques and response by this particular individual. Does it work? Does it not work? What are the problems?
Describe something. All this information is now in the memory. And this is something a human clinician should do but struggles because this takes hour and hours of preession preparation and so on. So before the next session the computes now here if you want from the all the possible options here some intervention priority using here this simple calculus. So this priority is simply a sum of the recency, the frequency and the severity of the of whatever goes on in the mental health status of this patient.
If you want to have here a simple example, what are the four things that this particular CD memory tracks now in real time on our simple example? So the cognitive distortion, what type of thinking trap did this human patient fall into? Everything is a catastrophe. I'm not able to do anything anymore. All or nothing thinking patterns. No. Then the trigger. What specifically caused this trap? An examination, break up with your girlfriend or whatever.
The alarm level, the severity is how bad is this distortion felt here by the human on a scale 1 to5 or whatever scale you have. And then the antidote. No. when now the human therapist or the eye therapist or whatever combination of both used a specific technique to help the human patient did it actually work the patient says okay I'm not now at level zero I'm already at level two yeah okay there's maybe a little light of hope at the end of the horizon here you have now a screenshot here this is here if you want to take an infograph here of the multi- aent CBD framework multi- intelligent body CBD system architecture.
And here we have our five agent that we went through the state transition, the evaluation, the soquality questioning, cognitive restructuring agent experiment and all the monitoring and the shared memory module. This is it. But there's one more thing that is really nice. And think about it, what we are missing, we are missing the training. We are missing how do we train this system to do this job? Because the classical AI system here at GPD6 solstra or fable or whatever you have OPUS is not able to do this.
So how do we we train this system? Why do the huge propriatory eye models fail? Think about the specific job the eye system has to perform with CPT. And I give you a hint. Think about Victor Frank on logootherapy. Do you understand? The job of the agent is not knowledge. It is not to provide you a solution, an abstract mathematical form or a logical deduction of something. The agent in a logo therapy in a clinical therapy has to understand the problem come up with an analysis and has to communicate with the human patient or provide some question for the human terapot.
So if you want the agent has to learn now how to talk to the human patient in a way that is maybe an established theory or clinical protocol from psychotherapy or whatever. So you have to kind of establish now trust and you have to guide here the self-reflection process of the human patient through some sublime question by the eye machine and more. Now I think this is really a challenge because I cannot imagine that I would trust an AI machine now because I understand this AI machine is just a pattern recognition machine and a pattern matching machine.
So there's nothing like empathy. There's nothing like trust. There's nothing like feeling in this machine but yeah it's replicating pattern and maybe it is replicating a pattern that might help me. I'm not sure that there or let's reformulate it. there is a theoretical possibility in statistics that maybe it will work with 1 percentage probability or 2%age success probability now whatever so this is really a leading and bleeding edge here of EI in clinical psychotherapy so we have to train this system to learn to ask the correct question given now the specific individual history of this particular patient.
Now this this is a challenge for you. You know the pattern the complexity of the pattern not just of the factual pattern but of the emotional pattern of the historic pattern of the educational pattern of the sort process pattern whatever this patient believes in has now to be identified classified find the pattern stored in the eye machine and train the eye machine with some training data. what to do in this particular case if you see this particular pattern in the mental health assessment of a human patient.
So therefore if you build this AI system, if you code this AI system, you have to clearly identify what is the job that you ask the eye machine to do. Here the job that you ask the eye machine to do is learn how to ask the patient the right question given your knowledge and given your logopetic approach your experience. You have read all the the text clinical textbooks on piootherapy. What is here your process? How to guide you this mental health development here for this human patient.
Now you have to train a model. You have to have training data. So to train now this new multi- agency model the authors had to invent a new way to generate training data. Now, of course, you might say, "Okay, this Chingua University uh cooperated with a hospital that was specialized on mental health and you got it." you know, so they have access to this papers and to this patient information and the the treatment plans and everything, but they decided to go here and have okay, we have 3,134 clinical reports where they say, okay, we cross out the names, we cross out the personal information, but we have to have here an understanding of the of the of the age, of the profession, of the problem, of the environment.
So 3,134 reports over a complete treatment and then they tried to extract what they called patient personas representative personas for a particular group and then they if you want initiated a dual role LLM simulation with an information asymmetry to generate the training data for the training of the QN3 EI model. So this is here my eye generated graphic where the idea is here what is the task of the eye to ask. The eye has to learn how to ask a human and the human is not willing to give you a logical answer because this human is in a deep emotional depression mode.
And you have as an AI machine to learn how to find the right question, ask the right question and respond with the next right question. So that this human opens up, it establishes some trust relation with the eye machine and is willing to be guided through this therapeutic process. So there's no help that the eye machine knows the answer. Okay, this patient is in a depression. Great. But you have actively to learn how to ask to help this patient.
So what is happening? You have an AI machine that has now fat 83,000 real therapy documentation and then you have an eye machine that is now the CPT counselor that learns now how to ask the question and you have a massive information asymmetry. You have 100% of the knowledge here on the simulated patient side and the CPD counselor has almost no idea who is this patient what is the problem with the patient and now the CPD counselor has to start to ask question and then it builds here the dialogues and from this dialogues those are the training data set so we can train our cubin 3 CPT machine.
So you put two agent in the same digital room. The patient agent here in orange has a secret persona but only revealing symptoms based on its emotional trigger. This is not a logical profile. This is an emotional profile and this AI machine has now to find the right emotional triggers so that this EI machine simulates a human patient that is now opening up based on the data of the 3000 real patient uh documentation. Yeah.
Why this methodology? You always have to answer the question, why did you choose this methodology? Now again, EI agent cheat. They cheat the hell out of it because it is cheaper. So this blindfolded approach forces now the eye machine to learn the contextual responsiveness to learn how to ask the right question. It cannot cheat. Yes, it has the knowledge in this situation. Maybe it is A, B or C multiple choice question.
But the learning process is you have to go through every listening step of the learning process. You have to really learn how to ask the right question. In this way, they generated 68,000 training samples. And you know exactly why we need 68,000. So everything about 10,000 you can use for supervised fine-tuning or reinforcement learning. And this is exactly why we need this data. Here you have it now in a screenshot here from the original paper from Chingua University.
And now it is interesting what happened next because the AR machine was still playing a psychopant and you had to apply DPO to break the cyto. So let's have a look. So once the order had this training data, they put the base molecule in 314B through some rigorous training process. No, but if forest they thought, hey, is it possible to just have a low rank adapter, an adapter on top of our LLM and they found this cheap shortcut like a Laura shortcut it failed and they said okay so let's go here with a clinical therapy.
Okay, it requires here the deep layer tensor weight modification of the ammo. We had a classical full supervised fine tuning SFT beautiful. They did supervised fine tuning and it was still not really clean this AI model. It was not really safe because an LLM naturally wants to be agreeable. Yeah, it has this psychopantic tendency as a chatbot. Oh, I'm so sorry you feel that way. You know, but you are absolutely right.
No, everything you want is correct. Everything you feel is correct. No, it is not able here to to oppose here something or be here a real clinical therapist. Yeah. So this psychopantic behavior is dangerous in therapy. Now because sometimes a therapist must challenge here the false reality here of a human patient. And to break this psychopantic habit of the supervised fine-tuning AI machine a QN3 the artist use DPO no or classical direct preference optimization we know this for 3 years classical training stuff and as it turns out with DPO yeah this is it this pushed TDI clinical professionalis score up by a massive margin so this was the way to do it results data here you see the session quality comparison to other competitive uh clinical systems and here in the last line you see our multi- aent CPT if you go for the mean it outperforms the other one now it is on you I mean think about this is a scale from 0 to three so this is a extreme reduced scale so therefore the step from 1.71 to 1.83 83 on mean performance is quite a significant step ablation.
If you go with no fine-tuning, no SFT, you have a mean of 155. If you go with a CPT specific supervised finetuning, but without the DPO, you add 1.92. And if you go with the full multi- aent CPT, this means supervised fine tuning and DPO, you have 2.10. Nice. And then what effect does the memory the specific cognitive distortion memory format play? And you'll see oh yeah it is absolutely important because if you go with our MA multi- aent CBT with no memory at all you have your longitudinal mean of 0.87 87 but if you go with the full session full CD cognitive distorted memory you jump up and remember the scale is 0 to three only to 2.29 29.
So absolutely this particular memory representation is highly beneficial here for the process of therapy. What are the insights? I think this MACBT is not trying to replace the human therapist here at the clinical environment. If you want this is a clinicianf facing AI exoskeleton optimized for therapy for mental health in clinical conditions. there. I think you can see it here in two different ways. Through the multi- aent logic, the separation into the five agents, it maps the structural geometry of the cognitive behavioral therapy.
And through the asymmetrical data simulation and data generation and data training for RL with DPO it internalizes here this professional boundaries and investigative patients required here in your real pure form of clinical psychology. So therefore absolute beautiful implementation of EI in a clinical psychology environment. I hope I could provide you some new information. Maybe you had a little bit of fun. new insight into this topic, how to coddi for this particular domain, for this sector.
I hope to see you in my next video.
The words are the caption track's own and nothing is reworded or re-transcribed. Paragraph breaks are placed between sentences so the text reads as prose.
Free tools for your own script: paste a draft and see where it stands before you record it.
Paste your draft and see where viewers are likely to drop off, with a rewrite for each weak line.
Paste the first 30 seconds of your own draft for a hook score and rewrites.
Check your draft against YouTube's advertiser-friendly guidelines before you record it.
Read this channel's public videos and transcripts, and download a writing brief for it.