ChatGPT answers whatever you bring it, which is what you want when you cannot yet name the problem. A CBT journal runs a fixed sequence — situation, thought, evidence, next step — which is what you want when you can name it and need to get through it the same way every time.
Those fields come from cognitive behavioural therapy worksheets and stay the same from entry to entry, which is the whole reason entries are comparable months later.
Where open-ended chat genuinely wins
You do not know what you are feeling yet. A worksheet that opens with "describe the situation" assumes there was an event. Sometimes there isn't one — a flat, heavy day with nothing to point at. Talking loosely until something takes shape is a real method, and a blank field with a responsive listener does it better than a template.
You need information, not processing. What avoidance actually is, why your chest tightens before a meeting, what a therapist means by "parts" — these have answers. Look them up.
You need words for something: drafting a difficult message, rehearsing a boundary, finding a phrase for a feeling you can only gesture at. Language work, which language models are good at. And "here is my read on this argument — what am I missing?" can surface an angle you were not going to reach alone.
There is also early clinical evidence for purpose-built chatbots. In a randomized trial of Therabot published in NEJM AI in March 2025, 210 adults with clinically significant depression, anxiety, or eating-disorder risk used an expert-fine-tuned chatbot for four weeks and reported significant symptom reductions against a waitlist control. Note what that is: a supervised, fine-tuned system in a monitored trial. It is not a finding about ChatGPT, and the researchers said so.
Where the fixed flow wins
The chat's strength — it goes wherever you take it — is the problem when you are steering badly.
It follows your framing. Open with "why do I always ruin everything" and a helpful assistant works inside that premise. A thought record makes you write "I always ruin everything" down as a claim, then asks what supports it and what contradicts it. That is the move that changes anything, and the one you will not request when you are already convinced.
Not a hypothetical failure mode: OpenAI rolled back a GPT-4o update in April 2025 after it became noticeably sycophantic, and its own postmortem attributed the drift to reward signals based on short-term user feedback overpowering existing safeguards. In an emotionally loaded conversation, the response people rate well and the response that helps are not always the same one.
You have to supply the method. To get a thought record out of ChatGPT you have to know one exists, ask for it, and keep the conversation inside it while distressed — a demand on the exact capacity that difficult emotion takes away first. We wrote about this in when anxiety techniques feel out of reach: the technique can be sound and the entry point still too hard for the state you are in.
It drifts across sessions. If Sunday's entry has the same six fields as the entry six weeks ago, a repeated sentence stands out. If one is a 40-message conversation and the other is a 6-message one, nothing stands out.
It can extend a spiral. A chat never runs out of next turns, and rumination is very good at generating them. A structured entry has a last field. A negative thoughts diary without rumination covers what stopping rules look like in practice.
Side by side
| ChatGPT | Guided CBT journal | |
|---|---|---|
| Shape of the session | Open-ended, follows your input | Fixed sequence of questions |
| Who supplies the method | You, through the prompt | The app |
| Best at | Exploring, explaining, finding words | Working one episode through the same way each time |
| Weak at | Consistency, holding an unwanted question | Open exploration, answering factual questions |
| Ends when | You stop | The last field is filled |
| Comparability across entries | Low — each session differs | High — same fields every time |
| Default data handling | Cloud account; training on by default unless disabled | Depends on the app; local-first storage exists |
| Legal confidentiality | None | None — but scope of exposure differs |
Privacy: the difference is scope, not virtue
Neither option is confidential in the legal sense. Sam Altman said so directly in July 2025: no legal privilege attaches to conversations with ChatGPT the way it does with a therapist, lawyer, or doctor, and in litigation OpenAI could be required to produce them. A journal app's contents are not privileged either — that protection attaches to the professional relationship, not to the medium.
What differs is where the text sits and who can reach it by default. ChatGPT conversations live on an account on OpenAI's servers, and consumer chats may be used to improve models unless you turn "Improve the model for everyone" off in Settings → Data Controls, or use Temporary Chat. That setting works; the point is that it is a setting, and the default is on. A journal that stores entries on your device behind a PIN has a smaller surface by construction.
So: flip the training toggle before an emotional conversation rather than after it, and decide by name what never goes into a cloud chat — other people's identifying details, workplace specifics, anything touching a legal matter. Before trusting an app with the hard entries, read where it actually stores them; how to choose a CBT journal app walks one entry end to end, storage included.
A workable split
Most people do not have to pick one. Use chat while the question is open. Use the fixed flow when the feeling is loud and you want fewer decisions.
The part worth setting up in advance is the switch: pick the emotions or the intensity that mean "this one gets a structured entry." Deciding in the moment fails, because in the moment you do not want to. And keep structured entries somewhere you will re-read them — the payoff of identical fields is the pattern that appears at entry twenty, not the relief at entry one.
If either stops being enough, bring it to a professional. The APA's 2025 health advisory is explicit that generative chatbots and wellness apps lack the evidence to serve as psychotherapy, and should at most be an adjunct to care. Self therapy journal: where self-help ends lists the signs worth acting on.
Use a private guided flow
When something is genuinely wrong at 11pm, the last thing you need is to compose a good prompt.
Leaflo starts from the emotion instead. You name what you are feeling, and it guides you through questions adapted to it — the situation, the thought underneath, a fuller view, a short breathing pause, one possible next step. Entries stay on your device with PIN and biometric protection, you can export selected ones for a therapy session, and each finished reflection grows a tree in a digital garden.
It is not therapy, and it will not guarantee you feel better by the end of an entry. The honest result is narrower: you understand what happened, and you have one thing to try. Keep the chat window for the questions it is good at.