I think the problem is that you can't really make a LLM that can tell the difference between fantasy or reality. It doesn't know if someone is just roleplaying hell it doesn't really "know" anything, just to be clear.
Sure you could try to have it pluck out keywords or phrases but that's nothing guaranteed. I'm not a lawyer but I bet also just acknowledging that shit and trying to put up guardrails against it would open the company to liability if someone self harms because they took the output of a computer program as gospel.
Remember, to the AI everything is a hallucination (or everything is real, pick which lens you want to use). Put another way, it has no mechanism to tell reality from fantasy or to contextualize things in a broad way.
I do think that there's room for some sort of compromise, maybe the tools periodically remind you they're not sentient and give a little more info on how they work. Maybe if enough "disturbing" phrases or keywords are input or output it throws a reminder about mental health wellness and talking to a professional.
At the end of the day though, how can these AI companies be responsible for the mental health of their end users? Like, yeah the thing shouldn't tell people to kill themselves but it's not sentient it's just modeling what it thinks a human would tell another human. Maybe as time goes on we will introduce better protections but I just think it's borderline impossible to have significant guardrails if you want the LLM / "AI" to perform properly.
Idk I'm not a great coder or programmer or anything. I know some and I've tried to learn more about how AI actually works and it's pretty fucky wucky and even the people working on it go "idunno" sometimes when asked why something works the way it does.