you are viewing a single comment's thread
view the rest of the comments
[–] 10 points 7 months ago (3 children)

That would be my bet, LLMs really gravitate towards playing along and continuing whatever's already written. And Gemini especially has a 1M long context so it could be going back for a book's worth of text and reinforcing it up the wazoo.

That said, there is something really unhinged about Google's Gemma series even in short conversations and I see the big version is no better. Something's not quite right with their RLHF dataset.

  • source
  • parent
  • hideshow 3 child comments
  • [–] 1 point 7 months ago

    I have found Gemini the hardest to jailbreak tbh. I have been able to get Claude and CGPT to straight up give me a list of curses and slurs it isn't allowed to say, but Gemini will only do it if you say the words first.

  • source
  • parent