We do see these issues in a lot of ways because llms train on bias. More polite answers get more verbose and angry answers get less verbose. There are a lot of these "hacks" that exploit this.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments