Not if it's done at the semantic level. If they instruct the model to only mention this brand of pasta, and give it a few arguments why it's the best, it will gladly incorporate that in its response and you won't have any way to detect that.
you are viewing a single comment's thread
view the rest of the comments
view the rest of the comments