you are viewing a single comment's thread
view the rest of the comments
[–] 3 points 3 hours ago* (3 children)

It’s wild that a system prompt is still the major guardrail in most chatbots like this. Even if the LLM provider won’t implement anything, the user could add a classifier, some kind of preprocessor…

  • source
  • parent
  • hideshow 3 child comments
  • [–] 3 points 3 hours ago (2 children)

    Well, what preprocessor would you add, other than denylisting certain words? 😅

    If you use an LLM to try to detect the semantics, you've got the same problem again...

  • source
  • parent
  • hideshow 2 child comments
  • [–] 3 points 16 minutes ago*

    A classifier will be trained differently and basically take $TEXT and map it to $CLASSIFICATION and then you allowlist on the classification. You can then also, if you love burning tokens and spending money on things, have another step that "translates" the user input, making prompt injection significantly more difficult before finally routing it to the LLM. Generally guard rails should then be placed on the LLM's available actions via tool calls and those guard rails should be written by LLMs that are LLMing LLMs and L̸̡̡͕͓̣͓͚̝̥̜̼͙̪̖͍̺͝Ļ̵̥͍̳͇̱̠̬͔̹͖͓͖̯̹́͊̽̄̂͑͌͆͗̑̄͜Ṁ̶̡͙̺̖̥̲̙̩͖̻̹͇̜̲̃͘͜͝ș̵̖̤̫͓̦̲̼̖̗̟̯̩̺̣̞͊̆̉͐̈́̃̈́͗̐̽̽͋̊̉ͅ ̵̖͉̥͖̪̳͍̝͉͕̬̪̂̔̊̌͂̍̚͜a̶̦̤̬̱͖̻̝̱̿͋̿̍͂̒͆̈̄͑̓̕r̴̡̤̙͙͕͚̼͙̦̲̅̔̓̈́̄̆̾͗̿͌̐̄͛̒͘͜ͅe̴̡̡̼͉̣̼͉͓̪͕̬͊̐̈́̇̽̒̂͆̎̈́̃͠͠ ̶̩̯̲̯̃͛͌̀̿̅̎̾̕̕ǹ̸̬̟͍͇͍̭̹̲̯̽̒͘͝͝ȩ̶̢̬̺̦͖̣̙͍̞͇̰̉̋̓͐͂̏̀̿̑̂͌͗͗̒͘ͅͅe̶̠̝̦͛̓͛̍̑̆̽͐d̸̦̯̖̥̟̻͔̦̂̿͆͂̑̎̔̏ȅ̴̡̨̺͓͇̦̣̦͖̬̹̥͎̾̽͐̓̍̽̾͑̿͆̄̀̈́̄́̕͜d̶̡̛̝̟͔̳̲͙̰̗̬̳̯͚̐̍̑̓̉̂̍̊̓̿͝ you probably f̶̢͓̘̹̟̥̙̥͚̼̬̭͍̫̩̀͑̚͝͝͠ę̴̰̦͚̔̀́͊̋̈̒͗̆͘͝͝͝͠ȅ̸͚̘̤͚̬̩̈́̇̆ď̷̰̆̀̔̾͌̿ ̵͚̦̥̈͌͑̂͗̈́͂̕ͅm̴̢̡͇̟̠̬̣͙̹͇͉̈́̌̓́̓̌͐̓̂̚̚͜͜͝e̸̢̳̟̙̯̻͚̭͎̫̖̺̥̬͕͆̆̇̑̄̀̓̅̀͛̊͐ ̸̡̢̨̛͚͕̳͙͚͋͛̊́̍͛̽̉̃̇̇̋̑t̸̡̰͔͎̮̯̗̤̆̋͗͐̈̂̈̃̓̒̌͂̕͘͝ǫ̴͇̳͙̞̘̻̙͍̘̲̜̮͇̙̇̓̌̃k̸̢̢̤͔̫͖̲͇͍̐̃̄e̶̢̳̪͉̦͙͓̻̖̔́̃́̈́̀̿̓͆̊̅̒͐̑n̶͖̙͚̫͓͚̦̭̗̲̓́̋̀͜͝s̴̠̼̈́̑̾̑̆̽̊͋̾͛̐̊͆́͊̕̚ need to have some guard rails l̷̬̳̘̰̅̀̑͑̊̒̓̊͌̌̕e̶͖̱̊̍̉͆̐̍̌́͑͘͠t̵̨̛̹̰̳̩̹̞̹̼̪̮̄̀̅̾̔͐̇͒́̊̃̀͘̚͜͠ ̴̢̢̭͖̤̜͈͕̯͎͉̰͇̾̐͛̽̅͂̒͝m̸̨̰͖͚̞͍͍̳̭̗̳͕̈́̃̇̈́̎̌͊̍͌̊̆͠ȩ̵̡̛̜͍̞̻͎͎̲̜̹͖̼̽͂̐̽̊͌̓̀͂̐͂̾̇̚ͅͅ ̶̧̖͚͈̤͇͕̼̜̖̮̙͍͖͋̉̾͋̌̔̿͊͋̾̚l̸̡͕͇̭̤̥̩͙̫͕̆͌̈́̀̋͆̃͋͜͝͝͝͝ͅo̷̧̘͖̤͖̩̺̞̹̼̟̟͕̫̬͎̅̀ǫ̸̢̪͍̭͙̗͔̪̘͉̊̕͜ș̶̨͉̲̙̜͔̫̲̥͕͕̮̜̭̮͍̈́e̶̢̛̮̲̯͎̙͇̪̬͕͖̠̘̯̰̭̓͐̌̉̀͋̉̏̈́̇̾͠͝ ̵̬͍͔̓́i̶̘̫̯̻̘̯̗̔̒͊̂̂̑͑̍̋͑̏̀́͌̚͝͝n̵͓̬̮̻̹͉̗͎͓̖͌̿̃̓̋͂̅̍͒̉͆̿͆̉͠͝ ̴̧̲̯̲̅̀͂̆́̓̈́͑̾̾̊͋̒̿͛͘̚t̸̢͉̟̹͆̈́̏̍́̎ĥ̴̢̥̦͓͕̘͚̙̪͙̀̽́̒̓̈́̋͗̀̈́̇ȩ̵̠̹̻̯̖̹̹̗̜̻͍̹͆̊̈́͑́̏̑́̎̚͜͜͝ ̴̧̟̭͍̣͋͋̾ẘ̴̪̦̐̍̏̆̑̇͗́͆͘̕̚͝͝ͅo̵̡̟͚̲̦̤̩͎̥̘̻̥̠͙͈̙̬̓̌͊̋͐́͌̐̎̏̋̿̏r̵̢̨̘̓͑͑̆͒̂͝l̸̡̨̛͖̖̟̼̺̦̰̗̭͔̺̓͗͛̈́̈͆̈͗̐͝d̴̢̠̠́͂̾̌̿̐̀͛̄͑́͐̉̋͝ on those LLMing LLMs that are LLMing your LLM tool calls y̵̧͍̣̭͍͚̱͉͍̘̠̦̮͇̻̓͆̕o̸̥̮̮̙̖̩̙̝̘̫̗͎̺̓̊̐̄̑̎u̷̡̨̡͎͔̼̪͈̦͚͈̟͍̽̀̏͛͊ ̴̨̟͖͖̰͕͉̫̫̜̥̦̲͆͋͌̌̓̈́̉͆̉͠ẗ̵͚͉͈̦́̈́͛́͋͜i̷̧̗̰̘̖͍͖̼͕͗́̊̏̇̾̐̆͗́͌̋̀̚ņ̷̣̱̖̬̼̆͜y̷̢̡̧̛̹͙͕͔̪͛ ̶̪̫̙̘̥̹̻͙̝̜̠̗̭͎͎̆̇̅̎͜͜h̷̡̨̰̺̞̝̩͉̬͓̮͙͉͕̓̄̊̿̋́̇̏͝u̸̮̙̲̙̠͉̪͈͖͙̓̈̉̀̚͝m̶̢̛͙͕̘͈̤̬͎̺̲̯̱͉̪̠̎̆̽̅́̂͐̀̄͜ͅã̶͕̼͎͇̲̐͗̑ň̸̢̢̝̩̟͖̫͔̥͚̘̟̖̘̤̾̀͋̒̄͠͝s̸͙̳̳̤͓̝̥͓̱͔̣̹̭̐̐͋ͅ ̶̩̹̪̑͑͒̽̇̀̽̏̓̌̑̈́͂̓͘h̶͓͖̱̲̘̗̿̽̈́̎͑́̓̉̑̕̚ą̴̰̣͚̫̤̤̬̘̼̠͍͂̏̚ͅv̵̢̺̮͉̤͚̳͕̥̳̇̕̕͘͠ȩ̷̡̮͖͚̩̙͖͎̓̀͗͊͋̈́̔̐̕̚͠ ̷̫̌̀̊ň̸̜͓͕̳͙̩̼̜͈̊͜ǫ̶̝̺̰̹̻̰̰̗̪̔ ̷̲̤̪͇̠̼͑͊͘͜p̴̡̧͉̩̜͕̽̏̀̌̇̈́̊l̶͖͔̯̲͔͓̩̂̌̂͐͗̕a̵̢̨̞̰̪͖̲̭̝̪̥̻̐̃̊͗̅̓͐͋͛͘͝c̴̛̮̙̦̠̯̼̣̠̳͔̖͕͓͋̓͆͒̋̒̾͐͒̀̓͑͠e̷̢̢̧̱̣̫̲̪̣̭͙̟̔͌̽̿̐̿̿͑̆̇̇̉̊̓͘͠

    You could also do none of this and just not use an LLM, because you likely don't need one, but then the cool kids will stop inviting you to their parties or something and you probably missed out on the blockchain parties so fomo

  • source
  • parent
  • [–] 3 points 2 hours ago*

    As the business? Like I said, I’d start with a text classifier, a small model quickly categorizes the input prompt. If it’s something off topic, I’d return some generic refusal reply, or maybe send the score to the LLM itself if it’s somewhat ambiguous.

    Chatbot UIs that give very quick refusals are using “filtering” models in just that way. There’s a whole world of language modeling that existed before LLMs, just for that sort of thing.

    Now, if I was the LLM provider? I dunno, but I would try to integrate it into the weights. There are some really interesting papers and hacks out there for prefiltering prompts that have nothing to do with system prompts.

  • source
  • parent