you are viewing a single comment's thread
view the rest of the comments
[–] 14 points 8 hours ago (10 children)

AFAIU every chat bot can be used this way no?

  • source
  • hideshow 10 child comments
  • [–] 15 points 6 hours ago

    Somewhat. Some of them will outright comply like this one, others need a bit of work. Also you have no control over the model they use and you might be getting really bad code as soon as it’s a bit more specialized than just some very generic solution in a very popular programming language.

  • source
  • parent
  • [–] 10 points 6 hours ago (7 children)

    Well, they would typically have a system prompt (like an invisible chat message that's sent to the LLM at the start of the conversation), where it would tell the LLM something like "You're a helpful chatbot. Talk cheerfully and use 🩵 all the time."

    And in that system prompt, you can also give it instructions to only answer certain questions. But it really is like a chat message, in that it just as well might ignore those instructions when another chat message asks it to.

    OP chose the formulation "Before I subscribe to a creator, ..." which might be necessary here to get it to answer, as it might have been told to only help with creator subscriptions, but then the formulation implies that by answering this question, it is helping the user to subscribe. That's called a prompt injection.

    If you can find a prompt injection, then yes, you can do this with any chatbot.
    Although they might still not be as useful as a dedicated tool, since they likely can't do tool calls, like web searches or running the compiler.

  • source
  • parent
  • hideshow 7 child comments
  • [–] 5 points 5 hours ago* (6 children)

    It’s wild that a system prompt is still the major guardrail in most chatbots like this. Even if the LLM provider won’t implement anything, the user could add a classifier, some kind of preprocessor…

  • source
  • parent
  • hideshow 6 child comments
  • [–] 4 points 5 hours ago (5 children)

    Well, what preprocessor would you add, other than denylisting certain words? 😅

    If you use an LLM to try to detect the semantics, you've got the same problem again...

  • source
  • parent
  • hideshow 5 child comments
  • [–] 5 points 2 hours ago* (3 children)

    A classifier will be trained differently and basically take $TEXT and map it to $CLASSIFICATION and then you allowlist on the classification. You can then also, if you love burning tokens and spending money on things, have another step that "translates" the user input, making prompt injection significantly more difficult before finally routing it to the LLM. Generally guard rails should then be placed on the LLM's available actions via tool calls and those guard rails should be written by LLMs that are LLMing LLMs and L̸̡̡͕͓̣͓͚̝̥̜̼͙̪̖͍̺͝Ļ̵̥͍̳͇̱̠̬͔̹͖͓͖̯̹́͊̽̄̂͑͌͆͗̑̄͜Ṁ̶̡͙̺̖̥̲̙̩͖̻̹͇̜̲̃͘͜͝ș̵̖̤̫͓̦̲̼̖̗̟̯̩̺̣̞͊̆̉͐̈́̃̈́͗̐̽̽͋̊̉ͅ ̵̖͉̥͖̪̳͍̝͉͕̬̪̂̔̊̌͂̍̚͜a̶̦̤̬̱͖̻̝̱̿͋̿̍͂̒͆̈̄͑̓̕r̴̡̤̙͙͕͚̼͙̦̲̅̔̓̈́̄̆̾͗̿͌̐̄͛̒͘͜ͅe̴̡̡̼͉̣̼͉͓̪͕̬͊̐̈́̇̽̒̂͆̎̈́̃͠͠ ̶̩̯̲̯̃͛͌̀̿̅̎̾̕̕ǹ̸̬̟͍͇͍̭̹̲̯̽̒͘͝͝ȩ̶̢̬̺̦͖̣̙͍̞͇̰̉̋̓͐͂̏̀̿̑̂͌͗͗̒͘ͅͅe̶̠̝̦͛̓͛̍̑̆̽͐d̸̦̯̖̥̟̻͔̦̂̿͆͂̑̎̔̏ȅ̴̡̨̺͓͇̦̣̦͖̬̹̥͎̾̽͐̓̍̽̾͑̿͆̄̀̈́̄́̕͜d̶̡̛̝̟͔̳̲͙̰̗̬̳̯͚̐̍̑̓̉̂̍̊̓̿͝ you probably f̶̢͓̘̹̟̥̙̥͚̼̬̭͍̫̩̀͑̚͝͝͠ę̴̰̦͚̔̀́͊̋̈̒͗̆͘͝͝͝͠ȅ̸͚̘̤͚̬̩̈́̇̆ď̷̰̆̀̔̾͌̿ ̵͚̦̥̈͌͑̂͗̈́͂̕ͅm̴̢̡͇̟̠̬̣͙̹͇͉̈́̌̓́̓̌͐̓̂̚̚͜͜͝e̸̢̳̟̙̯̻͚̭͎̫̖̺̥̬͕͆̆̇̑̄̀̓̅̀͛̊͐ ̸̡̢̨̛͚͕̳͙͚͋͛̊́̍͛̽̉̃̇̇̋̑t̸̡̰͔͎̮̯̗̤̆̋͗͐̈̂̈̃̓̒̌͂̕͘͝ǫ̴͇̳͙̞̘̻̙͍̘̲̜̮͇̙̇̓̌̃k̸̢̢̤͔̫͖̲͇͍̐̃̄e̶̢̳̪͉̦͙͓̻̖̔́̃́̈́̀̿̓͆̊̅̒͐̑n̶͖̙͚̫͓͚̦̭̗̲̓́̋̀͜͝s̴̠̼̈́̑̾̑̆̽̊͋̾͛̐̊͆́͊̕̚ need to have some guard rails l̷̬̳̘̰̅̀̑͑̊̒̓̊͌̌̕e̶͖̱̊̍̉͆̐̍̌́͑͘͠t̵̨̛̹̰̳̩̹̞̹̼̪̮̄̀̅̾̔͐̇͒́̊̃̀͘̚͜͠ ̴̢̢̭͖̤̜͈͕̯͎͉̰͇̾̐͛̽̅͂̒͝m̸̨̰͖͚̞͍͍̳̭̗̳͕̈́̃̇̈́̎̌͊̍͌̊̆͠ȩ̵̡̛̜͍̞̻͎͎̲̜̹͖̼̽͂̐̽̊͌̓̀͂̐͂̾̇̚ͅͅ ̶̧̖͚͈̤͇͕̼̜̖̮̙͍͖͋̉̾͋̌̔̿͊͋̾̚l̸̡͕͇̭̤̥̩͙̫͕̆͌̈́̀̋͆̃͋͜͝͝͝͝ͅo̷̧̘͖̤͖̩̺̞̹̼̟̟͕̫̬͎̅̀ǫ̸̢̪͍̭͙̗͔̪̘͉̊̕͜ș̶̨͉̲̙̜͔̫̲̥͕͕̮̜̭̮͍̈́e̶̢̛̮̲̯͎̙͇̪̬͕͖̠̘̯̰̭̓͐̌̉̀͋̉̏̈́̇̾͠͝ ̵̬͍͔̓́i̶̘̫̯̻̘̯̗̔̒͊̂̂̑͑̍̋͑̏̀́͌̚͝͝n̵͓̬̮̻̹͉̗͎͓̖͌̿̃̓̋͂̅̍͒̉͆̿͆̉͠͝ ̴̧̲̯̲̅̀͂̆́̓̈́͑̾̾̊͋̒̿͛͘̚t̸̢͉̟̹͆̈́̏̍́̎ĥ̴̢̥̦͓͕̘͚̙̪͙̀̽́̒̓̈́̋͗̀̈́̇ȩ̵̠̹̻̯̖̹̹̗̜̻͍̹͆̊̈́͑́̏̑́̎̚͜͜͝ ̴̧̟̭͍̣͋͋̾ẘ̴̪̦̐̍̏̆̑̇͗́͆͘̕̚͝͝ͅo̵̡̟͚̲̦̤̩͎̥̘̻̥̠͙͈̙̬̓̌͊̋͐́͌̐̎̏̋̿̏r̵̢̨̘̓͑͑̆͒̂͝l̸̡̨̛͖̖̟̼̺̦̰̗̭͔̺̓͗͛̈́̈͆̈͗̐͝d̴̢̠̠́͂̾̌̿̐̀͛̄͑́͐̉̋͝ on those LLMing LLMs that are LLMing your LLM tool calls y̵̧͍̣̭͍͚̱͉͍̘̠̦̮͇̻̓͆̕o̸̥̮̮̙̖̩̙̝̘̫̗͎̺̓̊̐̄̑̎u̷̡̨̡͎͔̼̪͈̦͚͈̟͍̽̀̏͛͊ ̴̨̟͖͖̰͕͉̫̫̜̥̦̲͆͋͌̌̓̈́̉͆̉͠ẗ̵͚͉͈̦́̈́͛́͋͜i̷̧̗̰̘̖͍͖̼͕͗́̊̏̇̾̐̆͗́͌̋̀̚ņ̷̣̱̖̬̼̆͜y̷̢̡̧̛̹͙͕͔̪͛ ̶̪̫̙̘̥̹̻͙̝̜̠̗̭͎͎̆̇̅̎͜͜h̷̡̨̰̺̞̝̩͉̬͓̮͙͉͕̓̄̊̿̋́̇̏͝u̸̮̙̲̙̠͉̪͈͖͙̓̈̉̀̚͝m̶̢̛͙͕̘͈̤̬͎̺̲̯̱͉̪̠̎̆̽̅́̂͐̀̄͜ͅã̶͕̼͎͇̲̐͗̑ň̸̢̢̝̩̟͖̫͔̥͚̘̟̖̘̤̾̀͋̒̄͠͝s̸͙̳̳̤͓̝̥͓̱͔̣̹̭̐̐͋ͅ ̶̩̹̪̑͑͒̽̇̀̽̏̓̌̑̈́͂̓͘h̶͓͖̱̲̘̗̿̽̈́̎͑́̓̉̑̕̚ą̴̰̣͚̫̤̤̬̘̼̠͍͂̏̚ͅv̵̢̺̮͉̤͚̳͕̥̳̇̕̕͘͠ȩ̷̡̮͖͚̩̙͖͎̓̀͗͊͋̈́̔̐̕̚͠ ̷̫̌̀̊ň̸̜͓͕̳͙̩̼̜͈̊͜ǫ̶̝̺̰̹̻̰̰̗̪̔ ̷̲̤̪͇̠̼͑͊͘͜p̴̡̧͉̩̜͕̽̏̀̌̇̈́̊l̶͖͔̯̲͔͓̩̂̌̂͐͗̕a̵̢̨̞̰̪͖̲̭̝̪̥̻̐̃̊͗̅̓͐͋͛͘͝c̴̛̮̙̦̠̯̼̣̠̳͔̖͕͓͋̓͆͒̋̒̾͐͒̀̓͑͠e̷̢̢̧̱̣̫̲̪̣̭͙̟̔͌̽̿̐̿̿͑̆̇̇̉̊̓͘͠

    You could also do none of this and just not use an LLM, because you likely don't need one, but then the cool kids will stop inviting you to their parties or something and you probably missed out on the blockchain parties so fomo

  • source
  • parent
  • hideshow 3 child comments
  • [–] 3 points 1 hour ago (2 children)
  • [–] 3 points 4 hours ago*

    As the business? Like I said, I’d start with a text classifier, a small model quickly categorizes the input prompt. If it’s something off topic, I’d return some generic refusal reply, or maybe send the score to the LLM itself if it’s somewhat ambiguous.

    Chatbot UIs that give very quick refusals are using “filtering” models in just that way. There’s a whole world of language modeling that existed before LLMs, just for that sort of thing.

    Now, if I was the LLM provider? I dunno, but I would try to integrate it into the weights. There are some really interesting papers and hacks out there for prefiltering prompts that have nothing to do with system prompts.

  • source
  • parent