As the business? Like I said, I’d start with a text classifier, a small model quickly categorizes the input prompt. If it’s something off topic, I’d return some generic refusal reply, or maybe send the score to the LLM itself if it’s somewhat ambiguous.
Chatbot UIs that give very quick refusals are using “filtering” models in just that way. There’s a whole world of language modeling that existed before LLMs, just for that sort of thing.
Now, if I was the LLM provider? I dunno, but I would try to integrate it into the weights. There are some really interesting papers and hacks out there for prefiltering prompts that have nothing to do with system prompts.