You could (and probably should) use a system-one style inference system for spam classification. Much cheaper and the structured output means it's impossible to go rogue and curl some malware or whatever. It can absolutely misclassify but its output is programmatically structured and just ranks a pre-selected set of output tokens.
In your case that's
Spam
Not_spam