I've tried giving screenshots of phishing emails to a local Qwen instance and so far it always correctly detected it as scam, even points out the exact elements that it based its judgement on. Sending screenshots to it ad-hoc isn't too scalable for family and friends. I'd like to be able to either forward emails for screening, or perhaps have it screen everything from a mailbox.

Has anyone done anything like this? Is there anything self-hostable that does this?

all 13 comments

sorted by: hot top controversial new old
[–] 3 points 1 hour ago

I would recommend using a classifier model instead of a text generation one.

As for the way to do it a simple program that connects to you mailbox via imap and calls the vllm or whatever api you are using and then moves it to spam could work?

  • source
  • [–] 8 points 5 hours ago

    You definitely don’t have to go full LLM for this. There are a wide variety of local spam filtering tools available, depending on what email application you use. If you really want to use modern ML, I’d go with something like Laya as it is tunable and runs locally.

  • source
  • [–] 23 points 8 hours ago (3 children)

    I'd be pretty concerned about prompt injection risks with feeding a LLM unsanitized data. You definitely need a good harness around it..

  • source
  • hideshow 3 child comments
  • [–] [S] 10 points 8 hours ago (2 children)

    Good point. It'll have to have no access to the internet or anything local outside of its container. Just text in, text out.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 3 points 4 hours ago

    You could (and probably should) use a system-one style inference system for spam classification. Much cheaper and the structured output means it's impossible to go rogue and curl some malware or whatever. It can absolutely misclassify but its output is programmatically structured and just ranks a pre-selected set of output tokens.

    In your case that's

    Spam

    Not_spam

  • source
  • parent
  • [–] 8 points 8 hours ago

    Yeah if it's just a basic input with a function call for spam or not spam the risk is low. What's the worst case outcome, it tricks it into saying no it isn't a scam and you have to delete it manually? Hahaha

  • source
  • parent
  • [–] 5 points 6 hours ago (1 child)

    Just be aware that those phishers are now also scanning their phishing emails to see how they can get better too! The wheels on the bus go round and round! Round and Round! Round and Round!

  • source
  • hideshow 1 child comment
  • [–] 10 points 8 hours ago* (last edited 8 hours ago)

    LLM plugins should be available in standard spamfilters like Rspamd. Some mail servers like Stalwart have it as well. You pick a prompt and provide it with an OpenAI-compatible endpoint and it'll ask the AI for every mail. Can be a local model.

    Not sure if it's in the Email clients as well. I found a few Thunderbird addons, but that was just a quick Google search.

  • source
  • [–] 9 points 8 hours ago

    If your email service offers an MCP connection (Google does), you can have your AI connect and examine emails and manipulate them. You can hook that up to qwen using maybe lmstudio. I haven't tried it self-hosted, but it sounds pretty straightforward.

  • source
  • [–] 5 points 8 hours ago (1 child)

    Probably better to ask this in an ai community. However fwiw there are a few skills for generic imap checking. I've set up a local openclaw gateway and model with an automated task to check my email in the morning and basically prod me into replying. The same concept should be applicable to spam detection too.

  • source
  • hideshow 1 child comment
  • [–] 2 points 2 hours ago

    OP if you don't mind running this through a Hermes / OpenClaw harness this is probably the easiest way with pre built email connectors. Also if you don't mind spending a few bucks a month DeepSeek v4 will blow out anything you run locally for pennies (or gpt luna as of last wed.)

    It can also set up its own custom connector if needed and run as needed via CRON.

  • source
  • parent
  • [+] 1 point 6 hours ago