Also /c/modabuse. Also /c/yepowertrippinbastards. Also lemmy.ml/c/worldnews, comrade, ACAB, and 552 individual accounts across 67 instances, about half of them on lemmy.world.
None of this is in the source code. It's downloaded at runtime from a file nobody has ever looked at.
If you're just tuning in
Tesseract is a third-party web frontend for Lemmy, maintained by asimons04 and licensed AGPL-3.0. Admins deploy it on their own servers alongside or instead of lemmy-ui, and there are public instances of it people use to browse Lemmy generally. If you've used a Lemmy site that didn't look like stock Lemmy, there's a fair chance it was this.
Last week db0 posted a PSA: Tesseract contains a blacklist of instance domains compiled directly into the application. 32 of them. Admins can't see it, can't configure it, and aren't told it's there. Connect to a listed instance and the app tells you it's "incompatible," which is not true.
I went through the code to see how that was implemented. The hardcoded list turns out to be the small half of the system.
There's a second filter policy fetched over HTTP every time the app loads. It isn't in the git repository. It's unauthenticated and world-readable, so anyone can pull it. Right now it carries 552 user accounts, 2,275 username patterns, 54 instances, 97 communities, 289 keyword patterns and 351 domains, with every category set to hide matches rather than flag them. Not collapsed behind a click. Simply absent, with no indication anything was removed.
Verify all of it in ten seconds
curl -s https://tesseract.dubvee.org/tesseract/api/system/policy \
| base64 -d | gunzip > policy.json
That's the live policy, base64-wrapped gzip, 111KB of JSON when it unpacks. There's a stale fallback copy at /data/policy.dat as well.
It filters criticism of moderators
lemmy.sdf.org/c/modabuse— listedlemmy.dbzer0.com/c/yepowertrippinbastards— listedlemmy.dbzer0.com/c/YPTBcirclejerk— listed- community regex
power ?tripping? - keyword
censoring me
Call the rest of it whatever you like. This part is not spam defence.
It filters words
The 32 community name patterns include Communis(t|m), Conservativ(e|es|ism), Leftis(t|m), Libertarian(ism)?, ^Green Part(y|ies), Zionis(t|m), (Police|Cops), guillotine and billionaire.
Keywords include comrade, ACAB, neoliberal, proletaria(n|t) and death to.
Filtered communities on instances that aren't blocked: lemmy.ml/c/worldnews, lemmy.today/c/news, lemmy.ca/c/politicalnewscanada, lemmy.ca/c/usa, infosec.pub/c/strategic_unions.
The 552 users aren't bots
67 instances. 272 on lemmy.world alone, 40 on sh.itjust.works, 19 on lemmy.ca, and 28 instances contributing exactly one person each.
355 of the 552 usernames are plain alphabetic, twelve characters or under, median length eight. Only 36 look like spam registrations. A bot list looks like the opposite of that.
Seven of them aren't even Lemmy. There are Mastodon and Friendica accounts in there: people who have never used Lemmy, hidden by a Lemmy frontend, with no possible way of finding out.
I have the list and I'm not posting it. Most of these are ordinary people who got pattern-matched, and 552 names on this comm is a harassment target inside an hour. Run the command above and grep for yourself.
And it lies about it
When the instance block fires you get: "Incompatible Instance. Not Supported. $instance is not compatible with Tesseract."
Nothing is incompatible. It's a policy decision dressed as an API error, and it's what had db0 chasing a version mismatch that never existed.
For the hidden users, communities and keywords, you get no message at all.
Admins can't switch it off
Tesseract has env vars for PUBLIC_DOMAIN_BLACKLIST, PUBLIC_FAKE_NEWS_BLACKLIST and the shortener lists. There is none for either blocklist. enableToxicMode bypasses the other filters and explicitly not this one.
Self-host it and you cannot disable this, nothing in your config admits it exists, and the contents can change without you pulling a commit.
Before someone says it
A lot of that domain list is real spam defence. It filters conservatism as well as communism. "It targets the left" doesn't survive the data and I'm not going to pretend it does.
The problem is that spam filtering and political editorial got welded into one undocumented, remotely-updatable blob, shipped hidden, to admins who've never read it and users who don't know it's there. The spam work is what makes the rest unauditable: "it's a spam list" answers every individual question and none of the whole.
And /c/modabuse is not spam.
Asks
- Publish the runtime policy in the repo, or kill the endpoint.
- Stop reporting a policy block as a technical incompatibility.
- Tell users when something's been hidden. One line.
- Give operators an off switch, like every other blacklist in the codebase has.
It's AGPL-3.0 and db0 already forked it. That's the licence working as designed. But forking isn't disclosure, and the admins who need this are precisely the ones with no reason to go looking.
If you run Tesseract, you are relaying a 111KB moderation policy you have never read, under your instance's name, to users who don't know it exists.
Full contents of every list, unedited, in the comments.
Like I would have zero problems with developers that do this, if they could come out clearly and document/announce: "I don't like these accounts or words, if you do then this is what you do to change it or shut it off."
Stuff like Lemmy's word list, Piefed's default blocklist, you can turn off or change the words. Even if I think the daily vote quota is a stupid default on setting, if you're a Piefed admin and don't like that, set the limit to 10000, 99999999 or whatever. I usually have tons and tons of sympathy for FOSS devs, take their side and give the benefit of doubt in grey areas. They work for free and share it for free so I don't expect perfection or how I would want things to be.
Undocumented shit like this, intentionally hidden is way past the line, and I have no patience for it whatsoever. You never know when this could be somehow used as an attack vector.
Lemmy’s slur filter is disabled by default. It’s opt-in at the admin level, and most instances don’t opt for it, including your own I believe.
I don’t know why this molehill keeps being made into a mountain.
I'd add on that stuff like this should NOT be the default (unless that is specifically what you advertise your software for, and even then you should give a clear option to disable or change it), because users (even admins using the software for their instance) will generally not configure things much.
What should be done is have it as an opt-in, or prompt the user to pick an option.
If one believes their blocklist/filter is in good faith and beneficial to others and to the world in general, then one shouldn't feel a need to force others to use it, only inform them that it exists.
Given that the list very clearly targets queer users for being queer (entire lemmy.blahaj.zone instance blocked), you are being far too charitable with an obviously bad faith actor.
If they come out and say they don't like blahaj.zone for bigoted reasons, let them, don't support their project, done. I'll defend blahaj.zone's strictly enforced any pronoun policy as much as a dev who clearly chooses to include or exclude certain instances by default, as long as the intentions are clear, not hidden away and easily rectifiable for your needs. What I am saying is a breach of trust is worse than bigotry, which itself is worse than mere ideological disagreements. A < B < C.
You can read into how you like the presence of certain instances/words/users on the blocklist, but putting BZ on a blocklist = bigot seems to me like a jump to conclusion.
Like I said, I afford FOSS devs lots of grace and benefit of doubt, that's just me, if the issue was solely over the choice of instances being blocked by default, I'd excuse it, but that isn't the main issue. Call me too charitable, sure, it's in my nature and you don't have to be me.
The issue is the blocklist was hidden away in base64, not documented anywhere, while constantly being updated, without any justification (like pulling from a public spamlist), and was secretly affecting every server's Tesseract frontend. It was the trust that the so-called "Toxic mode" could remove all filtering, but it didn't, and the trust that connection errors were simply configuration problems and not intentionally obfuscated filtering routines, these formed the trust that was broken.
I will, and I consider you willfully ignorant for not doing so. Consider the probability distribution of likely motivations for drawing up a secret blacklist that specifically targets queer communities and anyone to the left of "let's hunt immigrants for sport". The majority of that distribution is gonna be bigotry.
Interesting choice to give the developer who already violated user trust by imposing a secret blacklist the benefit of the doubt on their bigotry. Seems like that would lose them the benefit of the doubt if you were being intellectually honest.
Requiring users to enable "toxic mode" in order to access queer spaces isn't evidence of bigotry to you?
I never did that, sorry I kind of suggested I did in my last reply, when I meant in general to devs that don’t break peoples' trust. In a vacuum, a configurable blocklist with biased defaults is acceptable, even if I wouldn't set it myself that way. Their work, their choice. However, it is specifically the concealment of the biased defaults that is the breach of trust which calls into question the dev's motivations.
Alone, no, that would be mere speculation, if combined with the shady practice or if you have antiLGBTQ quotes then yes, I could say that fact potentially supports that conclusion. Consider that blahaj.zone is not the only LBGTQ+ safe space on Lemmy (see beehaw !lgbtq_plus@beehaw.org). Any admin or dev could have beef with Ada, not with queer folks, that could be motivation to block blahaj.zone. I never gave the tesseract dev that grace, because that blocklist with blahaj was never publicized, until the clandestine filter list was discovered and decoded.
That's exactly where I'm at. I don't align with a lot of the choices Rimu makes for Piefed, but I am 100% for him designing a tool that works for him, even when the defaults he chooses don't align with my preferences. As long as I have the option to disable/change them as an admin, then I'd rather see a passionate dev creating tools that they want to use than a burned out dev creating something their heart isn't invested in.
Unfortunately, tesseract isn't that...
One note/minor correction, it is in the git repo: https://github.com/db0/tesseract/commits/main/static/data/policy.dat
You can see when asimons04 changed the static base64 filter .dat file. The changes happen in completely unrelated or nondescript commits. Looking at the commit when it was implemented, it does say filter, but the error messages and code comments creating the folder indicate "cache" folder creation, when the variable name and function is for the policy, which might technically be correct, is misleading and is quite suspect to me.
Yeah, I really wanted to believe this was just a personal filter list that got included accidentally, but there's no fucking shot. The way it's implemented and intentionally obfuscated means there's no way in hell this isn't entirely intentional.