I was messing around with duck.ai because it doesn't need an account and I like to see how censored western AIs are. Specifically around the Iran war.

I was going to ask it to read Aljazeera to get their coverage of the war and then compare what it summarizes to the same thing from other news sites. To see if it changes what it says based on which news org I have it pull the info from. But then I ran into an interesting hiccup.

I tried this with multiple models and as soon as you specify to search for Aljazeera + specific news event it will get a 404 error. And the same for CNN.

I thought this might be the sites themselves blocking AI crawlers. So I tried it in deepseek. It had no issue pulling both Aljazeera and CNN. I also researched exa (the search plugin AIs use) and it does advertise a feature where you can block certain URLs from showing up.

It does seem possible to force it to see them. I was able to get it to access Aljazeera.com on the backend and read from their own documents to find articles but in the websearch tool itself there is no result coming up from either CNN or Aljazeera. Which is odd right?

I did some research. I thought maybe it had something to do with the Zionists. This smells of them. And the founder of duckduckgo does have ties to the founder of an "Israeli" tech firm (they seem to be cousins) and even said in 2013, I think, that he wanted to visit "Israel" in an interview he did with the Times of "Israel". So this might explain the Aljazeera part. It is banned in "Israel" after all. But why CNN? That's an American company. It tends to follow the state departments messaging.

I tried some more sites. Trying to think of any that might be getting blocked on purpose.

RT - not working Aljazeera - not working CNN - not working NBC - working ABC - working FOX - working QQ news - working CGTN - not working

After trying these I tried AP next. Then it told me "duckduckgo is temporarily unavailable". I refreshed. Same thing. It was acting like the site was not working. Until i changed my IP and deleted browser data. Then suddenly it worked again. Interesting huh? And AP is also not working btw.

So it blocks certain news sites, and if you keep trying to access them too often it blocks you from using the service by pretending to be offline I guess?

Would love to see if anyone can replicate this behavior. If you do it be careful not to let it trick you. What it will do is you'll ask for coverage from a specific site and it will find that site has nothing (404 error) and then it'll just make things up from other sites, and act like it did what you asked. You have to specifically tell it not to include results from any website other than the one you requested. And watch the thinking so you can see as it tries the searches and they fail.

all 17 comments

sorted by: hot top controversial new old
[–] 30 points 2 months ago

DDG did block Russian sites a couple years ago. The CEO even made tweets about how based they were for doing it.

  • source
  • [–] 21 points 2 months ago

    I haven’t regularly used DuckDuckGo in a long time but I did notice that it was suspiciously prioritizing conservative sources and omitted search results that even Google was still willing to show. The founder being a Herzlian explains a lot. I always found it odd that a search engine (supposedly) against tracking could afford so much advertising and I see now that DuckDuckGo was a scam all along.

    I miss Scroogle.

  • source
  • [–] 11 points 2 months ago (2 children)

    Might give this a try later. Always good to see what sort of censorship is being implemented thought these AI services.

  • source
  • hideshow 2 child comments
  • [–] 3 points 2 months ago (1 child)

    So I sort of gave this a try, but ran out of messages. What I tried was asking about a different topic. I asked about a comparison of articles on trade with China using Al Jazeera, BBC, CBC, and FOX. I actually go results from all. Took me a bit to really tune what I was asking for though. When I switched topic to the Iran War that's where it told me that it couldn't get reliable results from Al Jazeera and CBC. So I wonder if it might have more to do with the topic rather than the source you're trying to pull from. What I was going to try next was change 'Iran War' to 'Iran conflict' instead to see if it'd get me anything. I was checking Al Jazeera's hompage and they refer to it as the 'US-Israel war on Iran'.

  • source
  • parent
  • hideshow 1 child comment
  • [–] [S] 2 points 2 months ago

    Did you verify that it was actually pulling from the right place? When I asked it would sometimes just pretend. Like when a site didn't work it would use a different site and just pretend it had used the one you asked. But you can tell if you check the actual like thinking it does.

  • source
  • parent
  • [–] 11 points 2 months ago*

    ChudChudCringe

  • source
  • [–] 11 points 2 months ago* (2 children)

    There's a fuckton of legal back-and-forthery between news sites and big search engines over scraping info, I believe there are laws around it? I'd be quicker to wonder if this is some sort of licensing dealio.

  • source
  • hideshow 2 child comments
  • [–] [S] 3 points 2 months ago (1 child)

    Well this isn't a thing in the DDG main results. It's just with their AIs. It's easier to hide in an AI since it just tends to summarize info without always being forward about what sites it got it from. Like if I ask the same AI, "tell me about the current iran war" right now: It pulls from 4 sites.

    Notice "unitedagainstnucleariran.com" Not exactly an unbiased source huh? And it reads this one but then when it does the summary...

    Would you look at that it only mentioned Britannica and Wikipedia. And in it's summary it is pretty vague about where those thousands dying were or who killed them. I'm sure that has nothing to do with those other sources it didn't tell us about and how they might phrase things.

    Those are also simply not what would show up in a normal search. If you searched "Iran War News" in what world are those the first 4 links lol?

  • source
  • parent
  • hideshow 1 child comment
  • [–] [S] 3 points 2 months ago

    To confirm it I did a normal DDG search with the exact same syntax the AI uses. The first 2, Wikipedia, and Britiannica, are entirely normal. Those are the ones it doesn't hide. But the 2nd 2? Not even on the first page at all. I have no idea where it got those from.

  • source
  • parent
  • [–] 10 points 2 months ago* (4 children)

    it could be that Al Jazeera is blocking high-traffic or abusive crawlers... i just found out a couple weeks ago that Meta's AI crawler had been hammering our little site so hard for days that it was essentially DOS'd for upwards of 5 - 10 minutes at a time, over and over again day and night, just spamming invalid urls trying to learn every combination of possible pages that it could conceive of that might possibly exist on our site, 99.9% of which did not exist.

    Cloudflare (blech) allows you to see and block those crawlers and as soon as I put that in place the site came back to life almost immediately. DDG's crawler was in that list and so were several chinese AI crawlers, but only Meta's wasn't following any sort of rules so that one got blocked.

  • source
  • hideshow 4 child comments
  • [–] [S] 5 points 2 months ago (3 children)

    Well the thing is DeepSeek is much more well known than DDG's AI. So why would they block DDG but not Deepseek if that was the case? And it's not just Aljazeera. The same pattern emerges for all the other News sites I listed. They don't work on DDG. They work on Deepseek.

  • source
  • parent
  • hideshow 3 child comments
  • [–] 13 points 2 months ago* (1 child)

    I just looked at Al Jazeera's robots.txt and it contains:

    # Disallow Rules

    User-agent: anthropic-ai

    Disallow: /

    User-agent: ChatGPT-User

    Disallow: /

    User-agent: ClaudeBot

    Disallow: /

    User-agent: Claude-Web

    Disallow: /

    User-agent: cohere-ai

    Disallow: /

    User-agent: GPTBot

    Disallow: /

    User-agent: PerplexityBot

    Disallow: /

    User-agent: Bytespider

    Disallow: /

    It looks like duck.ai uses ChatGPT on the backend by default, so if it's behaving it won't scrape Al Jazeera's site. You could try switching it to one of the non-GPT, non-Claude models and see if it works better.

  • source
  • parent
  • hideshow 1 child comment
  • [–] [S] 2 points 2 months ago*

    Mistral doesn't use websearch and their Gemma model when asked will attempt to search, get a 404, and then just throws "Gemma 4 31B is temporarily unavailable. Please switch to a different model or try again later."

    What Gemma is thinking before it throws this error:

    The Claude model they have just says:

    But this is specifically the search tool. If you tell it something else it can access their site:

    It is only the search function, which has a built in method to censor which sites it can search from, that has the issue. I don't know if the ReadDocument tool would somehow be unblocked while searches would be blocked. But it strikes me as odd that the searches don't even show up. Like the search tool itself throws a 404 error when searching for these things. It's not just having no results or showing the URLs but being unable to access them due to the site blocking it from reading the content. It acts as if that URL simply does not exist and it's an invalid query.

    So either Aljazeera, RT, CGTN, AP, CNN, and probably more I didn't check, have all specifically blocked the search function used by DDG's various AI models, and also all have not blocked Deepseek, or it is something on DDG's end causing the fail. And considering that the more I tried to do it they eventually just had the entire site pretend to be down, and that Gemma instead of telling me it can't do the search pretends the model is down? It seems very fishy.

    What do you think? If you know a bit about how these work on the backend do you think this behavior lines up with the websites doing the blocking? I don't know enough to say for sure, but it doesn't seem like it at first glance to me.

    Edit: To clarify when the Claude model says it gets other results but not aljazeera results it is lying. You can see it do searches and when it searches for Al Jazeera it gets back a 404. But then it tries a more general search and gets other things back. And it then acts like these were a single search when talking about it. It confused me at first so wanted to point that out.

  • source
  • parent
  • [–] 7 points 2 months ago

    im kinda wondering about proton's lumo ai which uses qwen and some other stuff, maybe do a test on that? i know they have their own controversies i just wanna know what kind of weirdoes they are, libertarian weirdoes or fash weirdoes

  • source
  • [–] 1 point 2 months ago (1 child)

    I'm sure there's some law or razor about this, but I think it's much more likely that the AI is buggy and incapable, rather than a random smattering of western and eastern news sites being deliberately blocked.

  • source
  • hideshow 1 child comment
  • [–] [S] 1 point 2 months ago

    Well that's why I wanted other people to try it. if it's a bug it shouldn't be infinitely reproducible. It should only happen sometimes. For some situations. If it's consistent with what news sites it won't pull from across various users that's a different story.

  • source
  • parent