I was messing around with duck.ai because it doesn't need an account and I like to see how censored western AIs are. Specifically around the Iran war.

I was going to ask it to read Aljazeera to get their coverage of the war and then compare what it summarizes to the same thing from other news sites. To see if it changes what it says based on which news org I have it pull the info from. But then I ran into an interesting hiccup.

I tried this with multiple models and as soon as you specify to search for Aljazeera + specific news event it will get a 404 error. And the same for CNN.

I thought this might be the sites themselves blocking AI crawlers. So I tried it in deepseek. It had no issue pulling both Aljazeera and CNN. I also researched exa (the search plugin AIs use) and it does advertise a feature where you can block certain URLs from showing up.

It does seem possible to force it to see them. I was able to get it to access Aljazeera.com on the backend and read from their own documents to find articles but in the websearch tool itself there is no result coming up from either CNN or Aljazeera. Which is odd right?

I did some research. I thought maybe it had something to do with the Zionists. This smells of them. And the founder of duckduckgo does have ties to the founder of an "Israeli" tech firm (they seem to be cousins) and even said in 2013, I think, that he wanted to visit "Israel" in an interview he did with the Times of "Israel". So this might explain the Aljazeera part. It is banned in "Israel" after all. But why CNN? That's an American company. It tends to follow the state departments messaging.

I tried some more sites. Trying to think of any that might be getting blocked on purpose.

RT - not working Aljazeera - not working CNN - not working NBC - working ABC - working FOX - working QQ news - working CGTN - not working

After trying these I tried AP next. Then it told me "duckduckgo is temporarily unavailable". I refreshed. Same thing. It was acting like the site was not working. Until i changed my IP and deleted browser data. Then suddenly it worked again. Interesting huh? And AP is also not working btw.

So it blocks certain news sites, and if you keep trying to access them too often it blocks you from using the service by pretending to be offline I guess?

Would love to see if anyone can replicate this behavior. If you do it be careful not to let it trick you. What it will do is you'll ask for coverage from a specific site and it will find that site has nothing (404 error) and then it'll just make things up from other sites, and act like it did what you asked. You have to specifically tell it not to include results from any website other than the one you requested. And watch the thinking so you can see as it tries the searches and they fail.

you are viewing a single comment's thread
view the rest of the comments
[–] 11 points 2 months ago* (2 children)

There's a fuckton of legal back-and-forthery between news sites and big search engines over scraping info, I believe there are laws around it? I'd be quicker to wonder if this is some sort of licensing dealio.

  • source
  • hideshow 2 child comments
  • [–] [S] 3 points 2 months ago (1 child)

    Well this isn't a thing in the DDG main results. It's just with their AIs. It's easier to hide in an AI since it just tends to summarize info without always being forward about what sites it got it from. Like if I ask the same AI, "tell me about the current iran war" right now: It pulls from 4 sites.

    Notice "unitedagainstnucleariran.com" Not exactly an unbiased source huh? And it reads this one but then when it does the summary...

    Would you look at that it only mentioned Britannica and Wikipedia. And in it's summary it is pretty vague about where those thousands dying were or who killed them. I'm sure that has nothing to do with those other sources it didn't tell us about and how they might phrase things.

    Those are also simply not what would show up in a normal search. If you searched "Iran War News" in what world are those the first 4 links lol?

  • source
  • parent
  • hideshow 1 child comment
  • [–] [S] 3 points 2 months ago

    To confirm it I did a normal DDG search with the exact same syntax the AI uses. The first 2, Wikipedia, and Britiannica, are entirely normal. Those are the ones it doesn't hide. But the 2nd 2? Not even on the first page at all. I have no idea where it got those from.

  • source
  • parent