A sex offender convicted of making more than 1,000 indecent images of children has been banned from using any “AI creating tools” for the next five years in the first known case of its kind.

Anthony Dover, 48, was ordered by a UK court “not to use, visit or access” artificial intelligence generation tools without the prior permission of police as a condition of a sexual harm prevention order imposed in February.

The ban prohibits him from using tools such as text-to-image generators, which can make lifelike pictures based on a written command, and “nudifying” websites used to make explicit “deepfakes”.

Dover, who was given a community order and £200 fine, has also been explicitly ordered not to use Stable Diffusion software, which has reportedly been exploited by paedophiles to create hyper-realistic child sexual abuse material, according to records from a sentencing hearing at Poole magistrates court.

you are viewing a single comment's thread
view the rest of the comments
[–] 16 points 2 years ago (55 children)

Where does the training data come from to create indecent images of children?

  • source
  • parent
  • hideshow 55 child comments
  • [–] 51 points 2 years ago* (48 children)

    It doesn't need csam data for training, it just needs to know what a boob looks like, and what a child looks like. I run some sdxl-based models at home and I've observed it can be difficult to avoid more often than you'd think. There are keywords in porn that blend the lines across datasets ("teen", "petite", "young", "small" etc). The word "girl" in particular I've found that if you add that to basically any porn prompt gives you a small chance of inadvertently creating the undesirable. You have to be really careful and use words like "woman", "adult", etc instead to convince your image model not to make things that look like children. If you've ever wondered why internet-based porn generators are on super heavy guardrails, this is why.

  • source
  • parent
  • hideshow 48 child comments
  • [–] 3 points 2 years ago (7 children)
  • [–] 2 points 2 years ago (6 children)

    I'm not going to say that csam in training sets isn't a problem. However, even if you remove it, the model remains largely the same, and its capabilities remain functionally identical.

  • source
  • parent
  • hideshow 6 child comments
  • [–] 0 points 2 years ago (5 children)

    At that point it's still using photos of children to generate csam even if you could somehow assure the model is 100% free of csam

  • source
  • parent
  • hideshow 5 child comments
  • [–] 1 point 2 years ago (4 children)

    That would be true, it'd be pretty difficult to build a model without any pictures of children at all, and then try and describe to the model how to alter an adult to make a child. Is anyone asking for that though? To make it illegal to have regular pictures of children in these datasets?

  • source
  • parent
  • hideshow 4 child comments
  • [–] 0 points 2 years ago (3 children)

    No but it is a reason why generating csam should be illegal. You're using data trained on pictures of real kids

  • source
  • parent
  • hideshow 3 child comments
  • [–] 0 points 2 years ago (2 children)

    I'm not arguing whether or not it should be legal, I was just offering my first hand experience in regards to the capabilities of these local models since people seem to be confused as to how this actually works.

  • source
  • parent
  • hideshow 2 child comments
  • [–] 0 points 2 years ago (1 child)

    Is anyone asking for that though? To make it illegal to have regular pictures of children in these datasets?

    I was responding to this part of your comment which directly refers to legality

  • source
  • parent
  • hideshow 1 child comment
  • [–] 1 point 2 years ago

    I guess I just misunderstood what you were arguing then. For posterity: I believe datasets containing children is fine, datasets containing csam is not, and the legality of generating csam should be left up to psychologists on whether or not it is a societal net benefit. Whichever way is better for children that exist is my vote.

  • source
  • parent
  • load more comments (38 replies)
  • [–] 28 points 2 years ago (2 children)

    The whole point of diffusion models is that you can generate new concepts using training data. Models trained on any nsfw images can combine those concepts with any of its non-nsfw concepts. Of course, that's not to say there isn't CSAM in any training data, because there objectively has been in the past, but there doesn't need to be any to generate it.

  • source
  • parent
  • hideshow 2 child comments