assuming the training model didn't use real csam to begin with
This is a big contingency in your argument, and it's not an assumption we can make.
The stochastic copy and paste machine needs similar input to what it is expected to output. That's why there is an immense effort to collect as much data as possible from everything right now. If the relevant data isn't in the training set then it can't produce a related output.
You can train models on the output of other models, but at some point the original csam was produced and then processed for patterns by the machine. If you just copy those tagged patterns your still basing your output on very real csam.
Historically the US has banned all forms of csam including hand drawn or generated. This is a shift in that paradigm which now opens the flood gates for investigators as it dilutes real cases with facsimiles generated from real cases.
I'm not a free speech absolutist. There are categories of speech that should be protected and categories of speech that should not; and there are actions and products that absolutely should not belong in the "speech" category to begin with. Spending money is not speech, taking and owning pictures is not speech, and generating pictures of children being abused is not speech. There's no reason any of that should be bundled up in our overly broad web of "speech" just cause we want to err on the side of personal liberty over common good and protection of this vulnerable in society.