Why are you using the future tense when AI companies have already scraped the ever living crap out of every git forge it can access?
post
pretty much, with the caveat that code that has gone through an llm can't ever be licensed or copyrighted. it's basically a public domainifyer.
That doesn't seem to stop corporations assuming their software is still theirs even when an LLM wrote a lot of it.
No, code that has been purely written by an LLM is not copyrightable.
As soon as a human writes a prompt, a correction, a design guideline, a code review, it becomes a question of who has the better lawyer. Which I would bet the billion dollar Corp has the better chances.
so they would have to argue what counts as a transformative work of plagiarism. how much of a stolen painting you have to paint over before it's no longer stolen.
I'm not a lawyer, no idea what would they argue, I just know the lawyer price beats being right many times.
What you're thinking of is this: Malus.
It's an example of Clean-room design, basically you have a LLM read the code and write a specification, which a second LLM uses to rewrite the code without access to the original.
Although since the original code was most likely included in the LLMs training data, this might not really be really true.
Sure!
But I don't expect it to change much.
I could already do that, by hand as well.
It's a bit like how there's so many different superheroes who are obviously just off-brand SuperMan or off-brand Captain America.
Minor changes to avoid intellectual property law and branding has always been an option.
And I suppose all of these are easier now with the remixing slop-o-trons.
But they weren't terribly difficult, or particularly uncommon, even before.
You could always type it over and say you've recreated it, but that also didn't fly. Why would overfitting a machine learning algorithm on the data and then having it predict next tokens be any different?
Because machine learning is already basically a mass copyright infringement. The training data contains copyrighted material. The model is clearly a derivative of the training data. The output is clearly a derivative of the model. Yet somehow, it's legal (probably because they can afford good lawyers).
From my understanding, code is still covered by copyright. This means that copied code, even if run through an intermediary like an AI, is still copyright infringement. In the same way, even if an image generator recreates a character or movie frame, it isn't made public domain (the default state of AI Output), its just that the AI ingringed on someone else's copyright. If the code or image is then used, you can still be sued.
If you mean "can I just recreate an existing copyrighted work without it being copyright-infringing", no. You're still liable if you recreate text that would be considered a derivative work. You don't have a fantastic mechanism to avoid that with existing LLMs, though I would guess that you most-likely aren't going to generate infringing code randomly.
Same thing for images or other media.
all 24 comments