10
submitted 3 days ago* (last edited 3 days ago) by supersquirrel@lemmy.ca to c/techtakes@awful.systems

"It's very hard to collect data that trains the model to be better than humans, because there are very few humans who can create that data," Raj said. "You want to make it better than a Fields Medalist or a Nobel Prize winner. How do you collect that?"

...

While there, he worked alongside other researchers to understand the failure points of models hosted on cloud computing infrastructure and offer specific solutions, he said. He also created "synthetic" data that, unlike human-created writing or code, is generated artificially before being reused as training material for the LLMs, Raj said.

Creating synthetic data that matches the quality of human-created data is a tall task, Raj said. Computing power can be acquired relatively easily, but finding the data to support the project is uncharted territory, he added.

sigh

Garbage in garbage out, even if the garbage is synthetic that doesn't make it not garbage...?

you are viewing a single comment's thread
view the rest of the comments
[-] supersquirrel@lemmy.ca 1 points 3 hours ago* (last edited 3 hours ago)

Well said.

The synthetic data thing here pisses me off too, synthetic data has uses in science predicting sensor responses, comparing reality to expected findings, and analyizing related phenomena to the artificial data at a fine resolution with computer modelling but none of those things have to do with establishing a ground truth for what the model considers part of reality, part of its understanding of reality or part of the facts that supposedly underpin the reality.

If a scientist wants to analyze several types of algorithms and compare them maybe they might make a set of synthetic data that is artificially clean and simplified in order to compare and contrast the behavior of the algorithms especially at their edges and extremes. Note however that nothing about this process makes the algorithms smarter, the generative part is what the human scientist learns by observing what happens when the synthetic data is inputted into an algorithm. You need a human brain that understands context, understands the limits of a model vs the rest of reality, and understands things that aren't explicitly said about the framing context of what is being examined.

"A.I." is a lossy data compression algorithm, there is a fundamental "knowledge entropy" here where the end result can never be smarter than the raw data because the "A.I." can do nothing but apply a lossy data compression algorithm to the training data.

This is not a cynical take on the potential for artificial intelligence but rather a hopeful and heartfelt thanks to the professions of librarians and archivists, for surely it is the curation of a quality data set where the genesis of intelligence happens. If nothing else Machine Learning proves that with brute force...

this post was submitted on 11 Aug 2026
10 points (100.0% liked)

TechTakes

2640 readers
37 users here now

Big brain tech dude got yet another clueless take over at HackerNews etc? Here's the place to vent. Orange site, VC foolishness, all welcome.

This is not debate club. Unless it’s amusing debate.

For actually-good tech, you want our NotAwfulTech community

founded 3 years ago
MODERATORS