Even if the models don't collapse, it seems intuitively obvious that you can't get more "knowledge" out of synthetically generated data than went into it's generation.
A simple counter example: Bootstraping in statistics or any resampling method for that matter, can help us understand things about the underlying distribution, even though we do not know it.
There's other domains where you don't really need much input data to theoretically allow for complex deductions.
The best example is probably maths. A sufficiently intelligent AI could probably create every proof that will ever exist with just a small amount of input data. Similar deductions could be possible in physics, economics, computer science etc.
> The early stages of collapse typically involve a “loss of variance”, which means majority subpopulations in the data become progressively over-represented at the expense of minority groups. In late-stage collapse, all parts of the data may descend into gibberish.
If that is the case, why don't AI models have built in feedback loops, where the humans that use them can continuously work with the model and provide further human data. The edge models can then merge in their modifications to a public model, or be island and refined based upon the human's needs.
Because they need vast amounts of data. Humans can only type so quick. Current AI models are trained on, essentially, the totality of publicly available human writing output across the past few decades. That's a lot of text.
Understood, I'm not suggesting that your feedback would make an appreciable change to the model, but provide "human" data for refinement. Continuous feedback from each user would supply more data than having a static model.
Even if the models don't collapse, it seems intuitively obvious that you can't get more "knowledge" out of synthetically generated data than went into it's generation.
A simple counter example: Bootstraping in statistics or any resampling method for that matter, can help us understand things about the underlying distribution, even though we do not know it.
Update: typo
There's other domains where you don't really need much input data to theoretically allow for complex deductions.
The best example is probably maths. A sufficiently intelligent AI could probably create every proof that will ever exist with just a small amount of input data. Similar deductions could be possible in physics, economics, computer science etc.
We aren't getting the results predicted by the Dead Internet Theory, we are instead getting the Hapsburg Internet.
Gold.
Archive without paywall: https://archive.is/xh2fn
Why don't these companies spend a bunch of money digitizing old texts? There's bound to be tons of good training data that isn't online.
That's what Google Books was and then it all got mothballed for copyright reasons. The future of AI belongs to societies that DGAF about IP.
In that case, big corps will eat all the free data and than buy a regulatory lock in. I hate copyright, but in case if ai it is not a bad thing.
Summary what will happen
> The early stages of collapse typically involve a “loss of variance”, which means majority subpopulations in the data become progressively over-represented at the expense of minority groups. In late-stage collapse, all parts of the data may descend into gibberish.
Isn’t this essentially what people have been complaining has happened to Google search results?
If that is the case, why don't AI models have built in feedback loops, where the humans that use them can continuously work with the model and provide further human data. The edge models can then merge in their modifications to a public model, or be island and refined based upon the human's needs.
Because they need vast amounts of data. Humans can only type so quick. Current AI models are trained on, essentially, the totality of publicly available human writing output across the past few decades. That's a lot of text.
Understood, I'm not suggesting that your feedback would make an appreciable change to the model, but provide "human" data for refinement. Continuous feedback from each user would supply more data than having a static model.
Related ongoing thread:
AI models collapse when trained on recursively generated data - https://news.ycombinator.com/item?id=41058194 - July 2024 (64 comments)