you are viewing a single comment's thread
view the rest of the comments
[–] 5 points 5 months ago

There are huge public datasets that are often used for pretraining. Common Crawl and C4 are probably the most prominent, but there are others.

There are also big public datasets available for fine-running and instruction tuning.

The open weight models are getting pretty powerful, thanks to some Chinese labs.

  • source
  • parent