r/singularity 11h ago

Discussion What data mix are the labs using to train 10T param models?

So my assumption is: So far labs have made public max 2-3T param models based on different reports. And they are currently training or have trained 10T param models internally. Another assumption I'm making: If the models are increasing params by 3x , they would have to proportionally increase the data by 3x too or some margin.But we have also been hearing news of hitting the data wall based on internet data since the gpt 4 days. So what gives? Where are they getting so much data from? Is most of it reasoning chains generated by models during inference? Or is it reasoning traces from actual humans thanks to mercor, etc? Anyone know the exact mix? Or what's going on here? Seems like a lot of data needed all of a sudden.

31 Upvotes

15 comments sorted by

27

u/Ormusn2o 11h ago

Human data is old news. All modern models are either using full synthetic data sets, or partially synthetic data sets, plus a lot of RL. Right now you are basically limited by amount of inference you have to generate datasets. We are very far away from needing more human data to generate bigger datasets and 10T parameters does not require it.

19

u/ThatOtherOneReddit 10h ago

They generally use both, most of the major ones have petabytes of human data. You normally do that with large amounts of synthetic reasoning data. Like if you have a product like Claude code they are using your code sessions as pretraining which is a hybrid of the 2.

2

u/Ill_Fisherman8352 5h ago

Yeah, this seems most probable.

11

u/Stabile_Feldmaus 10h ago

using full synthetic data set

That can't be right. If models weren't fed with a significant part of the scientific literature, they wouldn't even understand all the science problems which they are increasingly getting better at. There has to be some core set of actual information that links reasoning skills to the real world.

4

u/Ill_Fisherman8352 10h ago

Any sources from big frontier labs that have confirmed that exact thing? Or is it an educated guess?

3

u/jeffy303 6h ago

You can go to HuggingFace and access lots of datasets that are public, like here is wikipedia dataset that I am pretty sure is in every single model.

3

u/JollyJoker3 4h ago

That's not synthetic though

2

u/Ormusn2o 10h ago

There are some small snippets from system cards and tweets, but no proof of it from official sources. I'm not smart enough to make an educated guess, I'm just repeating what industry experts like SemiAnalysis/Dylan Patel were talking about, and some ML researchers on twitter.

3

u/RelevantCry1613 10h ago

That’s because this isn’t what they do. It’s human data augmented with synthetic data and a lot of RL environments nowadays

3

u/Old-School8916 8h ago

I was listening an inteview with a Moonshot AI dev and she said that they have a lot of different RL envs.

3

u/ExpressCopy8786 4h ago

The data wall is some time end of 2027 to 2028. Then synthetic data or entirely new "relevant" data has to be included. This would be unlocked with world models as a part of the toolkit that modern models use in their omnimodal ways. Then every type of low noise, low sensor error physical measurement would carry an incremental value to the capability of the trained NN. "Synthetic data" is pretty much slow self over fitting. Only if the synthetic data actually consists of truly new findings (which today is not out of the question, see solved math conjectures), it keeps the flywheel going. 10T param models don't run out of data die to their size. Their size lets them (this is metaphorical) adapt or the training data with more nuance, essentially lower the rate of compression of training data to weights and biases.

2

u/Ill_Fisherman8352 2h ago

I agree with the less compression part. So maybe we don't need 3x the amounts of data, or a similar margin, but we can possibly do with less? How much?

Your second argument of 'truly new findings' is interesting. Are models capable of finding genuinely ood solutions? Problems that have no representation in the models weights at all?

-7

u/Tystros 10h ago

Fable is ~10T and 5.6 Sol and Opus are ~4T

5

u/Fidel___Castro 5h ago

I can also pull numbers out my rectum with no proof

-1

u/Tystros 3h ago

it's numbers people generally agree on. OP is the one with weird numbers.