Volume, and breadth of subject material. For language model purposes the specific content isn’t really important, it’s the wide variety of language samples from many sources, all in the same data format. You could get similar samples from other platforms, but you’d have to compile them from multiple sources and then standardize them somehow for input as training data.
Why anyone would train an AI on Reddit is beyond me, the only worse platform I can think of is 4chan.
Volume, and breadth of subject material. For language model purposes the specific content isn’t really important, it’s the wide variety of language samples from many sources, all in the same data format. You could get similar samples from other platforms, but you’d have to compile them from multiple sources and then standardize them somehow for input as training data.