This would be so cool, but we need to think more about how we could do it and make enough money in the future to train more models with even cooler features.
Neither of those look like they have a generative AI component.
We (as a society) desperately need a way to train these models in a federated, distributed manner. I would be more than happy to commit some of my own compute to training open audio / text / image / you-name-it models.
But (if I understand correctly) the current architecture makes this if not impossible, nearly so.