I don’t understand scraping infrastructure.
I can make 10 fake YouTube accounts and try to scrape Koala 36M but it’s not possible. It takes like 100-1000VMs to actually do this scraping in time
Large companies don’t publish anything. They have 10s of millions scale videos and don’t even put of 10M.
Does anyone have any advice on this? Im training video models and world models.
submitted by /u/lucidml_lover
[link] [comments]