Vent: that free AI art tool trained on 5 billion images and I found the exact number buried on their site
I was digging through the FAQ on a popular image generator (the one everyone posts on here) and it says right there that the model was trained on 5.85 billion image-text pairs, most of them scraped without permission. I checked three other big generators and they quietly admit the same thing in their fine print, which means half the stuff in any showcase thread might have a sketchy origin. If you post your work publicly, do you watermark it or just accept that it'll get scraped into the next model?
Wait, 5.85 billion pairs? That number is so huge it barely even means anything to me, like my brain just can't picture it. I get that scraping is kind of the wild west right now, but seeing it spelled out in their own FAQ would make me feel weird about posting anything I actually care about. I'd probably start watermarking stuff, though I know that barely slows anyone down.