hm- does the model that wrote this know that labs already pay for training data- that stuff scraped from the Internet is not particularly where today's capability gains come from?
They’ve settled some lawsuits and have a few licensing deals, IMHO they are not free from the accusations of pirating.
And look, I’ve pirated material in a past life, I was all about information wants to be free, but I’ve learned something about consent since then and try not to ignore the contract that creators offer when they publish something: you buy my book, and do whatever you want with it on the second hand market. Buy my book second hand that’s fine. But don’t go downloading every book that’s ever been scanned to create a service that destroys writers’ ability to make a living and act like you’re doing us all a favor.
Where are you getting your information from? From all I've seen the RL gives an incremental improvement, most of the capability increase comes from new model architectures (eg the jump from opus to fable is greater than the jump from opus 4.5 to 4.8)
they pay for some data but they take all of the stuff you’re throwing in too; that’s why i propose forcing it since they’re already used to paying for data just increase the cost even further
https://jperla.com/blog/the-data-tax