Okay, so you actually take issue with internet scraping itself? Let me get things straight, do you have a problem with LLMs themselves, or the data used to train them? It seems like an obvious “Yeah, the data is the issue,” to me but I want to clarify because I HAVE seen the other stance before.
And if it is just the data, does that mean your REAL issue is the act of internet scraping itself? Because then you’re also going to have to condemn every search engine and indexing algorithm in the world and go back to pre-indexing days when we just typed in the URL as the only method.
Or is it the monetisation that you don’t like? Because then the models I provided should be exempt because they’re 100% free.
It’s not as simple as one singular aspect, and I don’t know if I might’ve been fine with it if it was actually open and free from the start. But where I draw the line on ethics is with the usage of data without permission to create models intended to reproduce the data - which means laundering the data, reproducing the patterns from it to create competition/replacements for the people who made the creative work it’s trained on.
The monetization is a more complicated aspect, because it is also fact that at this point the big GenAI companies are dominating the market and consuming an unreasonable amount of resources to grow, and with that any use of LLMs is supporting that market. Using the services of those companies obviously supports them, publicly using LLMs supports the legitimacy of the market, and if the current trend continues, private use of selfhosted LLMs is making yourself dependent on them, and it sure seems like the tech market is doing its best to make the necessary hardware unavailable to people, meaning you’re setting yourself up to be dependent on the companies in the future.
But, that’s the fun thing - this is a short explanation of my opinion, because this is a complex topic, and I’m not writing this out every time, and people wouldn’t read it. My specific opinion is that governments need to catch up on copyright law, and hopefully GenAI is a bubble that bursts soon. If anti-GenAI sentiments become more widespread, maybe it’ll happen. If not, then I’ll have to eventually give up on my morals, but for now I can only hope it doesn’t come to that.
Okay, so you actually take issue with internet scraping itself? Let me get things straight, do you have a problem with LLMs themselves, or the data used to train them? It seems like an obvious “Yeah, the data is the issue,” to me but I want to clarify because I HAVE seen the other stance before.
And if it is just the data, does that mean your REAL issue is the act of internet scraping itself? Because then you’re also going to have to condemn every search engine and indexing algorithm in the world and go back to pre-indexing days when we just typed in the URL as the only method.
Or is it the monetisation that you don’t like? Because then the models I provided should be exempt because they’re 100% free.
It’s not as simple as one singular aspect, and I don’t know if I might’ve been fine with it if it was actually open and free from the start. But where I draw the line on ethics is with the usage of data without permission to create models intended to reproduce the data - which means laundering the data, reproducing the patterns from it to create competition/replacements for the people who made the creative work it’s trained on.
The monetization is a more complicated aspect, because it is also fact that at this point the big GenAI companies are dominating the market and consuming an unreasonable amount of resources to grow, and with that any use of LLMs is supporting that market. Using the services of those companies obviously supports them, publicly using LLMs supports the legitimacy of the market, and if the current trend continues, private use of selfhosted LLMs is making yourself dependent on them, and it sure seems like the tech market is doing its best to make the necessary hardware unavailable to people, meaning you’re setting yourself up to be dependent on the companies in the future.
But, that’s the fun thing - this is a short explanation of my opinion, because this is a complex topic, and I’m not writing this out every time, and people wouldn’t read it. My specific opinion is that governments need to catch up on copyright law, and hopefully GenAI is a bubble that bursts soon. If anti-GenAI sentiments become more widespread, maybe it’ll happen. If not, then I’ll have to eventually give up on my morals, but for now I can only hope it doesn’t come to that.