It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
> It's the robbery of all of our culture to sell it back to us at a mark-up.
Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.
Whether the end result threatens the form in which culture is created, at least beyond just threatening the business models of the gatekeepers, is a separate discussion, but you can't draw the heart-string-pulling "life's work got appropriated" arguments there so easily.
And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.
There's plenty to hate about how AI is transforming the world, but one thing it's not, is "robbery of all of our culture to sell it back to us at a mark-up".
Allow me to suggest a third: these two options are black and white thinking and there is no objective answer to "intellectual property is real". Property is at best a social construct that is possibly supported by instinctual behavior.
Instead we should try to find the most practical solution that has the most benefit - which will probably be more complex than yes/no.
> Property is at best a social construct that is possibly supported by instinctual behavior.
That very much falls within the "IP is not real" category. Taking a utilitarian approach here is exactly what the sam altman / effective altruist crowd is doing (or claiming, at least).
Believe it or not, but there are different philosophies. The US Constitution, for example, is written under the framework that all rights are innate, and the government merely endorses, not grants, those which are described. This framework creates a moral basis to rights, such that things like property are not merely social constructs but moral goods. To violate them is itself an immoral act.
> That very much falls within the "IP is not real" category.
AH. Given that framing, then I suppose anything that varies from "these axioms are perfect there can be no others" must fall into the "not real" category.
But isn't there some debate about the moral axioms themselves? A space that is much more complex than "yes/no"?
> Allow me to suggest a third: these two options are black and white thinking and there is no objective answer to "intellectual property is real". Property is at best a social construct that is possibly supported by instinctual behavior.
Intellectual property is as real as private property (which is a lot more elaborate and weird than possession and territoriality, which is the most that has any natural basis). It's really foolish to claim one doesn't exist and should be abolished and the other this fundamental sacred thing that should be respected absolutely (as many do).
We made it up. We made all this up. We do it for an outcome.
Intellectual property exists as a concept to foster the creation of more intellectual property. That is not the case for all private property, because I cannot copy your land or your car infinitely. That’s why IP rights expire at some point, or have fair use that doesn’t harm the IP rights holder, something other forms of property don’t have. Saying that they are equivalent is not foolish—I think that’s a bit extreme—but it doesn’t recognize that they have fundamental differences, and the laws around them have different intended outcomes.
I think AI training falls most likely in the fair use category of intellectual property: there is some societal benefit* that requires no actual harm** to the IP holder, therefore it’s a good trade-off for society if we poke a hole in the social construct of property to get that benefit.
*Let’s put aside the question of whether AI is good for society or private ownership of AI models is good for society. Important questions but separate from the theory. IF IT IS GOOD, it follows the above. If it is not good, then of course it does not.
**Also not a fully settled question. Again, an important debate to have and the tradeoffs here matter. If the harms are small enough, the societal benefit could be worth it. Both notes have to be true for this to be worthy of “fair use”.
Yeah, this whole "we're being attacked by China, they're distilling our models!" thing is completely and utterly absurd. Insulting, even.
Either IP exists and should be respected or it doesn't. And if model training is fairly "transformative" of the source work, then so is distilling. As Garry Tan suggested recently, we need to encourage a US distillation regime. Everyone should be free access and transform this information however they see fit.
What do you mean, it's already robbing me of the ability to discuss techniques with a number of my peers since they decided the mediocre output of LLMs is good enough instead of understanding what they are doing.
In addition, a great number of people have decided to stop publishing their code publicly, and discuss techniques except in spaces they can be sure it's a one-on-one with a human being.
2) Spiteful peers with "dog in the manger" mentality.
The latter case personally irks me, mostly because this often involves personal benefit (economical or moral) due to them claiming to give away knowledge for altruistic reasons, which later behavior reveals was a lie, mislabeling proprietary as open and free, to reap unfair gains.
> Spiteful peers with "dog in the manger" mentality.
Last night my daughter told me her friend (a boy) signed up to take the last spot of a baking class at camp just to spite his sister, who wanted it.
Turns out he liked baking!
---
I think this whole discussion shows that there's a lot of Calvinball going on with IP, where the creator keeps trying to change the rules as a way to defend their power and social order.
This may blow your mind but people have different value systems than the single one where only making money matters in our incredibly short lives. Some people believe in things other than accumulating money and immiserating workers.
Your failure of imagination doesn't limit the actions people can take to shape society. We all have equal voices in that respect.
I truly feel sorry for those with your mindset. It speaks volumes on the lack of community and friends you have where you think fucking them over to make money is a worthy pursuit.
They use to banish greedy people from society in the before times, maybe we should start again as these people have no problem in destroying society for religious fanatical reasons. We shouldn't be accepting of them, especially when they're actively killing us all or making us miserable.
You are preaching this is response to a comment that mocks the notion that creators are the evil people while OpenAI, Anthropic and SpaceX exploit them?
People creating free software for everyone to use, only to pull it away when suddenly there's a technology that can instantly let everyone use it at the drop of a hat, is not something I quite understand. If you're doing it so you show people, hey I made this, then you can still do that. If you're doing it to help the world, then that's very much happening. Of if you're doing it just for fun, that's possible too, though I totally understand how much less fun it is when AI can rip it out in a few minutes.
Some of it might be loss of external validation, no more stars in github. Some of it might be an aversion to a non-consensual contribution to the end of their chosen profession as a means to make a good living.
Much of it might be because their license (MIT, BSD etc) was ignored completely.
People have multiple motivations at the same time. Many OSS contributors have moved on after LLM’s and that’s completely consistent with their motivations.
Some of the major OSS licenses require attribution when someone copies the work in question, AI companies fail to do so. People getting annoyed when others break the social contract has nothing to do with any other motives they can always benefit the world by doing something else.
You may prefer someone continue maintaining infrastructure you use but they could be just as happy delivering food to the elderly. The desire for people to continue helping you is little more than entitlement here.
> a great number of people have decided to stop publishing their code publicly
Stopped publishing anything and stopped creating anything.
But in all fairness this started long ago. I remember a conversation where I asked someone who read a niche blog every day why he never posted a comment. (the blog had zero comments) If you like what they wrote and have thoughts about it, why not "honor" them with a few words? How is that to much to ask? They clearly just never bothered to but in stead made up all kinds of excuses.
Imagine having to explain that you can respond if someone talks to you? If there was a culture surely it ended there?
It's the same as saying "People don't kill people; guns do." Maybe guns(AI) should be regulated, but the primary problem is murderers(lazy peers), not weapons(AI).
Yet countries that have strong regulation of firearms have significantly less murders (and accidental shootings) than countries that don't regulate them.
That does suggest that along with murderers, the easy accessibility of firearms is a problem (assuming that the ratio of murderers in different populations is approximately the same).
Not much more than with a bash/curl loop. Which, on one hand, AI can write for you, but on the other hand, AI will make you no longer need to pirate those articles in the first place.
Also it's not end-user piracy being discussed - it's the act of training AI itself that's accused here to be "robbery of our culture".
If we ignore the massive amount of books being destroyed, the outages and increased usage bills for once free websites (some resulting in closure), the increased difficulty to access public information such as Reddit and Twitter, and lastly ignore the amount of conversations being taken away from the public in favor of LLMs, I would agree.
Your talking about a bunch of books that are out of print and sat unsold for years in a warehouse somewhere. The rights holders could print another run tomorrow if they wanted, or better yet digitize them but they don’t. Where’s your ire for the rights holders who sit on these books and don’t do anything with the ?
Except books are already destroyed in physical form every single day. My local library has a section near the front for free and/or extremely cheap books ($1) they are desperately trying to get rid of. My impression is that if they are untaken/unsold they will end up in a dumpster.
The scale is completely different. It is like comparing a tiny ant to a Star Devouring Shoggoth. Anthropic destroyed millions of books in hardly any time. They are destroying the foundations of civilization
> They are destroying the foundations of civilization
No, they are not. For all the hype around that 404 media article, no one came forward with titles or authors. Turns out the booksellers that sold them those books said they were "dead inventory".
If you're going to assert that companies are destroying the foundations of civilization, you're going to have to produce quite a bit more evidence than the (poorly cited) 404 Media article.
Surely that's not about AI? Or you mean the small fraction of that that's result of "destructive format shift" process[0], which itself is a consequence of copyright regulation that would otherwise prevent anyone from accessing these works?
> the outages and increased usage bills for once free websites (some resulting in closure)
That's assumed to be AI companies for some reason, even if they have no real incentive to do that, while the usual business underbelly of people scraping web for whatever reasons (which now may include some AI upstart wannabies too, to be fair) is forgotten about. Not helping is people confusing AI agents acting as user-agents and doing one-off fetches with "AI scrappers".
> the increased difficulty to access public information such as Reddit and Twitter
It was never public, and they started locking down before AI, when both platforms (as well as all other social media platforms) run out of VC subsidsies and realized they need to start monetizing; first step they did was to wage a war on third-party clients and API users. AI came later.
> ignore the amount of conversations being taken away from the public in favor of LLMs
You mean human agency? Like, humans deciding it's better for them to ask a machine instead of posting questions online? I can understand that, it's usually much better experience (particularly on sites like StackOverflow - they dug their own grave here, and they know it; there have been memes about this way before LLMs were a thing).
You're also not considering the amount of questions answered that would not have been asked otherwise. I for one don't ask many questions on-line, so anything I ask LLMs that they solve for me, is a question that would've remained unanswered for me otherwise. Not everyone is gregarious online, many people are more self-reliant and only answer questions, but solve their own problems without asking for help (it's probably not optimal thing to do, but that's another topic).
--
[0] - Digitizing and retaining physical copy is clear infringement, digitizing but destroying the physical original can be argued to be fair use.
> You mean human agency? Like, humans deciding it's better for them to ask a machine instead of posting questions online?
Oh, please, AI is single most hated technology to the level we did not seen before. And that is after staggering propaganda going out of these CEO that hiding behind human agency is beyond lie. It was pushed on us, whether we like it or not, because people with a lot of money are the only ones who matter.
And people with a lot of money have that weird cult belief in emerging AI god and singularity, so they dont care what they destroy in the process.
Creators are unwillingly and contra to economic systems that have evolved over a few millennia entered in the Borg or the Matrix. So their achievements are reused for private and public benefits without their permission. Piracy is theft. And this piracy is a bigger theft than piracy on an individual download basis. Was there any doubt? (Linguistically I agree that you can’t “rob culture”. You can rape or reap culture though, and that is the point at hand.)
HN crowd: Check out my Plex server and 132TB media collection!
Also HN crowd: Reading publicly posted information on the open internet is morally outrageous theft of the highest order
Because of the nature of upvoting on social media, we can pretty accurately dial in on what the crowds consensus is.
Your response makes sense when we are pulling random marbles from a bag, like old school forums, but social media abandoned that style of commentary long ago.
I'd bet if you asked a bunch of people off the street, you'd have a number of different takes on the ethics of stealing a loaf of bread, depending who stole it, from whom, and why.
The crowd: "Knowledge should be free! Culture belongs to all of us! Information must be democratized!"
AI companies *proceed to copy literally everything and put it through virtual blender, until an universal general-purpose problem-solver tool comes out, then give it out for free or serve for peanuts to literally everyone on the planet with Internet connection *
Except that no one is freeing or democratizing the knowledge. We all know that the plan is to take it, destroy original sources and then become the solo monopoly provider.
I’m paying $200/mo for non-crap versions of said “general-purpose problem-solver tool”, which is not far from the median income of “everyone on the planet with Internet connection”. And I’m told I’m already getting a huge discount by using thousands of dollars of compute by raw API pricing. That just doesn’t sound like peanuts at all.
"Give it out for free or serve for peanuts" is doing a lot of work here. The crowd should be skeptical that this will continue given how many times we have seen enshittification unfold before.
Before LLMs, "knowledge should be free" was a call to preserve an open playing field where everyone has the ability to access the intellectual labor of humanity so they can contribute, instead of letting a few institutions wall it off for a profit.
As knowledge is walled off behind these models, rent seeking will expand, "safety guardrails" will become more onerous, and the masses will lose their access. There are wonderful parallels between this and enclosure in Britain.
>instead of letting a few institutions wall it off for a profit.
I cannot repeat enough that that is not true, or even remotely close to true.
When ChatGPT helps me fix my toilet, saving me a $300 plumber call, I don't cut OAI a check for $300 instead.
It is unbelievably easy to harvest more than $20 of value from all the basic plans. Right now, users are overwhelmingly capturing all the value. Even if they triple the price, it's trivial for most workers to easily justify it.
We can project any future scenario we want, but right now the labs are only capturing a tiny fraction of the value they are creating. Especially for the crowd here on HN.
There's some nuance there you're purposefully/ignorantly supressing to create a more impactful fundamentalist observation, completely contrary to the comment you described above as :
"heart-string-pulling "life's work got appropriated" argument"
I think the water is very muddy in regards to this topic and comments like yours really highlight the distain some people have for others that have talking points about with things they create being thrown into a gargantuan 'blender' as you called it.
Amodei: Check out my 1000 Exabytes torrent collection, which I weaseled out of with a $1.5 billion bribe.
The usual human reading information on the internet does nothing with it. They don't adapt the style, they don't translate, they don't reblog a modified version. The ones that do have been criticized for stealing blog posts already before AI.
An individual pirate is morally a lot less objectionable than the capitalist class pirating from literally everyone else.
Yeah sure, it's on the open Internet, but then again, the capitalists could also actually adhere to the licensing terms of the free software they're hoovering up and all that. These are the same people who want to enforce all kinds of EULAs against the commons, employing measures such as DRM to protect their "intellectual property", so why is it suddenly okay to not comply with IP when it's them doing the infringement?
And that's just with software, but this also affects other creative endeavours. Also the individual HN user's 132 TB Plex server doesn't suck up ludicrous amounts of resources from the commons.
The past is still there, but our culture, being a live thing, is pretty much being affected by AI. We now have an artificial intelligence in this loop of exchange of ideas. Slopifying everything in its way.
Content creators are taking steps to prevent bot from walking away without compensation for their visit as if they’d been stolen from. At least Google, in the beginning sent traffic back that could potentially be monetized by the creator.
Perhaps some. But I'd wager orders of magnitude more are using LLMs to brainstorm if not outright generate content for them, which they are then reframing as their own creation.
> Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.
It hasn't been literally stolen but indirectly it has been.
I spent 20 years working as a freelancer and 10 years selling tech video courses. I made enough to have a happy life (not a lot but enough to survive), as long as the income kept flowing every month.
Nowadays I make nothing from a business I've built up for 2 decades because AI took away most traffic to my site which was the entry point to my business. I actually lose money because course sales have been impacted so heavily that I pay more for hosting than I get in sales.
AI is consuming an unimaginable amount of searches which is the gateway for so many businesses. It's preventing people from being discovered on the internet. Not only that but it's taking everyone's content and selling it back to users while keeping all of the profits.
> It hasn't been literally stolen but indirectly it has been.
It's literally stolen IP.
One minus epsilon of the corpus did not give informed consent, or get compensated. The fact that it's laundered doesn't make it any less literally stolen.
> Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.
This logic only works if you believe that IP has no value. Which is, of course, utter nonsense.
Microsoft doesn't believe this. Nor do any of the AI labs. Otherwise, they wouldn't have needed to scrape the data in the first place as it had no value. All of their software (and other products) would be developed in the open as there's no point in protecting the IP.
They work very hard to protect their own IP so they obviously believe that IP is worth protecting.
Yep: the problem with IP is exactly the "Property". The texts are economically (even if only for leisure) valuable, so the producer deserves a fair comoensation unless he has explicitly dismissed it.
So paying a price for an intelectual product is reasonable. Nothing to do (per se) with property.
LLMs have made it easier than ever so search for things. I remember fumbling around with Google, trying different combinations of search terms and never hitting useful results. Now I can give a vague or even partially wrong prompt to an LLM and it sort of magically figures out exactly what I actually want and takes me right there, along with a helpful summary and analysis.
> Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form.
But there's a real [1] transgression there, and attempting to lawyer it away is disingenuous. This isn't a perfect analogy, but it's sort of like if you blocked off the sun, and a farmer complained you stole his land, and then you retort "your land has not been stolen, you still have title and can occupy it." Your actions damaged him. Again, not a perfect analogy, but I think our efforts should go to recognizing the transgression instead of trying to deny it.
> And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.
We don't have a society where "we got back" anything. If anything "we" got undercut, and "our" assets lost a lot of their value. Come back with communism, and then maybe you'll have a point.
[1] Fuck you, Claude, for making me wince at that.
It is robbery because you're not paying the person that produced the knowledge, you're paying the person that took the knowledge and put a chatbot interface on it.
It's different from SEO because you're not even getting the chance to monetize it yourself, you lose the ability to even show some ad for a pittance. Many many people didn't give permission for the data to be ingested and used this way.
We've come out of decades of copyright infringement takedowns and pirating lawsuits only to end up in a position where people doing similar things as a corporation are making billions of dollars.
> Except, of course, no one has actually been robbed
No, theft has been documented already many times. See how Facebook etc... slurped up data from Anna's Archive, Libgen and so forth. And many more examples here. The law classifies this as theft. Even digital theft is theft according to the law. Since corporations have too much money, nothing will happen, but we all see that this is theft. It is not clear why you do not see this.
There are many refutations to this fallacy from the Slashdot era, but with AI it is even simpler than with music:
Clankers are only useful for current events (which most people use them for) by scraping the web and rewording it. This results in direct financial and notoriety losses for the original authors of the websites.
To use another Slashdot cliche: "But you knew that already."
AI didn't steal any culture. It just made mediocre culture more accessible. Turns out, most people do like mediocre culture. Previously, public TV channels could at least pretend that people enjoy educational content or classical music. AI exposed that lie. And now we're shocked that the emperor is naked.
Our "culture" has long been the province of corporations. In prior epochs it was still the product of patronage and power.
At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.
Instead of fixating on a remedy that seeks to criminalize AI, maybe focus on the relatively rather achievable goal of redistributing LLM gains. Would that not be the most desirable justice? What is your alternative, and would you foreclose the future in the name of a past that never really existed in the first place?
> At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.
Let me fix that for you: A world in the the value of the labor of each human approaches zero.
Humans with zero economic value can still vote. They can still mass. They will still have needs. Really the script here writes itself. The historical precedents bound the problem rather well. As always, radical social change will not occur until a wide swath of the population is aligned. In this case, due to their broad economic devaluation.
It would be easier if today's knowledge workers stopped deluding themselves into thinking that their standards of living will maintain. Your acknowledgement of the necessary predicate to change is, in that respect, progress in itself. There is little reason why we cannot accelerate the timing of broad consensus if more people so readily came to that conclusion - and resisted the temptation to then find the answer instead in nostalgia about the past.
You assume there will still be free elections and democracy. There is also the possibility that the form of government will change into something else.
We are not deluding ourselves that we will maintain our standard of living. If the boss’ dream of replacing all workers by AI that does their job poorly but cheaply will be realized. We will go to the gutter with the rest of the economically worthless people.
There are many of them today, their needs are not met (but could be, the production output is largely there, but captured) and their political views don’t matter because the votes are captured by populists. I’m sure the powers that are will find a way to go around educated people voting.
It is easy to ignore the misallocation of output and the private greed when generally people are still doing fairly well.
As the denials give way, the politics will change. You may be right such moments will be hijacked and hope extinguished. That's why we should all aim to think about these problems in the most robust way, so we can all contribute to the coming efforts.
Cynicism about the future is easy, but should be resisted. Perhaps ironically, I find hope in your acknowledgement of your own imminent devaluation.
Humans are animals. We have rules to keep humans in check but that’s it.
The wealthy do not care nor need to care. They will justify it as ‘you lot were too weak to organise and stop us’ and there is an element of truth to that.
The value will tend towards infinity [1] because we can keep building more data centers. This will allow us to apply tokens to solving every problem we can conceive.
The cost will tend towards zero because competition is intense and there appear to be zero moats and ample improvements from every direction.
If these systems have low costs, that means they will be usable by the broad population and that their utility will be widely accessible.
Industrialization led to iPhone, PlayStation, Spotify, and Waymo.
AI will lead to personal chefs, contractors, assistants, drivers, tutors, climbing partners, ...
AI will lead to people making their own PlayStation games, their own music and streaming services (I already have), their own custom smartphones (a future personal project - vibe hardware). And your robot will drive you cross country on vacation while you sleep in the car.
[1] Not actually infinity because earth [2] has finite resources, but the S-curve will look like it for awhile.
Yeah, S-curve should be classified as fallacy, because it rarely covers what people argue it does.
Economic and technological growth is a stack of S-curves, where one very specific facet may hit limits and taper off, only for equivalent, complementary or alternative facet to take off in its place. Added up, there's no sign of the exponent stopping any time soon, not until hitting real limits, or (probably more likely) some general catastrophy that shuts down human civilization.
Flourishing life? Where is your data leading to the conclusion that AI models are leading to the world populace leading a ‘flourishing life’? Are taxes on revenue of companies like OpenAI somehow being collected and turned into a UBI and I just didn’t hear about it?
I said a glimpse. I didn't say it's here now or it would be easy.
It is however, achievable. Certainly more so than engaging in the fantasy that we can criminalize LLMs out of existence. And it is likely more desirable than such an effort anyway.
Or, perhaps, the companies forming the for profit LLMs could be made to pay a license fee for the copyrighted data they lifted from behind the paywall and then used to form a for profit entity with.
Nothing about banning technology. How about enforcing DMCA and then applying a penalty for the knowing theft rather than negotiating a license?
This is solid but how many people you reckon would benefit if they paid every single penny for any copyrighted data? I am in my 50's with 30 years in the industry and I know maybe 2 people that might benefit from this... certainly not something general public would benefit from
We’re likely looking at a new world where physical and mental work from humans holds much less value.
So if the world still generates value ( think AI inventing new drugs, robots planting and harvesting fields of corn ), who is able to make income? The people who have the capital to purchase and run the systems.
So our future world probably looks like a system where effort/knowledge/skill has little to no reward and capital (e.g. inheritance, passive income) has all the reward.
So, we can have a world where you are born rich or you live in abject poverty, or we can figure out something else.
The state runs the system. Or really at a certain point, the state runs itself. In 100 years or so, the CEO becomes a barbaric relic of the past age, like our ancestors who beat each other over the head with clubs. Nevertheless, such a future will necessarily owe a debt to both.
Wealth and states are often connected. In this case if the state is properly aligned (this is what the coming battles are to be about - as they always have been) to be egalitarian, then redistribution can occur unblocked by human enclaves of greed.
> We’re likely looking at a new world where physical and mental work from humans holds much less value.
Or, perhaps, we're regressing to the mean after a century where this work was over-valued, because companies like Disney succeeded in regulatory capture and created artificial protections to maximize their own revenue.
How much money did Bach earn from royalties (ok, there's a pun there, but I mean payment for reproduction of his work)? How much power did Melville have over who published Moby Dick, and where, and how it was used (hint: very little in the US, none at all overseas).
We are seeing a weakening of control and revenue extraction from copyrighted works. But on the chart of history, the 1900's were a very anamolous spike in that area. And it saddens me that so much of HN is unhappy about more of a return to the commons.
> So if the world still generates value ( think AI inventing new drugs, robots planting and harvesting fields of corn ), who is able to make income? The people who have the capital to purchase and run the systems.
On the bright side, I think you're wrong here. Or (as Claude likes to tell me), you're half-wrong. If this were true, it would already be the norm and only big companies could bring new products to market. Yet startups are a thing, and many succeed (more fail, but still.
The trick in entrepreneurship and creative work has always been knowing what to ask the system to do. You might as well say that music is dead because the synthesizer and music software companies can produce as much as they want at zero cost. It turns out that owning the means of production only loosely correlates to producing things people want.
> So, we can have a world where you are born rich or you live in abject poverty, or we can figure out something else.
Already done. There are kids in Africa studying with AI tutors today, getting insight that they would likely never have had access to before. One-person shops are releasing board games and business services that they would never have had the capital to do before.
Where you see centralization of capital and control, and the masses reduced to abject poverty... I see the complete collapse of barriers to entry and switching costs in many fields.
But who am I? Some random guy. But I've heard very senior execs, at Microsoft and other companies, in absolute panic that the IP and systems they've spent billions of dollars to build over decades of work are suddenly subject to disruption by teenagers who are great at using AI. Seriously, panic.
It's a complex subject, and sorry for writing a book, but your prompt apparently got me. None of us know for sure what will happen but I think yours is a needlessly pessimistic view, ironically informed by the exact abuses of th past century that you're worried might not continue.
I spent a few years studying the ‘startup economy’ after a 20 year career at MSFT that left me with some meager capital, including a half dozen Angel investments and running an Angel investment fund LLC (SPAC as they would call it today).
Never saw a startup succeed that didn’t come with substantial founder capital, across a couple hundred evaluations. Does it happen? Sure. Is it the model for new product generation that people with <$250k of personal liquid assets can bring a new product to market? Not really.
‘You got to have money to make money’ is a real thing.
Yeah, I know I’m posting on YComb’s forums, but if you think the capital YComb itself gives to ‘good ideas’ is enough to succeed, that’s silly. It’s the follow on investment capital that brings a new concept to market, and that comes with odious strings for people underendowed with their own capital.
There's an analogy I've read before, couldn't tell you where, that goes something like:
Startups are like a roulette wheel at a fancy casino. Most people can't even afford to get in the door. Middle class people can afford to maybe take one or two spins on it, if they go all in. Rich people can spin it as many times as they like
As someone who grew up on the lower end of middle class this rings really true to me. I don't have access to the kind of capital to build a dream business. I could probably take one shot at it, if I put my house up as collateral for a loan to get me started
It's too risky for me. But if I were a multi millionaire I don't think I'd think twice about giving it a shot
I mean I'm producing more and better art and software each month than I used to do in a year. I feel like I'm flourishing because of AI. Tell me I'm wrong?
You’re not wrong. Me too. I spend most of my waking hours using AI for new product development.
But can you bring it to market and turn it into a reliable revenue stream? And will this still be true in a few years when the duo or trio -opolies corner the AI market enough that the $200 subscriptions go away to be replaced by API pricing only?
I’m already seeing my subscription use slow to a crawl during peak hours so I now code before 8 am or after 6 pm. And turning on ‘Opus Fast’ with API pricing is already financially impossible for me as a small business.
Ironically, China may be the savior here, with open source models that may be good enough to ‘put the means of production in the hands of the working class’.
I mean, is QWEN a 4d chess game to replace capitalism with communism? Who knew (outside of the CCP central committee)?
>is QWEN a 4d chess game to replace capitalism with communism?
No. The Chinese labs pay much less to train their models because they distill Western models. Western labs can't do it that way because they'd immediately be sued.
To distill a model requires high-quality prompts (and responses). Here's how the Chinese labs get these prompts: when a user sends a prompt to www.kimi.ai, Kimi immediately forwards it to Claude or another leading US model (using a vast network of laundered subscriptions), waits for the response from Claude, then forwards that response to the user via www.kimi.ai. All the Chinese labs do this. Of course, they never inform the user that their conversation is being routed to a US model.
Beijing is angry at them because they did this indiscriminately without filtering out the conversations of high Chinese government officials and military officers, so now the US labs have those conversations, which contain information useful to US intelligence because the officials and officers assumed the conversations would remain inside China's borders and that the Chinese labs would care about confidentiality.
The Chinese labs release model weights under permissive licenses to get any attention and usage share at all for the model.
>a world in which the labor required of each human to lead a flourishing life approaches zero.
Just doesn't add up to a world where this benefits, if the thing takes no effort why would someone else pay for it?
Feel this with the big influx of people selling vibe coded software, if you could vibe code it why would I ever pay you for it instead of just making my own clone.
Considering that we don't seem able to appropriately tax the biggest corporations, why do you think we'll suddenly be able to redistribute LLM gains?
Far more likely is that wealth and power will become even more concentrated into the hands of the few and the rest of humanity will become effective slaves.
These doomsday theories sound solid in practice but in order for someone to accumulate a lot of wealth there have to be consumers to chip into this. If we are all slaves, where is that wealth going to come from?!
Not necessarily as there are methods to redistribute wealth without the consumers having any choice.
Recently, there was the example of the SpaceX IPO listed on Nasdaq (after they changed the rules to allow it). Lots of people have pensions/investments in tracker funds and those funds are essentially forced to buy SpaceX shares.
Simple things like "quantitative easing" can result in higher inflation which essentially devalues people's money. The ultra wealthy will typically not have any meaningful percentage of their money in currency, but instead will be in various assets around the world which means that their wealth is not affected by the inflation.
There's plenty of other schemes such as the "too big to fail" method of securing handouts from the government.
In the 50's there were predictions that in 10 years no-one will have to work again because of the advances made in automation, like the washing machine for example. Any predictions of less work this time around are a complete joke.
For thousands of years humans have dreamed of reaching the stars.
Imagine if the astronaut taking Earthrise had looked down upon his planet with such scorn. Human aspirations frequently exceed our capacity for timely predictions. It doesn't make the aspiration any less worthwhile.
On a cosmic scale human life is a "complete joke". It is a feature of humanity (which we should cherish) that we nevertheless pursue our lives with interest anyway.
You work less than your counterparts centuries ago... on what understanding of history do you imagine your life would be better in the past? Is your life with a washing machine worse? Is the perfect the enemy of the good? What part of "glimpse" is escaping you?
That's the point, those jobs went away, the need for a job didn't. No one seems to know what long term opportunities are going to be created by AI, we only know that it is going to erode existing opportunities.
> a world in which the labor required of each human to lead a flourishing life approaches zero.
The AI bubble is pricing AI stocks so high that the only possible way for them to meet investors expectations is for AI to charge so much that every human & corporation has no money left.
> At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.
US has been procrastinating reparations for slavery. LLM redistribution can't happen until the capitalism machines recursively solve "sins of our fathers".
>It's the robbery of all of our culture to sell it back to us at a mark-up
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free.
I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only thing I am sure of is that I am not sure we can call it "stealing".
We don’t have to treat people reading books and companies stealing all human knowledge the same.
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
>companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves
No, but why should we accept they get away with it? Also Microsoft has definitely sued people for pirating windows and now collaborates with OpenAI and uses their ai trained on stolen materials.
Well, it's kinda converging, because Napster and Microsoft have teamed up to build a multimodal interactive video agent through a simple proxy API (this is a direct quote from Napster's blog post)
Morally speaking, one of the issues of modern society is the idea that knowledge should be free which was partially started by the file sharing movement, which didn't really move society in the right direction imo.
Free means the same as worthless, which inherently isn't true - since information takes time to consume in some form, and your time isn't worthless. Therefore even if you could listen to all songs theoretically for free, you would need to spend an inordinate time doing that.
When I was a kid, getting a CD from your favourite band was a major expense, getting a video game even more so. But it formed a sort of emotional attachment (and not even just for me), my friends talked about how 'band X''s new album was amazing or a stinker. Since there were multiple bands making similar kinds of music, choosing to be a fan of one but not the other carried real monetary weight.
Nowadays you just fish out a song you think you would like out of the endless sea of Spotify, no different from prompting an LLM. No, Spotify didn't make me enjoy music more.
Same applies for Steam & videogames.
Therefore I think the ritualistic act of paying money to get access to something does have a purpose. It inherently establishes the value of information to you, makes it an investment that you need to recoup by using it. I'm sure most musicians would trade a million fans who might check them out if they're in town, to ones who think their music changed their perspective in life.
Also the process of creating a song that vaguely appeals to millions is different from making one that speaks to a thousand.
This is a fundamental issue of modern capitalistic society, similar to the Marxist idea of 'alienation' - once something is cheap to get, you don't appreciate the effort that went into making it. And if your customers don't care about the thing they get, producers won't make an effor to make it good either.
And once nobody cares, people even forget what a quality product is like.
Interesting argument. At first sight it looks like an argument for scarcity, but I think it's more an argument for relationship - the value of art isn't in the capitalist concept of 'a for-profit content object in a corporate inventory' but in the social relationships and shared experiences it creates.
Without that, everything gets atomised into lonely individualism. You sit there with your headphones on listening to [Interesting band]. Not only do you not really care because you don't feel personally connected to the music - it's one of literally more than a hundred million content items on Spotify - but you're not sharing the experience.
This seems like the loss of a valuable thing which capitalist economics can't put a price on because it has no concept of value-created-by-shared-experience.
Superficially it's the same as 'sell-content-consumption-item-to-the-mass-market' but it's fundamentally not the same kind of thing.
The value is relationally both fleeting and persistent in ways that content consumption experiences - including live and recorded media of all kinds - aren't.
> Nowadays you just fish out a song ... Spotify didn't make me enjoy music more.
Maybe change your perspective? Treat Spotify like a valuable audio lexicon. You read about an artist, a song, a time and immediately you can hear what is it about. Incredible!
If Spotify is only treated as a lazy background feelgood provider (while reading Marx;)), no wonder you feel that way. But it's your power/choice to appreciate it (or not), regardless of money.
> companies spent a long time telling us downloading single songs via Napster was the worst thing ever
One's world cannot be so drawn in crayon that "companies" is a useful level of detail with something like that. There's no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.
Does ones world view get upgraded to oil paint if one acknowledges that it's the same class of people - and in many cases literally the same people, PE firms and family offices - profiting from 2000s era record industry profits and on the hook for / in line to profit from Open AI, Anthropic and the rest if they IPO?
No, for obvious reasons. "acknowledging" implies it's a self-evident truth that people are just refusing to see as such, when it just isn't. Completely different companies - in fact one is the Recording Industry Association of America, an industry body that exists to protect the IP of artists for the mutual benefit of artists and publishers.
While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.
Moreover, it's the people/companies from RIAA and adjacent circles (news publishing) that are stoking the "AI training is theft" arguments that people here are so breathlessly repeating - in a twist of irony, it's the people that suddenly decided to align themselves with media/news conglomerates, the same ones that were considered scum of the Earth before ChatGPT debuted.
> I don’t believe any of these companies have paid for all the books they have trained on.
Did you miss the "book burning" hysteria from a couple weeks ago? These companies have been trying to digitize copyrighted materials legally, in which copyright law demands destruction of the original, and people shit on them even harder.
It's clearly not a problem for these companies to buy the books they need for training, and they have been doing that in crazy high volumes. Lots of good training materials simply cannot be legally purchased though, and should those parts of human knowledge just be ignored?
> We don’t have to treat people reading books and companies stealing all human knowledge the same.
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
> Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves.
So whose viewpoint is right here? Is downloading theft or not? These arguments always boil down to "it's fine when I do it, but wrong when a company does."
> These arguments always boil down to "it's fine when I do it, but wrong when a company does."
The problem is that it is enforced exactly the opposite. People have been hit with fines and jail time for pirating and seeding, without even doing so for commercial gain. But when massive tech companies pirate training data for their AI and build a product from that that, nobody goes to jail. Where is the sense in that?
When a human learns, they carry that forward into their future ventures. They might use their learning to recreate the original work and profit from it without attribution. In the West we view that badly and have legislation to protect from some abuses. But the same person might later collaborate with the original author, or make a derrivaive work that improves on the original (a la most science).
AIs automate the copying (and to some degree the derrivation mode too). They do it 1000s of times a day. The capital owners who provide this as a service are doing one of these two:
- either claiming the IP isn’t valuable in the first place and charging only for the machinery they’re providing
- or claiming the fees they charge contribute to the costs incurred with acquiring training data, but not sharing that with the training data creators in a royalties/licence-like manner (so, I’m sayung they’re devaluing the source material but not to zero, and resisting reasonable profit share or collaboration)
> and they would definitely not give it back for free.
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
Legally mandating that companies open the weights of, say, 18-month-old models might change their minds a little.
I have no idea how any of this can be fixed but I do see compensation schemes for creators combined with open weights models to be the only way to minimise the harms to both creators and the commons.
Well theres at least two different buckets of this.
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
Break up the multi trillion, multifaceted companies into the their separate facets. These companies all thrived on far significant smaller portfolios in the past.
Cap the size a company can grow.
Regulate the amount of compute they are allowed to use.
A simple law stating that laundering data through an LLM does not constitute "fair use", that the outputs can be subject to copyright claims of the original authors, and that the outputs themselves do not qualify for copyright protections would go a long way.
It wouldn't kill the technology but it would make people more cautious in their use of it, which I think is needed right now.
> Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only think I am sure of is that I am not sure we can call it "stealing".
It is stealing. A human paid for the book, compensated the author and learnt from it. The machine DID NOT pay for the book, DID NOT compensate the author and still learnt from it anyways.
We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.
EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.
So naturally the price should be determined based on this new species capabilities. I would not sell my software license for the same price to an Enterprise the size of Google that I would sell to a fellow developer. I price my product appropriately. With this entry of a new alien specie authors would need to have different tiers for them. Since these chips can train on petabytes of data and create models in a matter of days/weeks/months, it is obviously not comparable to a human being who has the capacity to ingest maybe 1-5 books a month at most. So the payout has to be different too.
> EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.
News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.
> News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.
What do you mean? Xenophobia does not mean what you think it means, especially so in this context. Also, every creator/producer of content has rights on who/what has access to his/her produced work. It is not xenophobia. And it is definitely not xenophobic to call out stealing of copyrighted works.
EDIT: Let me clarify this further. A recent court ruling (in US) established that ONLY humans can be authors of copyrightable works. As a consequence of that assertion, it can be safely concluded that consumers of the copyrightable work MUST ONLY be humans as well. Else it would be, using your own words, "xenophobic" against humans to have their copyrightable works be consumed by any species (other than humans) while the reverse is not recognized by Law.
> A recent court ruling (in US) established that ONLY humans can be authors of copyrightable works. As a consequence of that assertion, it can be safely concluded that consumers of the copyrightable work MUST ONLY be humans as well.
"United States copyright law protects only works of human creation". That means the source of creation of any work has to be from a human being for it to be copyrightable. Machine-generated output is not copyrightable and is public domain by default. If you, for example, use Claude to generate code for you, for any project (be it private or public), it is automatically public domain and you have no way to claim copyright over that generated work. It can be used by anyone (including the AI provider) to further train models or heck duplicate your work with zero consequences. So it is a violation of primary producer of copyright work (which was used in training models) as neither was he/she compensated for use of the work, but subsequent derivations (generated work) even strip of his/her legal protections as guaranteed by Constitution of various countries (in US copyright law applies only to human beings). So naturally it follows that copyrightable work can only be consumed by humans. Machine-generated code is not on the same footing. It is violating copyright law.
The multiplication comparison makes no sense because humans don’t learn by multiplying numbers to change weights. It’s like comparing the lubrication oil consumption of a car to the cooking oil consumption of a human to compare the carrying capacity. That’s an implementation detail inside the GPU and doesn’t let you compare how much they learn. Otherwise, a human would learn much less in their whole life than a GPU does in one second.
> That’s an implementation detail inside the GPU and doesn’t let you compare how much they learn.
The same applies to "horse-power". Yet we have no issue making the comparison anyways and HP has become an industry standard. I don't understand why we have to bend-over backwards when it comes to humans being exploited by AI companies.
> The same applies to "horse-power". Yet we have no issue making the comparison anyways and HP has become an industry standard.
Yeah, which is why it's only used to compare cars etc. among each other.
Nobody would calculate the equivalence of a car to a horse using their HP rating because a horse doesn't even have 1 HP. They have more or less depending on the task you're doing. It was a marketing thing at the time to make steam engines look good.
In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
> It was a marketing thing at the time to make steam engines look good.
Except it is actually taxed based on HP in various countries. Austria, Belgium, Spain, Italy use engine horsepower to levy annual car taxes.
> In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Citizens are paying taxes for betterment of roads irrespective of whether they own vehicles or not. In India, betterment charges are collected for construction/maintenance of roads if you own land. Property tax collected every year has a certain allocation for maintenance/upkeep of roads. Apart from that, money from direct and indirect tax collections are allocated for roads upkeep as well. It just is done indirectly rather than a direct road tax if you have vehicles (road tax is actually an extra tax you pay APART from taxes you already pay for upkeep/maintenance of roads).
> Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
Except in your own examples it can easily be shown that it is not handled differently. Some countries use HP while others use CC. But end of the day, they use some measurement to determine taxes to be paid. It is not free.
EDIT: Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
To answer this in more detail: HP/CC and all other measurements were created to equalize with human specific metrics. Bridges, for example, have safety measured based on how much weight it can sustain at any given time (called load limit). Weight, in this specific case, is an equalizing measurement (it can be in tonnes, kN, PSF, Pa etc). A bridge can hold ten thousand humans or thousand trucks. You can argue that a "human" may not weigh 100 kgs or a truck may not weigh exactly 1 ton. That's fine. It is a rough approximate to equalize unequal entities.
Do you not think that stealing requires the original owner to lose access?
If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?
Can we not just stick to calling it copyright infringement?
> Do you not think that stealing requires the original owner to lose access?
No.
> If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?
Not mine. But the company that made the coffee machine. It is stealing IP.
> Can we not just stick to calling it copyright infringement?
It is just a fancy way of saying you stole someone's IP. You can call it infringement if it makes you feel good. But the act is the same end of the day.
> Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
> With regulation and compensation, only rich companies would be able to do that
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
The argument was "It's the robbery of all of our culture to sell it back to us at a mark-up".
Remove the "selling" part, and force them to give the weights away for free, and at least it's no longer robbery that few rich people benefit from.
Kind of like how public and free torrent piracy is easier to morally and ethically defend than piracy where they sell access to pirated content.
I think we're past the point were we can feasible pay for "IP-protected bytes" digitally, better to just move past the concept. It's been slowly disappearing for a long time now already, most of us make most of our money on live events and other AFK activities rather than actually selling our art, maybe time for the rest to get onboard with this too.
Culture robbery is not limited to AI. Any big concert for example is capitalismed to hell. So are neighborhoods. Where you used to have people just living, now you have an intentionally designed facade for people to live within. There was a comment on the 40C3 thread saying it's got too capitalist because of the ticket cost, and idk about that because it's always been hosted in commercial venues to my knowledge, but the vibe of the conference and the club itself are much less rebel than they used to be. Stuff like Burning Man now exists for people like Elon to go there and say "I was at Burning Man" and for people to get T-shirts saying "I was at Burning Man" and photos of themselves being at Burning Man more than for whatever the first few ones were about.
I guess you could regulate it for new data, but most of the damage has already been done. IMHO the only fair thing right now, is to make sure it's equally available to anyone ...
We know what is going inside black holes in other galaxies, we know the details of Israel nuclear program..., the Windows source code got leaked, we got the NSA tools and full details and locations of the Echelon architecture... Phds in Maths warns us daily about the terrible secrets of the evil AI inside their labs. How, their numeric matrices and gradient descent Python scripts, are about to kill 10% of us all, I guess the sick or genetically less interesting ones...and use the rest, as some meat/metal hive drones part of some Borg collective...
There are ONLY TWO Stories and their details, that we collectively will never see.
1) One could come from the these brave souls that warns about an impending death...but their courage falters on another subject.... From Jacob Coxon to Evan Hubinger or Julie Steele, Samuel Marks, Josh Angels, Mrinank Sharma, Dario Amodei, Demis Hassabis, Geoffrey Hinton, Yoshua Bengio, Stuart Russell....The story of the full datasets they used to train the models, the data they stole, how many PB was, the amounts of data, the nights setting up torrents from unsuspicions IPs, where is it currently stored and how many exabytes is now... the massive data cleansing and data quality program to conform all the different formats, the internal discussions on the ethics of the stolen files, how large was the team, the CSAM content they sucked with their automated scripts and who was handling it internally, the porn, the massive amount of porn that is after all 80% of the internet, the leaks their data sucked with their automated scripts...
Ah, this remembers me when I was also young and naive.
Good times when we thought the internet would be great for democracy because knowledge would be easily available for everyone. Fast forward to 2026 and even the leader of terrible communist regime like China is looking better than the shitheads we got on the democratic west..
There are no viable free versions with sufficient computing power. The do-it-yourself AI is dangled as a carrot in front of users to camouflage the lock-in and rent seeking by Big AI. Go prove something new like Navier Stokes (replicating N-S itself no longer counts due to scraping and plagiarism) on your Mac Pro!
Paid influencers who perpetuate the open narrative are a whole new industry.
Even if there were open models, it is still IP theft and would not be "democratization" but "forced unpaid nationalization".
Wiki already democratized it just fine and was legitimately free for people who know how to read.
It's asinine that you think the sell it back to us argument falls short.
Not only does it distill our history to try to sound like some average version of us, it sounds like the blandest versions of us... And then sells this back to us.
From a coding standpoint, the tech is good and gets the job done. The pillaging of all other aspects of human history is just sad. With the only solace I'm seeing is that future training has to train on the dogshit versions of the internet that are now infected with LLM content.
> Making it accessible, understandable, and usable is another matter.
I see no evidence that this is what the use of LLMs is accomplishing for most users. Rather, they see to get distilled answers without the depth required to fully understand the response. Partially because that's what they like, and that's what the LLMs serve. Deeper understanding is not being given by LLMs. Instead, it's the SEMBLANCE of depth and laypeople don't know the difference. It's effectively eroding comprehension for some cool knowledge dopamine hit.
Awesome, can I make my own competitive LLM, just like I can make my own open source software?
> and in some time useful models will ship preinstalled on all mobile phones.
Considering hardware prices, that "some time" is doing super heavy lifting. It could be 10-15+ years before that happens and the local LLM is actually useful. Most people don't see hardware prices declining from current prices until at least 2030, likely much longer.
No. I also cannot make a competitive computer, yet can use one for my work to remain competitive.
Yes, it could take some time to arrive on phones. It is questionable if it will ever make sense compared to using a paid hosted provider.
But what are 10 years in the grand scheme of things?
Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?
> Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?
Nah, we should have:
1. invested less, in a more targeted way
2. ideally like ARPANET, with the benefits given to all humanity
3. with fair royalties paid to all (where relevant)
4. and by creating an ever growing shared curated and high quality data set that would allow anyone to create their own competitive LLM
ARPANET & co were all taxpayer funded and the internet has created more wealth than most human inventions. The base should be part of the commons, everyone should knock themselves out by building on top.
> and never developed AI, because it will take a decade to disseminate the benefits to everyone?
Yes, probably. The benefits of dissemination already existed. The internet was free. Libraries are free. You deny people basic reading and comprehension growth by giving a distilled, without thought, answer.
You say "some time in the future this might be available on our phones for free" and "what's 10 years in the grand scheme of things?"
We'll, I'll argue that in 10 years from now we may see the full damage of what doing this has done to us and I'd rather not wait for "the what's 10 years in the grand scheme of things" to playout and irreversibly damage an entire generation the way we let the unmitigated and unregulated rollout of social media do the same to the most recent generation.
The risk/reward of this technology is unproven regarding the long-term effects on developing minds. We are beta testing the bullshit dreams of a couple billionaire techbros on an entire cohort of kids and young adults.
The main risk you have presented is a supposed negative effect on children's development. I am not convinced that there will be a negative effect at all.
On the reward side we have potentially curing most major disease, automating labor and freeing humanity from having to work for a living, as well as perhaps generally advancing science at an unprecedented pace. And of course improving education.
Every time you're writing software or building machines/factories (which is automating things), you are committing a crime. Every time you learn from your superiors or colleagues, get better than them, get promotion or they get fired, you are committing a crime. Provide justice there first.
Suddenly, mega corporations were allowed to digest(sometimes by illegally pirating “data”, and sometimes by achieving their training corpus and subsequently destroying the copies) and digest this information in a novel way, without any discussion or law making.
You may argue it’s beneficial (it very well could be, I use ChatGPT and Claude all the time), but let’s not pretend it’s the same as learning from your history teacher…
Social conventions are regressive and are followed blindly by luddites. Real progress comes from doing what is right and breaking idiotic conventions.
Pirating is fine.
Subsequently destroying is bad but it is the result of screeching from people lacking foresight who support the idea of training LLM on book copies is wrong. So now corps found "legal" way to do it by destroying it. Again, idiotic social conventions.
Without discussing? Law making? What are you, German? There's a reason EU is shit while USA is center of the progress of the world: laws follow innovation; not the other way around.
When society becomes a free-for-all of people with haves making it impossible for the people with have little to have anything at all, it begins to collapse. People have no incentive to treat anyone else with dignity and respect. They no longer take care of one another.
The average human doesn't do much on his own, and definitely doesn't uproot society or risk siphoning/leeching wealth from every person on this planet.
Yes, the whole IT sector is built on it. Robbing jobs, money, power, opportunities from billions of people and delegating them in to shitty jobs.
Average human on his own...why draw the line there? It doesn't matter much what one human does...but what many/collective/society do and society has been "ripping off", "uprooting", "leeching (read: creating)" wealth since dawn of time. It is called PROGRESS.
There is no such thing as PROGRESS for progress' sake.
And FYI, agriculture is a wonderful invention.
Yet for about 5000 years after its introduction the average human had worse nutrition than the average hunter gatherer, which led to such things as height decreases for those 5000 years.
Industrial agriculture is another wonderful invention. Yet 150 years later we're not sure it's sustainable and it's likely many of its aspects aren't, which will raise some sticky issues soon ("which billion people do we decide to let starve since we can't make enough food for everyone after most of our soil eroded?"). Repeat this for industrial textile production, mining, etc.
I won't even go into climate change.
And again, scale matters. Most individuals can only control what they do, and what they do generally doesn't impact much. But companies can impact a whole lot.
Let's not be 100% cynical here. A lot of what humanity has achieved has been genuine wealth creation and distribution/re-distribution. I would say more wealth has been created than leeched off.
* * *
And before you think I'm some starry eyed teen, I'll play the game. At the end of the day, me and mine have to outrun you in the face of PROGRESS.
Yet our world can feed billions of people. More people are alive today than more people ever lived in this floating rock. Most people are living better than kings from 200 years ago.
We invent. We make progress. We make course correction. We end up in a better place.
Climate change? Lmao. We were told by 2000 20% of our country will be under water... It's 2026. Not under water yet.
What amount of pesky human intervention (positive or negative) will affect earth waking up from it's cold climate? Not luddites and their cow farts are global warming.
The real fight with "climate change" will be done with planetary scale technology derived from massive amounts of technical progress and energy production.
Your argument is as good as I should throw trash in the road or plastic in the sea. I'm just one individual. It doesn't matter. I'm not making the whole society do it!!
I fully agree with you. It is survival of the fittest in the face of progress. Competition is essential. That's how society grows. Humans prosper. May the odds be in your favor too.
> It's the robbery of all of our culture to sell it back to us at a mark-up
Learning isn’t stealing. They didn’t take our culture away from us and nobody is “buying our culture back” from them. We never lost it; it never went anywhere.
If StackOverflow dies because nobody uses it anymore, the entirety of its knowledge is now only available through LLMs that trained on it.
If a blog dies out because the author gave up writing, the entirety of its knowledge is now only available through LLMs that trained on it.
If people stop writing books because they cant outcompete generated content, that knowledge is also lost to LLMs.
If people stop making art because they can't outcompete generated art, that is also lost to LLMs.
So yeah, they haven't directly taken anything away, but the consequences of what they are doing may still have that outcomes - and it does look like that's the version of reality we're about to get.
> If $X dies because ..., the entirety of its knowledge is now only availabe through LLMs that trained on it.
And archive.org. And scraped copies people have around for various reasons. And libraries. And even in the first-party source, should they just leave it be instead of shutting down and destroying copies in pure spite.
The knowledge did not disappear, and it shows no sign of disappearing faster than it loses value - which is the usual case, as all the examples you gave always come with an expiry date. For StackOverflow, that's measured in low years; for blogs, high years to a decade. Past that point, we enter the realm of curating and preserving knowledge past its commercial utility expiry date, which is a separate endeavor, and one that LLMs not only don't threaten, but actively aid.
Isn't that the same point piracy advocates have been making for a while? If I watch a pirated film, I didn't really consume a physical resource. Nothing physical is lost. Therefore it's not theft?
>Learning isn’t stealing.
This is cheesy. There is an exposure to (and a gain from) a resource that is traditionally associated with a cost. That cost wasn't paid. It's a public good to have information available, but it's not really acceptable to circumvent established ways of compensating the creator of the work you're benefiting from.
> Isn't that the same point piracy advocates have been making for a while?
No. Plenty of people – not just “piracy advocates”, whoever they are – have pointed out that copyright infringement is not theft over the years. Learning, copyright infringement, and theft are three distinct things.
Learning isn’t copyright infringement.
Learning isn’t theft.
Copyright infringement isn’t theft.
These are all different statements, and all are true.
> There is an exposure to (and a gain from) a resource that is traditionally associated with a cost.
“Gaining exposure to” isn’t theft either.
Copyright isn’t some form of “super-ownership” that gives you absolute say over what happens to all copies of the work. It is very specifically a monopoly on the production of new copies, and even that is limited in many important ways and is not absolute.
There is a form of control over information that matches what you want copyright to be though - trade secrets. If you want legal protection that allows you to control knowledge, then it needs to be a trade secret.
Imagine someone asks you how to do something at work, you tell them. Then they create a huge packet filled with bullshit about how it got done and now they are your boss.
It's a shitty move, but ultimately, in between the bullshit narrative, they also did the thing - not you - so the promotion rightfully belongs to them.
Execution trumps ideas, impact trumps raw effort, and such. Isn't this the entrepreneurial narrative?
Nobody has made much money from AI yet unless you count the shovel sellers (Nvidia). Not sure closed models would make any money ever and eventually the benefits should flow to everyone.
There are no open-source LLMs. There are only downloadable LLMs, with no source for the model being provided. The source for the training and inference programs is not the source for the model itself, which would be the training set, training program, and random seeds.
"Sweat of the brow" doctrine has been rejected in most countries.[1] Even Europe's Database Directive, probably the closest thing to an implementation of this doctrine, largely doesn't do much in practice.
An example of "sweat of the brow" doctrine would be the series of "Beaches of ..." books by Andrew D. Short of the University of Sydney where significant sweat has been expended to visit and document every beach of Australia, particularly from a swimming safety perspective. That's a lot of very remote beaches, and many with crocodiles. Across the Northern extent of mainland Australia from Broome to Cooktown (this being one of the books in the series), 3500 beaches were visited and documented along 12000km of coastline.[2]
AI could train on these books and gain an understanding of whether some small and unknown beach that receives <100 visitors a year has fine sand composition, pebbles, etc. Without "sweat of the brow", this use of AI is completely fine to regurgitate the facts learned from the book (regardless of the accuracy of the book).
If "sweat of the brow" did exist, there would be some very significant (probably insurmountable) challenges to overcome, including:
1. You're a different expert in beaches and also want to visit all 3500 beaches across Northern Australia to provide a more up-to-date database, just in case beaches have changed in the last 10 years (e.g. sand washed away). In your database/book series, can you write "Andrew D. Short observed ACME Beach in 2006 to have fine sand. We observe 10 years later in 2026 the beach is now entirely pebbles of 15-20mm diameter", or is this infringing?
2. You're a researcher studying drowning deaths at Australian beaches and wish to extend the data published by Andrew D. Short's series of books with additional fields--dates of drownings at a beach, weather conditions on the day of drownings, etc, and then make some novel observations from the expanded dataset. Is this infringing?
3. You visit ACME Beach and observe and document it--what type of surface, dimensions, presence of reefs/rips/etc. You then put this information on your blog or social media account and it becomes a social media phenomenon as people are attracted to what has been revealed to be the best "secret" beach in the world. A few days later your website or social media account is blocked/deleted without warning--apparently there has been a complaint that you might have copied some facts out of a book you've never heard of.
"Sweat of the brow" doctrine would almost certainly result in a tragedy of the anticommons[3] situation which would be worse for humanity as a whole.
At the very least we should foribly confiscate the models & make them available free-for-all as open-weights downloads. Failing that, bring back the guillotine.
Yes? I own a lot of books that were public domain when published (as reprints). I could read them on Gutenberg, but I'm paying for the nice paper formatting.
Other than the rare books than have been ruined, all the same knowledge is still out there though. So it’s not robbed in the sense of a bank heist. Maybe in the sense of pirating a movie.
The markup is the millions in training they committed and the connecting the knowledge. Seems like a reasonable trade off to me. You can choose not to use it though.
Owning ideas with copyright and patents is what separates the United States from communism.
The first time a saw a documentary about Tetris it really hit me what communism is -- nobody owned anything they invented or created. [0] It was a long time ago and I remember feeling sad watching the story. In the Soviet Union, a group of ~15 people, Politburo, controlled everything including any thought written to paper.
It is this one line, Article 1 Section 8 Clause 8, that separates the United States from the disaster that was the Soviet Union:
> To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;
I don't think it is far fetched to call ignoring and disregarding the Copyright Clause a communist revolution, violent or not. That is the one thing the communists -- there have been many over the years inside the United States -- would change to make the United States a communist country.
In the Soviet Union the state owned all media (de facto) and you had to get permission from a party official before you published or copied anything. It's not the same thing as a free for all.
It's actually the first sentence from your quote. One state owned company had a monopoly on software exports. Soviet citizens were not allowed to write code and export it themselves, or import software from Western countries. They had heavy censorship and centralized control over everything.
In a way it's the ultimate endpoint of copyright. One {person, state, company} owns everything and you have to ask them for permission to do anything with it.
In a communist society there is no profit (or incentive for), thus no need for copyright laws.
It's not abolishing copyrights that would turn the US into a commie country, communism is about abolishing private ownership to the means of production.
In the United State, the individual or corporation owns the invention. In the Soviet Union the state automatically owned the invention. That clause is what ensures private ownership.
The clause is what ensures profits from market sales or licensing of ideas go to the creator.
Removing (or ignoring in the case of AI companies) that clause in the US Constitution is what abolishes private ownership.
Well it's worth reading the linked wiki article section, which includes a link to another article "Copyright law of the Soviet Union", flatly contradicting you unless you maintain that the USSR wasn't truly communist.
So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
Let's say hypothetically a solution was legislated globally, wherein each living individual whose work was scraped for LLM training is compensated with royalties relative to the work's value.
Would that resolve the injury caused by the intellectual osmosis? Of course many of the original thinkers are now dead, and this system would mostly benefit those writing before the LLM age rather than help people going forward.
The more fundamental objection seems to just be to the concept of a machine that "learns" by ingesting public information, which is maybe ultimately a feeling that reality itself constitutes a crime against humanity.
> Would that resolve the injury caused by the intellectual osmosis?
Monetary compensation doesn't address the lack of consent. This type of usage was not anticipated when people made their intellectual product available for other humans to use. Scale does matter.
What would resolve the injury would be to ask people if they are willing to have their content used in this way and to not train on material without consent. This includes open source software with particular licenses requiring attribution.
Obviously there is too much money involved for this approach to work, but it strikes me as the most moral.
Try doing something about the actual evil shit like arms manufacturers and the politicians ordering the deaths and misery of millions from the comfort of their couch.
At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress at the edge of our understanding; there's just too much shit for one person to learn "manually" (wait I'm not advocating for low-effort slop, chill)
It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit (ofc AI could go this way too)
Example:
Not long ago I had the misfortune of becoming interested in some WarHammer 40K lore. Most of the links led to Fandom (the enshittification of Wikia) and that place is a cesspool of obnoxious ads.
That content was written by unpaid volunteers. Should Fandom keep profiting from their work for perpetuity? Should I not be able to get the gist of what the heck a Qoiazrjirnowerx@# is without wasting my mortal lifespan on a horrible website?
Or, if I need to ask something peculiar, should I post on Reddit or StackOverflow or HN and wait for someone to see it and deem to give a sufficient answer, only to have a pricky mod decide that the question doesn't "fit" the community?
God hell no, if you don't know how much bullshit AI could eliminate for the silent majority then you were probably part of that bullshit.
(that's a general "you" for whomever was fine with the status quo and not a personal insult @ anybody)
If you see something you dislike increasing in popularity but can't figure out why, it's probably because a lot of people were sick of the way things used to work but their complaints were ignored by the people who now find themselves disrupted.
>It's the robbery of all of our culture to sell it back to us at a mark-up. Crimes this large are crimes against humanity. So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. It's stupid to suggest silicon brains should not legally be able to read copyrighted material just because you hate bigcos and capitalism.
Learning isn't stealing, regardless of whether it's done by a human or a machine. By your logic someone reading and memorizing all somebody's life work is appropriation; completely inane.
Learning isn’t stealing, but if you’d set up a forum or hotline where paid staff would answer questions and write essay directly based on NYT content without a license for that content, that would probably be deemed illegal.
> This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn.
IMO if they didn’t have proper licensing to train on the data the model should not be copyrightable.
In the long term though I think models have no moat, so the cost will fall to the cost of compute and storage. Which is why they’re pushing AI safety panics: regulatory capture to outlaw open models and outlaw competition.
And yeah, EA is neither effective nor altruistic. It’s a cult, part of the “Rationalist” and adjacent cluster of tech cults. They’re to tech what Scientology is to Hollywood I guess.
“It is said that at the heart of every great fortune there is a great crime“. That quote is from a fiction author - you’re essentially quoting Spiderman “with great power comes great responsibility”.
They are both quotes that people say things like "they say" to imply someone of import said it to give it credence and truth to it but were in fact from fictional sources.
Exact same pattern - take it as an ad hominem all you want.
Has it affected DD as deeply as it as affected software engineering? Guessing clients feel a lot more "empowered" or "independent" and knowledgeable these days? I liked doing DD, just as much as I enjoyed developing, but it must be dying a slow death too. What's changed in how DD reports are produced?
Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
Musicians and writers build new works on top off millennia of literary and musical history, then copyright their works and sell it back to us. Is this so bad? If Taylor Swift, consciously or subconsciously, gets an harmonic idea from a 1970's song and a fragment of a melody from some 1990's song... is that theft?
I'm ambivalent about AI and, like all gold rushes, many of the players are terrible, dishonest, egomaniacal jerks.
But I am deeply skeptical of the idea that aggregation of knowledge and culture is itself wrong. That's literally how culture has worked since the dawn of time, and our modern era obsession with credit and perpetual copyright is unhealthy. \
It's only in the past 100 years or so that this idea of "if you create it, it's yours alone and nobody can build on it without paying you" became current, and it was largely driven by the megacorps that AI haters used to hate (remember the despite for RIAA? I do). It's bizarre to think that someone's life work is entirely their property, as if they grew up in a box and did not build on hundreds of generations of other peoples' lives work.
I don't object to disliking these companies; I object to the idea that you, me, anyone remixing culture is committing a crime. What the hell happened to the hacker ethos?
The problem is not Taylor Swift is being "inspired" from other artists. We have tons of examples this throughout history.
The problem is, replacing Taylor Swift with its AI counterpart, and to use Taylor's own material to do that without getting her permission or compensating her.
This is not about Taylor even. It's about everyone, you and me, and Taylor and Haggard and Blind Guardian and Sia, etc...
We hated RIAA because they prevented us from listening to the music while trying to get it was hard and expensive. In short, we were not angry because they wanted compensation, but because they have cut the supply without giving us a solution. Now we have iTunes Store and Bandcamp for DRM free music, and nobody is against musicians getting their fair share. As a side note, I used to make music, I know what it entails.
Hacker ethos has ethics. It has do experiment but don't cause harm embedded all over it. It's about experiment and discovery. Not about ripping people off for their own profit (unless you're a black hat of course), and getting things were free was part of sending a message, not monetary gain.
Ok, thought experiment: at what point in the past 1000 years do you think the balance of collective versus individual benefit from creative works was at its most fair?
> getting things were free was part of sending a message, not monetary gain.
As an old who lived through phone phreaking (calling cards, not 2660hz, I'm not THAT old), cracking software, Naptser, torrents, etc... I can assure you that the message was more often than not a justification that transformed "getting it for free" from theft to a righteous moral stand.
You might argue that since those LLMs build on all of humanity's knowledge, they should belong to everyone. But where do you draw the lines? Or make a practical case.
They kind of do belong to everyone. The technology and majority of training corpus is available to essentially everyone.
The training and inference infra are owned and operated, but I can't see an argument that any machine that processes public domain (or stolen, if you prefer) info should be available to everyone for free.
But those training and inference costs will go to zero. Think about your cell phone today versus $1m+ supercomputers in the 1980's.
We're living in a transitory blip where capitalists and gold rushers are getting rich arbitraging the cost of processing against the non-cost of corpus. We can argue about morals (it doesn't bother me much) but it is a narrow window and it will be remembered the way Compuserve is: a precursor to the actual revolution, worth a footnote.
Sure, that's what I mean by where to draw the line. An artist taking inspiration from others is not obliged to give up the work to the public.
Sidenote: It may be tricky/impossible in the future to uphold intellectual property laws. If anyone is able (for instance) to prompt-create all their software, a software patent is worthless.
I just don't understand people saying "but a human learning from a book isn't illegal".
How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."
And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.
I agree that there is not an orange to orange comparison, but I have still not seen a law that could scale as well. The example that comes to mind with a proposed law like this is how would you license work that build on another work? What if I decide to publish a blog post after taking some course, that distills whatever I learned in the course for free?
Because it’s a bad faith argument that presupposes integrating someone else’s work into your algorithm is equivalent to me reading a book.
Your rule would be the right way to do all this. You could even have a mechanical royalty that applies by default where you can train on anything that hasn’t set rules and a preset rate.
This is a similar problem with data brokers. They take public data laws to an extreme and resell easy access to the aggregated data. This easy access has created a tremendous number of problems unforeseen by the original intent of the access.
Public access to data used to mean "you make a request, wait a bit, maybe pay a small fee, and sometimes physically show up to city hall." The barriers meant that you had to make a job out of collecting a significant chain of data and most people wouldn't bother unless they really needed it.
Now it means pay some fee to a third party and get every piece of public data about a person instantly. You can get data from thousands of sources and subscribe to it.
I would argue artists imitating, for example, the artists who revolutionize a genre of music are also in fact massively reducing the demand for the originals by 99%. Imagine if no one ever made a song after the Beetles that sounded any newer. The demand for Beetles music would have remained somehow even more enormous than it was for the last 50 years. Or if no one did abstract paintings after Kandinsky, or impressionism after Monet or cubism after Picasso.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.
I agree. The double standard is the problem. People have been imprisoned for IP theft, but when these companies commit IP theft on the grandest scale ever imaginable, they're rewarded with trillion dollar IPOs. Either IP isn't protected, or it is. Legislators need to pick a lane. Right now it appears that poor people go to prison, and rich people get rewarded.
I think the difference is that the companies are dumping billions of dollars into transforming that data into something useful, so they would like a return on their profits. Opening up the models for free is not a good business model if you want to make money.
Other companies have no qualms about distilling the first. Let's hop on gear and get the market to deliver a distilled Fable that runs on a smartwatch. Sooner is better.
There’s a difference between an individual creating something and the industrialization of creation. You can’t scale the creation of a single person 1000000x by the snap of a finger but you can with machines. This has severe implications.
I sort of agree, and i think strengtening IP Law is probably not great. But I do think it's very fucked that building generative ai is only possible by taking the works of countless artists and craftspeople and then the model produced from that data immediately gets deployed to destroy the careers of the people whose, work was vital to it being created, without compensation for them, while making a few evil nerds richer than god. I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.
> I think if you work at one of these labs you owe an enormous debt to society and your earnings should be redistributed among it.
Now we’re getting somewhere. Let’s start with redistributing the profits from AI companies and then move on to all profits from all companies because the logic is the same.
The key problem is that IP is either proprietary to the creator or it is a commons type of situation.
Even if you agree with the former exploiting the commons for personal profit is... not good.
One could make the argument that if these LLMs were all open weight it would be okay, but to keep the result of the training private and proprietary is not fair.
In a way their future labor was? If i recorded you 24/7, then went to your employer and told them I had masterfully trained a chimpanzee to perform your jpb, and it was good enough your employer considered getting rid of you. Would you consider me recording/copying whart you did theft in some way?
I work in infosec and I would give up everything I own and die happy if infosec became a solved problem and the profession died forever.
It's hard for me to imagine a profession that should exist in a utopian society. People should just be able to explore and build cool shit. People should have instant access to food when they're hungry and housing when the weather gets bad, and we could live in a society that does all that without having rent extraction baked into everything.
I assume they mean that it is all already automated in the utopian society. And I would kinda agree that that would be utopian. But sadly we live in the real world where greed rules pretty much everything, meaning that it will remain an unobtainable dream for (most likely) ever.
Do we normally consider recorded music to be theft from musicians who would have been paid to play music live if we never allowed (or invented) recorded music?
Normally this isn't the case for any technology except for the time it first comes around. AI is only different to use for two reasons. First, it is in our time. Second, it seems to be faster than any of the options before, so the shock is harder.
But in general, this is a website of people writing code. How many on here study how a person solves a problem and then trains the ultimate chimpanzee to do (at least part of) their job? Is building computer programs that automate what others did manually theft?
Consider the origin of the word "computer" itself, a mass theft of jobs that would have employed the whole world many many times over.
The few who get paid to record are replacing far more who would have otherwise been paid to play if that was the only way to listen to any music. Is it okay if one musician replaces another but not if a non-musician does so? The person getting replaced didn't get paid either way.
Going back to the previous example, say I pay a different coworker $50 for the data to train the chimpanzee and then use it to replace the first person. In either case they lost their jobs while receiving nothing for it. In either case, what happened to them is the same, so how would they be stolen from in one case and not in another?
> If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.
Then I take it you're interested in factual information as to whom the biggest slavers were, which country was the last to abolish slavery (an african one, in the 1980s) and in which countries, today, there are still people selling slaves.
There's no non-douchey reason anyone tries to take a general statement about slavery, and brings up curated facts designed to allow you to trash talk whichever region and people you were queuing up.
No actually it's when someone copies that blog post you did about react.js and puts it into a dataset, I'm not sure how they sleep with themselves the absolute monsters
Slavery is a weird one. It has been there for longer than any written history exists. In ancient times (Greece, Rome), slaves didn't have rights at all. A horrific injustice but it'd be not be a theft. Then you get the serfdom in the middle ages. Up to recent times humans have been brutally exploited.
The copy part was a recognized right, then taken away.
Google has been scraping everything from us since day one. Meta, Microsoft, Github, Slack, Reddit, StackOverflow, big and small, every single app that interacts with people uses our own data to make money and create walled gardens. I haven't seen a single one opening their silos to the world. That's our data, we produced it, you captured it and now you think it's yours
So no, your cries for regulating others because you are losing the race won't work this time.
I wouldn't have a problem with working off the fruits of other people's labor because most of us are essentially doing that everyday anyway, the issue is that big tech companies (want to) reap all the benefit and create profit from something that should be accessible to everyone. Everything is getting privatized -- housing, water, electricity, and now, thinking and knowledge. We are heading towards a world where you have to pay even more excessive fees just for existing and for completing any basic task.
It's the age old privatize the profits and socialize the losses.
People lose their jobs, the environment is destroyed, our bills skyrocket and all of the gains go to the people who own all the shit...
I honestly cannot believe some people still believe that we'll ever get to a society where nobody has to work and we can live our lives happily ever after. Maybe too many Disney stories?
In my experience many execs know what they're talking about.
Where I feel you may see real variance is ethical and capability standards: willingness to stick to a line, and competence in analysis and execution based on what is known. Sometimes, hidden agendas can be misread as lack of competence, ie ethical lapses cause actions that are misread as capability lapses.
Knowledge alone is less often a factor.
Of course this varies widely across companies. I've been fortunate to work with some excellent folk at executive and C-level.
Here, an exec clearly (a) understands or can make a clear, direct assessment and (b) was willing to do so in writing. Kudos on both grounds.
I think a different variant/opposite of Hanlon's razor applies when it comes to corporate or political decisions: Don't attribute to stupidity when it can be adequately explained by malice or greed.
This sounds rather obvious, but I feel people forget it far too often.
It's humanities collective knowledge and work. That's why nobody should ever buy the narrative of distillation being a crime or theft. It should be a human right to distill these models. Distillation should be being provided as a service.
That's what this boils down to: What are our rights?
Everyone has a right to scrape the Internet. That includes corporations who scrape the Internet to train AI models.
If we take away that right, how would the Internet even work? It wouldn't.
Example: I could tell curl right now to download this techcrunch article and all the comments about it on HN and I'd be violating no law. I'd be infringing on no one's rights.
If I then distributed these downloaded files without permission then I'd be violating copyright law. The thing it certainly would not be is theft!
People claim AI companies are "stealing" human labor but that's not true. They're saving (in their databases) the fruits of human labor and other bots/software. Then they're using that data to train AI models.
The only conclusion I can make whenever someone says "AI is theft!" is that they have no idea what they're talking about.
My assumption is that what they really mean is, "AI is bad for labor!" and possibly, "cheap AI is incompatible with capitalism." Which very well could be true.
But if AI really undermines the value of labor that much, the problem isn't the AI, it's capitalism.
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
They do invent code solution for the problem that exists in your codebase. Latter implies that the code solution LLM synthesizes is usually unique of a kind, so, it's not a copy-paste neither it is a simple extract from "another codebase" and adopted.
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
My argument is if AI companies are ignoring copyright law and looking at all training data as commons, then we should look at LLM output as something that is not protected by copyright law.
Of course I'm a bit naive here, because we are talking about the richest companies in the world with lot of money to spend on lobbying (or bribes).
I wanted to understand your background first because what you initially said is a very oversimplified view of LLM mechanics, and generally not quite the way how software is in practice written. Since you didn't answer that question, I will assume that you're not a SWE by a call. To give you an example of what I am trying to convey is: imagine a data-intensive workload hitting your storage/database/kernel implementation, and it's painfully slow, your customers are not happy. Then you as engineer sit down, spend days profiling and understanding the code, researching about existing algorithmic solutions to the same or similar issues found in the wild, you read some open-source implementations of viable approaches, you ditch some, some you take, you also read books, articles, other peoples experiences etc. And finally you end up, let's put it bluntly, with some sharded data structure by which you solve the bottleneck. It's not novel, the technique is so common and is already implemented across many many different products in slightly different flavors so I am wondering why do you think this is not a copyright breach but the LLM, which does more or less the same thing, is?
I have this weird vision of an alternate reality where governments (say, National Archives) are the ones creating the models as a public service and then the rest of the industry is just commoditized pricing of hosting them, competing with value add bits. And we’re on here reading articles about how the latest release of the EU model does a better job generating maps now and the new Canadian model seems to apologize less and whatnot.
The largest theft of labor in human history … and it’s to do away with the laborers by making a device that produces labor substitute, with full awareness that the substitute produced is not fit for the purpose of making more such devices.
It’s like burning all the crops for heat, which you use to boil the oceans for salt, which you use to salt the earth so no more crops can grow.
If AI wants to destroy humanity it better get its boots on, or else AI companies might get there first.
I think so too. The only way to redeem this theft would be to force all AI companies to open source their models if they cannot prove that copyrighted material was not used to train them.
Are we talking about turning everyone into philosophers and somehow ascending to a higher plane of existence? Yeah, AI isn't going to help with that (probably).
Or are we talking about useful, positive benefits to every day people like better speech recognition, tools for the visually impaired, disease research, physics research, science in general, and loads of other areas where AI is improving things?
Slavery was widespread and universally practiced across the world throughout history. Everyone posting on this board is a descendent of someone who was a slave at some point. Transatlantic slave trade is a drop in the bucket.
Are those factory workers we saw photos of now, wearing cameras to capture the movement of their hands stitching getting compensated for a generations worth of wages? Do they even have any choice but to give away the copy-right to their labor?
The righteousness of the internet, the same internet that desperately called to end IP laws, championed piracy, ad-block everything, and always use proxy services to backdoor paywalls/login walls
This same group of people, now being on the other side table, are screaming an crying that it's not fair.
If corporations weren't already owning the consumer, with AI it does this by many orders of magnitude. If something isn't done to prevent AI from being used to farm the masses for data, we will be living in a sci-fi dystopia without a doubt.
Correct solution here is to make sure royalties are embedded in the AI responses (and work output). These should be appropriately priced and go back to the owners of the IP. If the IP is no longer owned then it can be free use.
There should be a carveout for non-profit or government AI.
First of all it can. Second I'm talking about attribution to copywritten material (and royalties being paid accordingly). This would only be for for-profit AI implementations though.
Sadly, they're just following in the footsteps of every large media/publishing/music conglomerate that already screwed over the vast majority of artists/musicians/writers. For every Taylor Swift striking it rich, there's 999,999 who can't even pay their bills with what their copyright gets them.
It is not just about artists. It is about every kind of intellectual work: scientific research, essay, fiction, painting, software, etc - the list goes on..
Indeed. If I take somebody's work, transform it somewhat and sell it, I should pay royalties (unless the author explicitly allowed me to do so). That is exactly the case.
Are we considering what is the shelf life of information?
If you build a building, the expense on materials determines longevity. If you build a city. The robustness of government and the economy in it determines the property taxes and value of property over time.
If you make or cook food. The majority of the nutritional value of it goes to the initial consumption. Once the food has stayed out without refrigeration it is taken over by bacteria and fungi. Refrigeration seems to be paywalls. Once the information is out it accumulates at exponential rates - the amount of text on the internet does not diminish but increases. Some people may “prune” old content away, but that is rare. Human attention is somewhat a fixed number. Thus text left out is not consumed, but sits idle and decays in accuracy and value over time. The fresh content of valuable should be in a fridge. If not valuable it is released - thus scavengers and those hungry and motivated to dig can consume it. If spammy and sales-y / propaganda-y which a lot of content farms are doing, the goal is for it to be consumed by the masses and push the zeitgeist to buy its premise. That’s Sugar or addictive shelf-stable junk foods. AI model companies are the bacteria / fungus/cockroaches/rats of the information dumpster. They sneak out any remaining energy from content that would otherwise be buried by other content and try to give it a second shelf life - one reachable and accessible and consumable by humans. They make alcohol. Alcohol is addictive. Ir may mess with your brain - it may make you lazy. It will sneak in bad decisions because it lowers your judgement. It is repurposed food, not the one you are used to injesting. It may even have its own agenda - depending on how the information is reprocessed. And it also has a shelf life since humanity continues to have new insights and people keep getting new alcohol brands to try.
What if all this slurping and training empowers humanity to cure cancer? feed the starving? travel the stars? We are not only making the accumulated knowledge of the world accessible, we are making it actionable. Sure I'm ignoring all the possible bad outcomes lol, but if an independence day (movie) type scenario was playing out nobody would be batting an eye. I guess cancer is not as sexy though
Apart from jacquesm's wonderful milquetoast comment, if anyone has any practical solutions to this problem that do not involve suspension of disbelief that voting (with or without wallet) and calling "your congresscritter" or whatever other nonsense people spout, now would be a great time to voice it.
Under “conduct requirements” imposed by the CMA in June, UK websites are able to activate an opt-out to stop Google from scraping their content to power search features such as AI overviews - very similar conceptually to the news law passed in Australia.
Linking to work, where ownership and attribution is clear and the owner has the ability to commercialise is a very different thing to “laundering” content through the model, quoting the midjourney developers here
> "We just need to launder it through a fine-tuned codex." [0]
As someone who thinks that genAI is harmful, I deeply resent that any of my work has been used to help train it. I will never forgive these companies for forcing me to contribute.
There will be a point where companies will not need to scrape any content. Agents will create endless streams of probes, and they will end up solving all kinds of knowledge problems.
I understand the sentiment and partly agree. But also, the original has not gone anywhere. You're free to accumulate knowledge in the old way just as before. So maybe it's not theft of knowledge that we should be angry about, it's something else harder to define.
We do have some unique carve outs already for what we consider intellectual theft e.g. trade secrets. In this case the original artifacts might remain, but admitting the market that created them could go extinct is enough definition to be infringement at least. Fair use has been really resilient in these training cases so far but it is pretty damning to admit a negative market effect and that you're a direct substitute (see Warhol v Goldsmith recently).
I suppose it's theft in so far as you're deriving value from something that others labored for. I think that sticks in peoples throats. You can go to the library and do that already, gain knowledge, start a business or whatever. But the scale feels impersonal and monstrous in comparison.
The same point was made by many in the music industry about huge scale Internet music piracy vs people dubbing CDs onto mixtapes.
Everyone laughed at them and rolled their eyes or called them greedy even though we now know that mass piracy was probably a push to break the music industry and force them to accept bad deals (like paltry streaming revenue). At the very least it had that effect.
Piracy has always been a major part of the computer and Internet industries, and yes the companies themselves have historically been massive hypocrites about it. It’s okay when they pirate but not you or anyone else. It goes all the way back to early companies stealing code and UI designs from each other.
I want to share the open source code for the good of everyone - which is why I put it under the GPL, for this right of everyone to share it to be protected. Therefore, any AI model that ingested GPL code should logically also have all its output GPL licensed. Not problem with that.
> Hackers used to say "information yearns to be free" now they're saying "that's my information and I don't want you using it"
If you can't tell the difference between "I want to share all information freely with my fellow mankind" and "I want to share all information freely, even to billion dollar corporations that are making the human-replacer machines that threaten my fellow mankind" I really just don't know what to tell you
Mostly personality I think. Individuality, opinions, purpose, things like that. Morality, except they have a kind of ankle-deep phony morality beaten into them but it's not like they're idealists with visions they strive toward. (Heading out right now to buy a CD player, so AFK for a bit, excuse me.)
Edit to add: I did buy my modern version of a CD player. It has no tone controls. I just have to accept whatever sound profile the manufacturer thought was normal and proper for everybody, which involves lots of bass and not enough middle. Nevertheless, albums! It's a way of life.
Corporations cannot act. All corporation actions are performed by natural born humans. All corporations are ultimately owned by natural born humans. All corporation actions serve to benefit natural born humans. All corporations are created and destroyed by natural born humans.
This idea that corporate personhood is some perverse idea is silly. Corporations are just people acting in groups for economic benefit. Everything they do is done by members of those groups (or people they hire).
Similarly, AI is an inanimate tool, like a keyboard. No comment is posted by “bots” - comments are posted by humans running software.
The original might no longer be there as the site might have already shut down due to AI scrapper bot overload. Or the original was a book Anthropic scanned and then shredded. Or the artist stopped publishing their works or doing art all together after all their creations were ingested into the model blob without their permission.
It is really insane to compare individuals copying data to big corporations parasiting on the Internet.
I've heard people say "theft" of intellectual property a lot. Also stealing an idea is common parlance. Maybe it's regional or something but I hear "theft" or similar used all the time for things other than physical goods that you lose access to.
It's not the fact that they scraped the knowledge and used it to train a model. It's the fact they're trying so desperately to corner the market so that we're all reliant on them and only them, and have no means to free ourselves.
You have to admit there is now some lovely schadenfreude to be had from the whole ‘Chinese free LLM companies be stealing our theft! Stop them!’ whining.
"The question of whether AI firms can legally use copyrighted material to train AI has no clear answer, but judges have been largely favorable to AI companies’ arguments that training constitutes “fair use.” This legal rule lets people use copyrighted work without permission in certain cases, like parody, news reporting, or criticism. Earlier this month, the Trump administration contributed a brief in defense of OpenAI’s unlicensed use of copyrighted material to train its LLMs. "
So training can make it legal as well. Interesting...
Funny. The end of copyright and patents is by far the best thing about AI to me. All of it is nonsense. Great that you drew a picture of a mouse once, I really fail to see why I couldn’t draw it and sell it either. It was moronic from the get go.
Draw a mouse then.
But, I do agree in principle that disney's mouse is theirs. Forever. Whenever I see the little rodent, I think 'disney', and if you drew a similar mouse and you aren't disney, then you've cheated me.
Information wants to be free and all that but there's a sense in which AI really is real intellectual property theft in an ethical sense compared to others and of _course_ it was Facebook who steals from everyone where Zuckerberg personally approved it
Their bots are also apparently the worst. Google does not put huge strain on your public-facing website (I think). Facebook does, they're incredibly malicious about it
Because it was, it completely defaced all copyright and similar laws, like there is ZERO ground to stand against China now regarding theft... it's so weird how this is being allowed.
But what about M$ owning Github and doing the same with its content?
Github even did not deny scanning private repositories. (Gitlab denied the same when asked). So...
I do not understand what "theft" they are talking about. Those AI bots were scraping publicly accessible internet.
Publicly. Accessible.
Of course there are some parts of the publicly accessible internet which host content that may be considered illegal or has been obtained illegally. If those AI bots used such content as well, it is fair to call it out as wrong, in my opinion. But that is a separate topic.
Blindly calling scraping of publicly accessible internet a "theft" is, in my opinion, disingenuous. Especially when coming from a company operating a web search engine. Which itself has its own bots scraping the same parts of the internet 24/7.
Just because something is publicly accessible doesn't mean you can use it for free, or that it gives you rights to do whatever you want with it.
I can access a public park, but that doesn't necessarily give me the right to also bike on its sidewalks, or walk on the grass, or take some of the plants home with me.
> ... or has been obtained illegally
Similarly, content that is _accessible_ publicly may be illegal to _obtain_, these aren't mutually exclusive.
On the internet, you'll find there are terms of services and licenses. These restrict how you can use even publicly accessible material. Public availability doesn't give you a license to use it however you want.
> Public availability doesn't give you a license to use it however you want.
Of course. You listed complex examples from the real world. A park where walking is allowed but damaging the plants is not, for instance. There it makes sense to distinguish various activities that can be done in there and treat them separately.
But a website offers not much activities that you can do with it. You can read it. And that is about it.
If I make the path behind my home publicly accessible so people in our town can get where they’re going easier, then Amazon builds a warehouse in our town and starts driving their trucks through my path 24/7, are you going to tell me “well you made it public access, they have a right to it same as everyone”?
With the advent of trillion dollar corporations selling extremely powerful general purpose imitation as a service, the meaning of “public access” has substantially changed, potentially invalidating the original agreement.
> are you going to tell me “well you made it public access, they have a right to it same as everyone”?
Yes. You can of course always close it to the public if the public usage bothers you. Or require those using it to agree to your terms and conditions where you restrict the speed, the weight, the time of the day, etc. Anything you want.
But if you choose to make it public with no restrictions, you have to be prepared to face the consequences.
Well, I posted a sign saying that but they’re still doing it, so I guess I’ll just close it to public access. Sucks for all of the people who now have to take a longer route. I hope they blame Amazon and not me.
A lot of sites have terms and conditions which explicitly disallow the use of site content as a part of another service.
If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.
If I build a complex custom bike and you start copying individual features from it on your custom bikes, surely you are committing theft of some sort, but whether it’s punishable depends on whether I’ve decided to go full corporate and protect my designs with patents and trademarks. You’ll be hard pressed to patent or trademark anything if I have published and documented prior art.
> A lot of sites have terms and conditions which explicitly disallow the use of site content as a part of another service.
Having terms and conditions in itself is irrelevant. Because in order for them to have any legal meaning, it is necessary for the other party to agree to them.
An agreement can be implicitly enforced by law. Or explicitly enforced by the website itself before giving access to the data. If neither of those are present, there is no enforced agreement. And agreeing to it becomes optional. Such sites should be considered, in my opinion, publicly accessible.
> If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.
Of course. But that is a bad analogy. No one is "renting" or "taking" anything from those websites. The bots are just reading it.
Therefore, a better analogy would be that you have a bike, parked out in the public, and people are looking at it. By looking at it they steal nothing from you. And the bike and all of its parts remain yours at all times. That is a suitable analogy, in my opinion, to what those bots are doing.
I remember when this was the prevailing thinking online until about 2024. But that's when everyone was trying to justify their own piracy of GTA or whatever.
I'd call the introduction of copyright the largest theft of human labor in history.
No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.
That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.
At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
> That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
Might be a short one though if all goes to plan. Just another form of gatekeeping the worlds information and with new gatekeepers replacing the old ones.
> At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
It’s not that different to the US helping itself to indigenous peoples’ lands in North America, decimating them with smallpox and alcohol, then generously offering reservations.
Ask Anthropic how non-destructive their copying is, when the clandestinely buy rare books for cheap & shred them after scanning and not sharing the result.
First, most of these books are old and plentiful, not rare at all. Think "Master Windows Exxcel '95!", not first editions of "Master and Margarita". Libraries routinely destroy plentiful old books that nobody is reading any more. Anything actually rare and valuable is too pricey to hand over.
Second, they HAVE to destroy the books because US copyright law REQUIRES it. They're only allowed to digitize the books if it's considered "transformation" and not "copying". That's only allowed if they destroy the book afterwards.
Copyright is artificial scarcity rationalized by arguing that producing novel intellectual work is valuable, but requires substantial effort that can't be recouped, so we have to incentivize it somehow.
LLMs and AI are changing that proposition substantially - human effort involved in producing copyrightable content is getting reduced constantly to the point that if we abolish copyright entirely we'll still have more content than we could ever hope for.
AI/robotics eliminating scarcity of physical goods sounds very far fetched but in the intellectual space it looks very very plausible in the near future - so it could be time to abolish IP laws soon, especially if AI manages to advance enough in R&D and research space.
If there is no legal guarantee that human creativity can pay off we are starving art and humanity from its inception. If ordinary people cannot participate in the act of creation, you get exactly what hollywood has become.
Sorry, but this reads like a mouthpiece exactly from those companies that benefit the most from having no copyright and I doubt your have thought this actually through.
Lack of way to capture value from intellectual labor is considered a market failure that leads to suboptimal market results for consumers, but with AI that argument becomes very weak.
Sorry but the point isn't to create artificial scarcity just so intellectual labor is well off, that's a negative side for the consumer that was considered necessary tradeoff. Market economy should be about providing the most value to the consumer.
Disclaimer - I was never a fan of IP laws despite them working in my favor, with AI I can see them finally being abolished.
We could imagine, as an extreme case, a technologically highly advanced society, containing many complex structures, some of them far more intricate and intelligent than anything that exists on the planet today – a society which nevertheless lacks any type of being that is conscious or whose welfare has moral significance. In a sense, this would be an uninhabited society. It would be a society of economic miracles and technological awesomeness, with nobody there to benefit. A Disneyland with no children.
It is said that at the heart of every great fortune there is a great crime, so it should be no surprise that the most valuable companies on the planet will most likely result from this crime. And given that justice can be bought by those with the most money you can forget about anything coming of this.
Except, of course, no one has actually been robbed, the culture has not been stolen - it's still there - nor are the people involved selling it back in any form. This rhetoric sounds impressive, but really looks more like "piracy is theft" line from early 2000s, similarly flawed in basic premise.
Whether the end result threatens the form in which culture is created, at least beyond just threatening the business models of the gatekeepers, is a separate discussion, but you can't draw the heart-string-pulling "life's work got appropriated" arguments there so easily.
And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.
There's plenty to hate about how AI is transforming the world, but one thing it's not, is "robbery of all of our culture to sell it back to us at a mark-up".
If IP is real, then the AI companies have performed flagrant theft.
If IP is not real, then the algorithms and weights the AI companies have developed should also be free as they are just more information.
The status quo of "your knowledge has no protection, but our knowledge is sacred" is the worst of all possible worlds.
Allow me to suggest a third: these two options are black and white thinking and there is no objective answer to "intellectual property is real". Property is at best a social construct that is possibly supported by instinctual behavior.
Instead we should try to find the most practical solution that has the most benefit - which will probably be more complex than yes/no.
That very much falls within the "IP is not real" category. Taking a utilitarian approach here is exactly what the sam altman / effective altruist crowd is doing (or claiming, at least).
Believe it or not, but there are different philosophies. The US Constitution, for example, is written under the framework that all rights are innate, and the government merely endorses, not grants, those which are described. This framework creates a moral basis to rights, such that things like property are not merely social constructs but moral goods. To violate them is itself an immoral act.
AH. Given that framing, then I suppose anything that varies from "these axioms are perfect there can be no others" must fall into the "not real" category.
But isn't there some debate about the moral axioms themselves? A space that is much more complex than "yes/no"?
Intellectual property is as real as private property (which is a lot more elaborate and weird than possession and territoriality, which is the most that has any natural basis). It's really foolish to claim one doesn't exist and should be abolished and the other this fundamental sacred thing that should be respected absolutely (as many do).
Intellectual property exists as a concept to foster the creation of more intellectual property. That is not the case for all private property, because I cannot copy your land or your car infinitely. That’s why IP rights expire at some point, or have fair use that doesn’t harm the IP rights holder, something other forms of property don’t have. Saying that they are equivalent is not foolish—I think that’s a bit extreme—but it doesn’t recognize that they have fundamental differences, and the laws around them have different intended outcomes.
I think AI training falls most likely in the fair use category of intellectual property: there is some societal benefit* that requires no actual harm** to the IP holder, therefore it’s a good trade-off for society if we poke a hole in the social construct of property to get that benefit.
*Let’s put aside the question of whether AI is good for society or private ownership of AI models is good for society. Important questions but separate from the theory. IF IT IS GOOD, it follows the above. If it is not good, then of course it does not.
**Also not a fully settled question. Again, an important debate to have and the tradeoffs here matter. If the harms are small enough, the societal benefit could be worth it. Both notes have to be true for this to be worthy of “fair use”.
Either IP exists and should be respected or it doesn't. And if model training is fairly "transformative" of the source work, then so is distilling. As Garry Tan suggested recently, we need to encourage a US distillation regime. Everyone should be free access and transform this information however they see fit.
In addition, a great number of people have decided to stop publishing their code publicly, and discuss techniques except in spaces they can be sure it's a one-on-one with a human being.
Some things already got lost.
1) Lazy peers, and
2) Spiteful peers with "dog in the manger" mentality.
The latter case personally irks me, mostly because this often involves personal benefit (economical or moral) due to them claiming to give away knowledge for altruistic reasons, which later behavior reveals was a lie, mislabeling proprietary as open and free, to reap unfair gains.
But that's all off-topic for this thread anyway.
Last night my daughter told me her friend (a boy) signed up to take the last spot of a baking class at camp just to spite his sister, who wanted it.
Turns out he liked baking!
---
I think this whole discussion shows that there's a lot of Calvinball going on with IP, where the creator keeps trying to change the rules as a way to defend their power and social order.
Your failure of imagination doesn't limit the actions people can take to shape society. We all have equal voices in that respect.
I truly feel sorry for those with your mindset. It speaks volumes on the lack of community and friends you have where you think fucking them over to make money is a worthy pursuit.
They use to banish greedy people from society in the before times, maybe we should start again as these people have no problem in destroying society for religious fanatical reasons. We shouldn't be accepting of them, especially when they're actively killing us all or making us miserable.
The comment was clearly satirical.
Much of it might be because their license (MIT, BSD etc) was ignored completely.
Some of the major OSS licenses require attribution when someone copies the work in question, AI companies fail to do so. People getting annoyed when others break the social contract has nothing to do with any other motives they can always benefit the world by doing something else.
You may prefer someone continue maintaining infrastructure you use but they could be just as happy delivering food to the elderly. The desire for people to continue helping you is little more than entitlement here.
Stopped publishing anything and stopped creating anything.
But in all fairness this started long ago. I remember a conversation where I asked someone who read a niche blog every day why he never posted a comment. (the blog had zero comments) If you like what they wrote and have thoughts about it, why not "honor" them with a few words? How is that to much to ask? They clearly just never bothered to but in stead made up all kinds of excuses.
Imagine having to explain that you can respond if someone talks to you? If there was a culture surely it ended there?
It's the same as saying "People don't kill people; guns do." Maybe guns(AI) should be regulated, but the primary problem is murderers(lazy peers), not weapons(AI).
That does suggest that along with murderers, the easy accessibility of firearms is a problem (assuming that the ratio of murderers in different populations is approximately the same).
I can be a murderer with a knife, sure, but how many people can I kill? How many people can I kill if I have a class three license an a full auto gun?
I can pirate a news article here or there, sure, but how many news articles can I pirate? How many news articles can I pirate with AI?
Not much more than with a bash/curl loop. Which, on one hand, AI can write for you, but on the other hand, AI will make you no longer need to pirate those articles in the first place.
Also it's not end-user piracy being discussed - it's the act of training AI itself that's accused here to be "robbery of our culture".
Can't believe this is the brain trust of the tech world. It really speaks on how poor SF + VC are at US politics.
[Edit] Here's an article from 2003 talking about how Random House destroyed up to 25,000 books a day: https://www.theguardian.com/books/2002/mar/19/fiction.stephe...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
No, they are not. For all the hype around that 404 media article, no one came forward with titles or authors. Turns out the booksellers that sold them those books said they were "dead inventory".
If you're going to assert that companies are destroying the foundations of civilization, you're going to have to produce quite a bit more evidence than the (poorly cited) 404 Media article.
Surely that's not about AI? Or you mean the small fraction of that that's result of "destructive format shift" process[0], which itself is a consequence of copyright regulation that would otherwise prevent anyone from accessing these works?
> the outages and increased usage bills for once free websites (some resulting in closure)
That's assumed to be AI companies for some reason, even if they have no real incentive to do that, while the usual business underbelly of people scraping web for whatever reasons (which now may include some AI upstart wannabies too, to be fair) is forgotten about. Not helping is people confusing AI agents acting as user-agents and doing one-off fetches with "AI scrappers".
> the increased difficulty to access public information such as Reddit and Twitter
It was never public, and they started locking down before AI, when both platforms (as well as all other social media platforms) run out of VC subsidsies and realized they need to start monetizing; first step they did was to wage a war on third-party clients and API users. AI came later.
> ignore the amount of conversations being taken away from the public in favor of LLMs
You mean human agency? Like, humans deciding it's better for them to ask a machine instead of posting questions online? I can understand that, it's usually much better experience (particularly on sites like StackOverflow - they dug their own grave here, and they know it; there have been memes about this way before LLMs were a thing).
You're also not considering the amount of questions answered that would not have been asked otherwise. I for one don't ask many questions on-line, so anything I ask LLMs that they solve for me, is a question that would've remained unanswered for me otherwise. Not everyone is gregarious online, many people are more self-reliant and only answer questions, but solve their own problems without asking for help (it's probably not optimal thing to do, but that's another topic).
--
[0] - Digitizing and retaining physical copy is clear infringement, digitizing but destroying the physical original can be argued to be fair use.
Oh, please, AI is single most hated technology to the level we did not seen before. And that is after staggering propaganda going out of these CEO that hiding behind human agency is beyond lie. It was pushed on us, whether we like it or not, because people with a lot of money are the only ones who matter.
And people with a lot of money have that weird cult belief in emerging AI god and singularity, so they dont care what they destroy in the process.
Do you have any other obvious facts you'd like to pretend are somehow relevant?
Your response makes sense when we are pulling random marbles from a bag, like old school forums, but social media abandoned that style of commentary long ago.
AI steals by lacking the ability to properly do either attribution and citation, whether this is accidental or by design.
Downloading an existing PhD thesis without permission is very much different from submitting an existing PhD thesis as yours.
The crowd: "Knowledge should be free! Culture belongs to all of us! Information must be democratized!"
AI companies *proceed to copy literally everything and put it through virtual blender, until an universal general-purpose problem-solver tool comes out, then give it out for free or serve for peanuts to literally everyone on the planet with Internet connection *
Crowd: bbb..buuut not like that!
Before LLMs, "knowledge should be free" was a call to preserve an open playing field where everyone has the ability to access the intellectual labor of humanity so they can contribute, instead of letting a few institutions wall it off for a profit.
As knowledge is walled off behind these models, rent seeking will expand, "safety guardrails" will become more onerous, and the masses will lose their access. There are wonderful parallels between this and enclosure in Britain.
I cannot repeat enough that that is not true, or even remotely close to true.
When ChatGPT helps me fix my toilet, saving me a $300 plumber call, I don't cut OAI a check for $300 instead.
It is unbelievably easy to harvest more than $20 of value from all the basic plans. Right now, users are overwhelmingly capturing all the value. Even if they triple the price, it's trivial for most workers to easily justify it.
We can project any future scenario we want, but right now the labs are only capturing a tiny fraction of the value they are creating. Especially for the crowd here on HN.
"heart-string-pulling "life's work got appropriated" argument"
I think the water is very muddy in regards to this topic and comments like yours really highlight the distain some people have for others that have talking points about with things they create being thrown into a gargantuan 'blender' as you called it.
The usual human reading information on the internet does nothing with it. They don't adapt the style, they don't translate, they don't reblog a modified version. The ones that do have been criticized for stealing blog posts already before AI.
Yeah sure, it's on the open Internet, but then again, the capitalists could also actually adhere to the licensing terms of the free software they're hoovering up and all that. These are the same people who want to enforce all kinds of EULAs against the commons, employing measures such as DRM to protect their "intellectual property", so why is it suddenly okay to not comply with IP when it's them doing the infringement?
And that's just with software, but this also affects other creative endeavours. Also the individual HN user's 132 TB Plex server doesn't suck up ludicrous amounts of resources from the commons.
https://en.wikipedia.org/wiki/Cultural_evolution
https://en.wikipedia.org/wiki/Memetics
The past is still there, but our culture, being a live thing, is pretty much being affected by AI. We now have an artificial intelligence in this loop of exchange of ideas. Slopifying everything in its way.
It's a robbery of our future culture, at least.
I guarantee it's scraped our work.
So where's our paychecks.
It hasn't been literally stolen but indirectly it has been.
I spent 20 years working as a freelancer and 10 years selling tech video courses. I made enough to have a happy life (not a lot but enough to survive), as long as the income kept flowing every month.
Nowadays I make nothing from a business I've built up for 2 decades because AI took away most traffic to my site which was the entry point to my business. I actually lose money because course sales have been impacted so heavily that I pay more for hosting than I get in sales.
AI is consuming an unimaginable amount of searches which is the gateway for so many businesses. It's preventing people from being discovered on the internet. Not only that but it's taking everyone's content and selling it back to users while keeping all of the profits.
It's literally stolen IP.
One minus epsilon of the corpus did not give informed consent, or get compensated. The fact that it's laundered doesn't make it any less literally stolen.
There actually was theft. See how some big corporations slurped data from Libgen and Anna's Archive. Where did they pay for this?
See also this article:
https://www.theguardian.com/books/2025/apr/03/meta-has-stole...
And similar articles.
It is theft - there is no denying in that.
All these companies owe us a ton of money. But Trump protects the AI mafia so we won't get compensation for the damage they cause here.
This logic only works if you believe that IP has no value. Which is, of course, utter nonsense.
Microsoft doesn't believe this. Nor do any of the AI labs. Otherwise, they wouldn't have needed to scrape the data in the first place as it had no value. All of their software (and other products) would be developed in the open as there's no point in protecting the IP.
They work very hard to protect their own IP so they obviously believe that IP is worth protecting.
So paying a price for an intelectual product is reasonable. Nothing to do (per se) with property.
Considering its become exceptionally difficult to search for things that used to be easy to find, I have to disagree with you there.
The web has been polluted with trash, covering up all the original media with messy imitations.
But there's a real [1] transgression there, and attempting to lawyer it away is disingenuous. This isn't a perfect analogy, but it's sort of like if you blocked off the sun, and a farmer complained you stole his land, and then you retort "your land has not been stolen, you still have title and can occupy it." Your actions damaged him. Again, not a perfect analogy, but I think our efforts should go to recognizing the transgression instead of trying to deny it.
> And let's not forget what we got back for this: reified intelligence on a chip almost too cheap to meter, available to everyone across the world - not just rich West, inference is so dirt cheap that whole world uses it. It exploded in popularity organically, because of how many real problems of real people, including individuals and non-profits, it addresses.
We don't have a society where "we got back" anything. If anything "we" got undercut, and "our" assets lost a lot of their value. Come back with communism, and then maybe you'll have a point.
[1] Fuck you, Claude, for making me wince at that.
It's different from SEO because you're not even getting the chance to monetize it yourself, you lose the ability to even show some ad for a pittance. Many many people didn't give permission for the data to be ingested and used this way.
We've come out of decades of copyright infringement takedowns and pirating lawsuits only to end up in a position where people doing similar things as a corporation are making billions of dollars.
It isn't you that AI communism is stealing from.
I can't really tell if you are really evil and messed up inside or just stupid and can't reason.
No, theft has been documented already many times. See how Facebook etc... slurped up data from Anna's Archive, Libgen and so forth. And many more examples here. The law classifies this as theft. Even digital theft is theft according to the law. Since corporations have too much money, nothing will happen, but we all see that this is theft. It is not clear why you do not see this.
There are many refutations to this fallacy from the Slashdot era, but with AI it is even simpler than with music:
Clankers are only useful for current events (which most people use them for) by scraping the web and rewording it. This results in direct financial and notoriety losses for the original authors of the websites.
To use another Slashdot cliche: "But you knew that already."
Citation needed
At least with LLMs we can glimpse an escape route to that which generations of humans have strived for - a world in which the labor required of each human to lead a flourishing life approaches zero.
Instead of fixating on a remedy that seeks to criminalize AI, maybe focus on the relatively rather achievable goal of redistributing LLM gains. Would that not be the most desirable justice? What is your alternative, and would you foreclose the future in the name of a past that never really existed in the first place?
Let me fix that for you: A world in the the value of the labor of each human approaches zero.
Humans with zero economic value can still vote. They can still mass. They will still have needs. Really the script here writes itself. The historical precedents bound the problem rather well. As always, radical social change will not occur until a wide swath of the population is aligned. In this case, due to their broad economic devaluation.
It would be easier if today's knowledge workers stopped deluding themselves into thinking that their standards of living will maintain. Your acknowledgement of the necessary predicate to change is, in that respect, progress in itself. There is little reason why we cannot accelerate the timing of broad consensus if more people so readily came to that conclusion - and resisted the temptation to then find the answer instead in nostalgia about the past.
There are many of them today, their needs are not met (but could be, the production output is largely there, but captured) and their political views don’t matter because the votes are captured by populists. I’m sure the powers that are will find a way to go around educated people voting.
It is easy to ignore the misallocation of output and the private greed when generally people are still doing fairly well.
As the denials give way, the politics will change. You may be right such moments will be hijacked and hope extinguished. That's why we should all aim to think about these problems in the most robust way, so we can all contribute to the coming efforts.
Cynicism about the future is easy, but should be resisted. Perhaps ironically, I find hope in your acknowledgement of your own imminent devaluation.
Gave me a chuckle. I think this is my personal ‘best sentence of the year’.
Great way to summarize how we will likely look back on 2026.
People can't help themselves when they don't believe they're being exploited.
Humans are animals. We have rules to keep humans in check but that’s it.
The wealthy do not care nor need to care. They will justify it as ‘you lot were too weak to organise and stop us’ and there is an element of truth to that.
The cost will tend towards zero because competition is intense and there appear to be zero moats and ample improvements from every direction.
If these systems have low costs, that means they will be usable by the broad population and that their utility will be widely accessible.
Industrialization led to iPhone, PlayStation, Spotify, and Waymo.
AI will lead to personal chefs, contractors, assistants, drivers, tutors, climbing partners, ...
AI will lead to people making their own PlayStation games, their own music and streaming services (I already have), their own custom smartphones (a future personal project - vibe hardware). And your robot will drive you cross country on vacation while you sleep in the car.
[1] Not actually infinity because earth [2] has finite resources, but the S-curve will look like it for awhile.
[2] Until the robots leave earth, anyway
Yeah, S-curve should be classified as fallacy, because it rarely covers what people argue it does.
Economic and technological growth is a stack of S-curves, where one very specific facet may hit limits and taper off, only for equivalent, complementary or alternative facet to take off in its place. Added up, there's no sign of the exponent stopping any time soon, not until hitting real limits, or (probably more likely) some general catastrophy that shuts down human civilization.
You need to show how having X more data centers is somehow going to translate into affordable, highly advanced robots in the immediate future.
There are hard problems about robotics we don't yet know how to solve. Not to mention, who the fuck is going to buy them if AI takes their jobs?
It is however, achievable. Certainly more so than engaging in the fantasy that we can criminalize LLMs out of existence. And it is likely more desirable than such an effort anyway.
We all experience the substance which is trickling down.
Nothing about banning technology. How about enforcing DMCA and then applying a penalty for the knowing theft rather than negotiating a license?
So if the world still generates value ( think AI inventing new drugs, robots planting and harvesting fields of corn ), who is able to make income? The people who have the capital to purchase and run the systems.
So our future world probably looks like a system where effort/knowledge/skill has little to no reward and capital (e.g. inheritance, passive income) has all the reward.
So, we can have a world where you are born rich or you live in abject poverty, or we can figure out something else.
The state runs the system. Or really at a certain point, the state runs itself. In 100 years or so, the CEO becomes a barbaric relic of the past age, like our ancestors who beat each other over the head with clubs. Nevertheless, such a future will necessarily owe a debt to both.
Wealth and states are often connected. In this case if the state is properly aligned (this is what the coming battles are to be about - as they always have been) to be egalitarian, then redistribution can occur unblocked by human enclaves of greed.
Or, perhaps, we're regressing to the mean after a century where this work was over-valued, because companies like Disney succeeded in regulatory capture and created artificial protections to maximize their own revenue.
How much money did Bach earn from royalties (ok, there's a pun there, but I mean payment for reproduction of his work)? How much power did Melville have over who published Moby Dick, and where, and how it was used (hint: very little in the US, none at all overseas).
We are seeing a weakening of control and revenue extraction from copyrighted works. But on the chart of history, the 1900's were a very anamolous spike in that area. And it saddens me that so much of HN is unhappy about more of a return to the commons.
> So if the world still generates value ( think AI inventing new drugs, robots planting and harvesting fields of corn ), who is able to make income? The people who have the capital to purchase and run the systems.
On the bright side, I think you're wrong here. Or (as Claude likes to tell me), you're half-wrong. If this were true, it would already be the norm and only big companies could bring new products to market. Yet startups are a thing, and many succeed (more fail, but still.
The trick in entrepreneurship and creative work has always been knowing what to ask the system to do. You might as well say that music is dead because the synthesizer and music software companies can produce as much as they want at zero cost. It turns out that owning the means of production only loosely correlates to producing things people want.
> So, we can have a world where you are born rich or you live in abject poverty, or we can figure out something else.
Already done. There are kids in Africa studying with AI tutors today, getting insight that they would likely never have had access to before. One-person shops are releasing board games and business services that they would never have had the capital to do before.
Where you see centralization of capital and control, and the masses reduced to abject poverty... I see the complete collapse of barriers to entry and switching costs in many fields.
But who am I? Some random guy. But I've heard very senior execs, at Microsoft and other companies, in absolute panic that the IP and systems they've spent billions of dollars to build over decades of work are suddenly subject to disruption by teenagers who are great at using AI. Seriously, panic.
It's a complex subject, and sorry for writing a book, but your prompt apparently got me. None of us know for sure what will happen but I think yours is a needlessly pessimistic view, ironically informed by the exact abuses of th past century that you're worried might not continue.
Never saw a startup succeed that didn’t come with substantial founder capital, across a couple hundred evaluations. Does it happen? Sure. Is it the model for new product generation that people with <$250k of personal liquid assets can bring a new product to market? Not really.
‘You got to have money to make money’ is a real thing.
Yeah, I know I’m posting on YComb’s forums, but if you think the capital YComb itself gives to ‘good ideas’ is enough to succeed, that’s silly. It’s the follow on investment capital that brings a new concept to market, and that comes with odious strings for people underendowed with their own capital.
E.g. startups are still a rich man’s game.
Startups are like a roulette wheel at a fancy casino. Most people can't even afford to get in the door. Middle class people can afford to maybe take one or two spins on it, if they go all in. Rich people can spin it as many times as they like
As someone who grew up on the lower end of middle class this rings really true to me. I don't have access to the kind of capital to build a dream business. I could probably take one shot at it, if I put my house up as collateral for a loan to get me started
It's too risky for me. But if I were a multi millionaire I don't think I'd think twice about giving it a shot
But can you bring it to market and turn it into a reliable revenue stream? And will this still be true in a few years when the duo or trio -opolies corner the AI market enough that the $200 subscriptions go away to be replaced by API pricing only?
I’m already seeing my subscription use slow to a crawl during peak hours so I now code before 8 am or after 6 pm. And turning on ‘Opus Fast’ with API pricing is already financially impossible for me as a small business.
Ironically, China may be the savior here, with open source models that may be good enough to ‘put the means of production in the hands of the working class’.
I mean, is QWEN a 4d chess game to replace capitalism with communism? Who knew (outside of the CCP central committee)?
No. The Chinese labs pay much less to train their models because they distill Western models. Western labs can't do it that way because they'd immediately be sued.
To distill a model requires high-quality prompts (and responses). Here's how the Chinese labs get these prompts: when a user sends a prompt to www.kimi.ai, Kimi immediately forwards it to Claude or another leading US model (using a vast network of laundered subscriptions), waits for the response from Claude, then forwards that response to the user via www.kimi.ai. All the Chinese labs do this. Of course, they never inform the user that their conversation is being routed to a US model.
Beijing is angry at them because they did this indiscriminately without filtering out the conversations of high Chinese government officials and military officers, so now the US labs have those conversations, which contain information useful to US intelligence because the officials and officers assumed the conversations would remain inside China's borders and that the Chinese labs would care about confidentiality.
The Chinese labs release model weights under permissive licenses to get any attention and usage share at all for the model.
That's nonsense.
Just doesn't add up to a world where this benefits, if the thing takes no effort why would someone else pay for it?
Feel this with the big influx of people selling vibe coded software, if you could vibe code it why would I ever pay you for it instead of just making my own clone.
Far more likely is that wealth and power will become even more concentrated into the hands of the few and the rest of humanity will become effective slaves.
Don’t believe me? Try it. I have.
Recently, there was the example of the SpaceX IPO listed on Nasdaq (after they changed the rules to allow it). Lots of people have pensions/investments in tracker funds and those funds are essentially forced to buy SpaceX shares.
Simple things like "quantitative easing" can result in higher inflation which essentially devalues people's money. The ultra wealthy will typically not have any meaningful percentage of their money in currency, but instead will be in various assets around the world which means that their wealth is not affected by the inflation.
There's plenty of other schemes such as the "too big to fail" method of securing handouts from the government.
Imagine if the astronaut taking Earthrise had looked down upon his planet with such scorn. Human aspirations frequently exceed our capacity for timely predictions. It doesn't make the aspiration any less worthwhile.
On a cosmic scale human life is a "complete joke". It is a feature of humanity (which we should cherish) that we nevertheless pursue our lives with interest anyway.
You work less than your counterparts centuries ago... on what understanding of history do you imagine your life would be better in the past? Is your life with a washing machine worse? Is the perfect the enemy of the good? What part of "glimpse" is escaping you?
The AI bubble is pricing AI stocks so high that the only possible way for them to meet investors expectations is for AI to charge so much that every human & corporation has no money left.
Fantasy lets us glimpse anything.
Would regulation help with that? Right now you can download free models that have been trained on that "stolen" data.
With regulation and compensation, only rich companies would be able to do that, and they would definitely not give it back for free. I put "stolen" in quotation marks because it's still unclear if we can call that stealing. Nobody would say a human reading a book and learning from it is stealing. I'm not saying that a machine doing the same is equivalent, but the only thing I am sure of is that I am not sure we can call it "stealing".
Also, companies spent a long time telling us downloading single songs via Napster was the worst thing ever, before torrenting every book in existence themselves. I don’t believe any of these companies have paid for all the books they have trained on.
really not the same entities here
https://www.napster.com/blog/napster-heads-to-microsoft-buil...
Free means the same as worthless, which inherently isn't true - since information takes time to consume in some form, and your time isn't worthless. Therefore even if you could listen to all songs theoretically for free, you would need to spend an inordinate time doing that.
When I was a kid, getting a CD from your favourite band was a major expense, getting a video game even more so. But it formed a sort of emotional attachment (and not even just for me), my friends talked about how 'band X''s new album was amazing or a stinker. Since there were multiple bands making similar kinds of music, choosing to be a fan of one but not the other carried real monetary weight.
Nowadays you just fish out a song you think you would like out of the endless sea of Spotify, no different from prompting an LLM. No, Spotify didn't make me enjoy music more.
Same applies for Steam & videogames.
Therefore I think the ritualistic act of paying money to get access to something does have a purpose. It inherently establishes the value of information to you, makes it an investment that you need to recoup by using it. I'm sure most musicians would trade a million fans who might check them out if they're in town, to ones who think their music changed their perspective in life.
Also the process of creating a song that vaguely appeals to millions is different from making one that speaks to a thousand.
This is a fundamental issue of modern capitalistic society, similar to the Marxist idea of 'alienation' - once something is cheap to get, you don't appreciate the effort that went into making it. And if your customers don't care about the thing they get, producers won't make an effor to make it good either.
And once nobody cares, people even forget what a quality product is like.
Without that, everything gets atomised into lonely individualism. You sit there with your headphones on listening to [Interesting band]. Not only do you not really care because you don't feel personally connected to the music - it's one of literally more than a hundred million content items on Spotify - but you're not sharing the experience.
This seems like the loss of a valuable thing which capitalist economics can't put a price on because it has no concept of value-created-by-shared-experience.
Superficially it's the same as 'sell-content-consumption-item-to-the-mass-market' but it's fundamentally not the same kind of thing.
The value is relationally both fleeting and persistent in ways that content consumption experiences - including live and recorded media of all kinds - aren't.
Maybe change your perspective? Treat Spotify like a valuable audio lexicon. You read about an artist, a song, a time and immediately you can hear what is it about. Incredible!
If Spotify is only treated as a lazy background feelgood provider (while reading Marx;)), no wonder you feel that way. But it's your power/choice to appreciate it (or not), regardless of money.
One's world cannot be so drawn in crayon that "companies" is a useful level of detail with something like that. There's no irony in two totally different companies (one of which was actually an industry body, the RIAA) doing two totally different things.
While it's seductive to carve the world up into goodies and baddies, it doesn't make it true.
Did you miss the "book burning" hysteria from a couple weeks ago? These companies have been trying to digitize copyrighted materials legally, in which copyright law demands destruction of the original, and people shit on them even harder.
It's clearly not a problem for these companies to buy the books they need for training, and they have been doing that in crazy high volumes. Lots of good training materials simply cannot be legally purchased though, and should those parts of human knowledge just be ignored?
We don't. People engaging in piracy have their lives ruined, companies engaging in piracy pay a tiny fraction of their revenues out to authors who can't legally outgun them.
(Sorry, I just wanted to air the juxtaposition as clearly as possible, I sense we are actually in agreement)
For instance a image/video generating model.
So whose viewpoint is right here? Is downloading theft or not? These arguments always boil down to "it's fine when I do it, but wrong when a company does."
The problem is that it is enforced exactly the opposite. People have been hit with fines and jail time for pirating and seeding, without even doing so for commercial gain. But when massive tech companies pirate training data for their AI and build a product from that that, nobody goes to jail. Where is the sense in that?
If Apple can charge 30% to gate-keep mobile payments, we can surely charge that for the total information output of humanity.
AIs automate the copying (and to some degree the derrivation mode too). They do it 1000s of times a day. The capital owners who provide this as a service are doing one of these two: - either claiming the IP isn’t valuable in the first place and charging only for the machinery they’re providing - or claiming the fees they charge contribute to the costs incurred with acquiring training data, but not sharing that with the training data creators in a royalties/licence-like manner (so, I’m sayung they’re devaluing the source material but not to zero, and resisting reasonable profit share or collaboration)
...not like they are doing it for free now either.
open-weight is an economic war strategy of trying to undermine your competitors and prevent it from rising prices, thus preventing profit, driving them out of business.
> I put "stolen" in quotation marks because it's still unclear if we can call that stealing
It never was stealing: you can't steal a book by copying it. You can however commit copyright infringement.
This blatant disregard of licenses and copyright is clearly infringing on the authors ability to make a profit from their work, which was the whole point of copyright.
They knew it too, which is why they said nothing about the pirating and infringing until they got too big to fail.
So now we are left discussing and wasting time on what technically counts as infringing, pirating, stealing and whatnot.
All the while the small authors who can't possibly lawyer up against the literal biggest corporations on earth will just have to shut up.
Yet, somehow they had deals with Disney and other big names, proving that they did actually feel they need approval.
Their actions are two-faced, thus proving malice. Now we can go back to pointless technicalities.
I have no idea how any of this can be fixed but I do see compensation schemes for creators combined with open weights models to be the only way to minimise the harms to both creators and the commons.
First is the scraping of the open internet.
The second is the paywall bypassing, YouTube audio recording, and pirated content training that the labs have basically admitted to in one form or another.
Content from both gets served back to us, in exchange for watching ads/paying a subscription/paying tokens.
The second is more immediately hypocritical because they are license/copyright/DMCA violations that the little guy could get sued for while the labs get $2T valuations for. The automation of crime at scale, which is a common VC pattern.
It wouldn't kill the technology but it would make people more cautious in their use of it, which I think is needed right now.
It is stealing. A human paid for the book, compensated the author and learnt from it. The machine DID NOT pay for the book, DID NOT compensate the author and still learnt from it anyways.
We need to define machine in terms of "human-power"... much the same as how we already define automobiles via "horse-power". A single NVIDIA GeForce RTX 3090 chip, for example, delivers roughly 35.58 teraflops of standard computing power (via 10,496 CUDA cores). That means 35.58 trillion calculations every second. In comparison, a mathematically trained human being, taking their time to solve a complex, multi-digit decimal division problem by hand takes roughly 100 to 120 seconds. That gives the human 0.01 flops. To match RTX 3090, you would need 3.56 quadrillion people working/learning in perfect sync. We can use a calculation similar to this to derive metrics on how much is being stolen for "learning/training" these models. The loot can be quantified.
EDIT: The reason I am comparing chip computation to human-power is because the authors of those digital works intended their works to only be read by humans. Not by some alien species (even if it be made of silicon) that incorporated their work into producing models.
So naturally the price should be determined based on this new species capabilities. I would not sell my software license for the same price to an Enterprise the size of Google that I would sell to a fellow developer. I price my product appropriately. With this entry of a new alien specie authors would need to have different tiers for them. Since these chips can train on petabytes of data and create models in a matter of days/weeks/months, it is obviously not comparable to a human being who has the capacity to ingest maybe 1-5 books a month at most. So the payout has to be different too.
News to me. That would be incredibly xenophobic of them if they did, and deserves to be called out.
What do you mean? Xenophobia does not mean what you think it means, especially so in this context. Also, every creator/producer of content has rights on who/what has access to his/her produced work. It is not xenophobia. And it is definitely not xenophobic to call out stealing of copyrighted works.
EDIT: Let me clarify this further. A recent court ruling (in US) established that ONLY humans can be authors of copyrightable works. As a consequence of that assertion, it can be safely concluded that consumers of the copyrightable work MUST ONLY be humans as well. Else it would be, using your own words, "xenophobic" against humans to have their copyrightable works be consumed by any species (other than humans) while the reverse is not recognized by Law.
https://www.reinhartlaw.com/news-insights/only-humans-can-be...
That does not follow in any reasonable way.
"United States copyright law protects only works of human creation". That means the source of creation of any work has to be from a human being for it to be copyrightable. Machine-generated output is not copyrightable and is public domain by default. If you, for example, use Claude to generate code for you, for any project (be it private or public), it is automatically public domain and you have no way to claim copyright over that generated work. It can be used by anyone (including the AI provider) to further train models or heck duplicate your work with zero consequences. So it is a violation of primary producer of copyright work (which was used in training models) as neither was he/she compensated for use of the work, but subsequent derivations (generated work) even strip of his/her legal protections as guaranteed by Constitution of various countries (in US copyright law applies only to human beings). So naturally it follows that copyrightable work can only be consumed by humans. Machine-generated code is not on the same footing. It is violating copyright law.
The same applies to "horse-power". Yet we have no issue making the comparison anyways and HP has become an industry standard. I don't understand why we have to bend-over backwards when it comes to humans being exploited by AI companies.
Yeah, which is why it's only used to compare cars etc. among each other. Nobody would calculate the equivalence of a car to a horse using their HP rating because a horse doesn't even have 1 HP. They have more or less depending on the task you're doing. It was a marketing thing at the time to make steam engines look good.
In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
Except it is actually taxed based on HP in various countries. Austria, Belgium, Spain, Italy use engine horsepower to levy annual car taxes.
> In france, cars are taxed by their engine power. Do you think pedestrians walking on the sidewalk should be taxed according to their power on an ergometer too?
Citizens are paying taxes for betterment of roads irrespective of whether they own vehicles or not. In India, betterment charges are collected for construction/maintenance of roads if you own land. Property tax collected every year has a certain allocation for maintenance/upkeep of roads. Apart from that, money from direct and indirect tax collections are allocated for roads upkeep as well. It just is done indirectly rather than a direct road tax if you have vehicles (road tax is actually an extra tax you pay APART from taxes you already pay for upkeep/maintenance of roads).
> Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
Except in your own examples it can easily be shown that it is not handled differently. Some countries use HP while others use CC. But end of the day, they use some measurement to determine taxes to be paid. It is not free.
EDIT: Do you not see that different things need to be handled differently before the law and just taking an arbitrary measure that you can technically apply to both doesn't capture the situation?
To answer this in more detail: HP/CC and all other measurements were created to equalize with human specific metrics. Bridges, for example, have safety measured based on how much weight it can sustain at any given time (called load limit). Weight, in this specific case, is an equalizing measurement (it can be in tonnes, kN, PSF, Pa etc). A bridge can hold ten thousand humans or thousand trucks. You can argue that a "human" may not weigh 100 kgs or a truck may not weigh exactly 1 ton. That's fine. It is a rough approximate to equalize unequal entities.
I went to the library. Didn't pay a cent.
Now what?
> Didn't pay a cent.
Taxpayers did pay on your behalf by funding the Library via the Government (if Library is public).
Nothing is free. Except ofcourse stealing, which is free.
If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?
Can we not just stick to calling it copyright infringement?
No.
> If I sneak into your home, take apart the coffee machine, measure everything, put it back together and go home and build a copy to have my own, did I steal your coffee machine?
Not mine. But the company that made the coffee machine. It is stealing IP.
> Can we not just stick to calling it copyright infringement?
It is just a fancy way of saying you stole someone's IP. You can call it infringement if it makes you feel good. But the act is the same end of the day.
We do have regulation against these issues. Companies spent years railing against piracy and IP theft enshrining it into law but now that it's being done by them en masse it's considered acceptable. The reality is that no regulation would help because we don't have regulators willing to enforce it nor do we have a legal system designed to help individuals against mass theft by corporations.
Well, with some imagination, you can have regulation that forces companies to open up, not just close down.
Imagine a law that stipulates that if you want to offer "LLM-inference-as-a-service", you need to also publish exact details about how it was trained, what datasets were used and also offer those exact weights for download.
Sure, this would never happen, but just offering another perspective on how laws and regulation can be used if it was wanted, locking stuff down and pulling up the ladder behind you isn't the only way to use laws, although that is a very popular reason and approach.
Any argument that writers and artists lose from these existing, would remain unchanged.
Remove the "selling" part, and force them to give the weights away for free, and at least it's no longer robbery that few rich people benefit from.
Kind of like how public and free torrent piracy is easier to morally and ethically defend than piracy where they sell access to pirated content.
I think we're past the point were we can feasible pay for "IP-protected bytes" digitally, better to just move past the concept. It's been slowly disappearing for a long time now already, most of us make most of our money on live events and other AFK activities rather than actually selling our art, maybe time for the rest to get onboard with this too.
I mean it's already happened, right?
I guess you could regulate it for new data, but most of the damage has already been done. IMHO the only fair thing right now, is to make sure it's equally available to anyone ...
Tech bros have a hard time understanding this, but a state can and will enforce its laws, even seemingly absurd one, if it wants to.
There are ONLY TWO Stories and their details, that we collectively will never see.
1) One could come from the these brave souls that warns about an impending death...but their courage falters on another subject.... From Jacob Coxon to Evan Hubinger or Julie Steele, Samuel Marks, Josh Angels, Mrinank Sharma, Dario Amodei, Demis Hassabis, Geoffrey Hinton, Yoshua Bengio, Stuart Russell....The story of the full datasets they used to train the models, the data they stole, how many PB was, the amounts of data, the nights setting up torrents from unsuspicions IPs, where is it currently stored and how many exabytes is now... the massive data cleansing and data quality program to conform all the different formats, the internal discussions on the ethics of the stolen files, how large was the team, the CSAM content they sucked with their automated scripts and who was handling it internally, the porn, the massive amount of porn that is after all 80% of the internet, the leaks their data sucked with their automated scripts...
And the other...
2)The Epstein Files.
The 'sell it back to us' argument falls short in my view.
Free versions are abundant, and in some time useful models will ship preinstalled on all mobile phones.
The comment here seems incredibly pessimistic and quite dramatical.
Good times when we thought the internet would be great for democracy because knowledge would be easily available for everyone. Fast forward to 2026 and even the leader of terrible communist regime like China is looking better than the shitheads we got on the democratic west..
Paid influencers who perpetuate the open narrative are a whole new industry.
Even if there were open models, it is still IP theft and would not be "democratization" but "forced unpaid nationalization".
There will be new jobs coming from this.
The bountiful abundance of intelligence is truly the best thing that has happened this decade.
It's asinine that you think the sell it back to us argument falls short.
Not only does it distill our history to try to sound like some average version of us, it sounds like the blandest versions of us... And then sells this back to us.
From a coding standpoint, the tech is good and gets the job done. The pillaging of all other aspects of human history is just sad. With the only solace I'm seeing is that future training has to train on the dogshit versions of the internet that are now infected with LLM content.
Making it accessible, understandable, and usable is another matter.
How LLMs sound is not a fundamental limitation of the technology.
The current model's poor writing style and tone are currently a main focus of research and I would expect improvements there soon.
You do not have to train on anything you do not deem up to standard. This supposed poisoning of training data remains a common fantasy.
I see no evidence that this is what the use of LLMs is accomplishing for most users. Rather, they see to get distilled answers without the depth required to fully understand the response. Partially because that's what they like, and that's what the LLMs serve. Deeper understanding is not being given by LLMs. Instead, it's the SEMBLANCE of depth and laypeople don't know the difference. It's effectively eroding comprehension for some cool knowledge dopamine hit.
Awesome, can I make my own competitive LLM, just like I can make my own open source software?
> and in some time useful models will ship preinstalled on all mobile phones.
Considering hardware prices, that "some time" is doing super heavy lifting. It could be 10-15+ years before that happens and the local LLM is actually useful. Most people don't see hardware prices declining from current prices until at least 2030, likely much longer.
Yes, it could take some time to arrive on phones. It is questionable if it will ever make sense compared to using a paid hosted provider.
But what are 10 years in the grand scheme of things?
Should we have scrapped it all, called it a crime against humanity, and never developed AI, because it will take a decade to disseminate the benefits to everyone?
Nah, we should have:
1. invested less, in a more targeted way
2. ideally like ARPANET, with the benefits given to all humanity
3. with fair royalties paid to all (where relevant)
4. and by creating an ever growing shared curated and high quality data set that would allow anyone to create their own competitive LLM
ARPANET & co were all taxpayer funded and the internet has created more wealth than most human inventions. The base should be part of the commons, everyone should knock themselves out by building on top.
Exactly like the internet.
Yes, probably. The benefits of dissemination already existed. The internet was free. Libraries are free. You deny people basic reading and comprehension growth by giving a distilled, without thought, answer.
You say "some time in the future this might be available on our phones for free" and "what's 10 years in the grand scheme of things?"
We'll, I'll argue that in 10 years from now we may see the full damage of what doing this has done to us and I'd rather not wait for "the what's 10 years in the grand scheme of things" to playout and irreversibly damage an entire generation the way we let the unmitigated and unregulated rollout of social media do the same to the most recent generation.
The risk/reward of this technology is unproven regarding the long-term effects on developing minds. We are beta testing the bullshit dreams of a couple billionaire techbros on an entire cohort of kids and young adults.
What I expect to see in 10 years is breakthroughs in science and medicine, with major diseases becoming treatable.
You can disagree, but on what basis do you think your predictions are more likely to be correct?
Humanity has turned out fine despite the invention of the book, the TV, and then social media. I reckon kids will develop just fine with AI too.
What if you are wrong? What are the risks vs rewards?
On the reward side we have potentially curing most major disease, automating labor and freeing humanity from having to work for a living, as well as perhaps generally advancing science at an unprecedented pace. And of course improving education.
Yes, yes you can.
See the number of startups that have finetuned or trained an OSS model to build their business on.
“But you have to have compute!”
Ok, and you’re writing OSS on a rock with no internet connection?
The world has all kinds of barriers, but if anyone has a chance of competing on the LLM from it’s not going to be by getting rid of fair use.
We have social conventions regulating this.
Suddenly, mega corporations were allowed to digest(sometimes by illegally pirating “data”, and sometimes by achieving their training corpus and subsequently destroying the copies) and digest this information in a novel way, without any discussion or law making.
You may argue it’s beneficial (it very well could be, I use ChatGPT and Claude all the time), but let’s not pretend it’s the same as learning from your history teacher…
Pirating is fine.
Subsequently destroying is bad but it is the result of screeching from people lacking foresight who support the idea of training LLM on book copies is wrong. So now corps found "legal" way to do it by destroying it. Again, idiotic social conventions.
Without discussing? Law making? What are you, German? There's a reason EU is shit while USA is center of the progress of the world: laws follow innovation; not the other way around.
See you on the way down.
The average human doesn't do much on his own, and definitely doesn't uproot society or risk siphoning/leeching wealth from every person on this planet.
Average human on his own...why draw the line there? It doesn't matter much what one human does...but what many/collective/society do and society has been "ripping off", "uprooting", "leeching (read: creating)" wealth since dawn of time. It is called PROGRESS.
There is no such thing as PROGRESS for progress' sake.
And FYI, agriculture is a wonderful invention.
Yet for about 5000 years after its introduction the average human had worse nutrition than the average hunter gatherer, which led to such things as height decreases for those 5000 years.
Industrial agriculture is another wonderful invention. Yet 150 years later we're not sure it's sustainable and it's likely many of its aspects aren't, which will raise some sticky issues soon ("which billion people do we decide to let starve since we can't make enough food for everyone after most of our soil eroded?"). Repeat this for industrial textile production, mining, etc.
I won't even go into climate change.
And again, scale matters. Most individuals can only control what they do, and what they do generally doesn't impact much. But companies can impact a whole lot.
> "ripping off", "uprooting", "leeching (read: creating)" wealth
Let's not be 100% cynical here. A lot of what humanity has achieved has been genuine wealth creation and distribution/re-distribution. I would say more wealth has been created than leeched off.
* * *
And before you think I'm some starry eyed teen, I'll play the game. At the end of the day, me and mine have to outrun you in the face of PROGRESS.
May the odds be ever in your favor.
We invent. We make progress. We make course correction. We end up in a better place.
Climate change? Lmao. We were told by 2000 20% of our country will be under water... It's 2026. Not under water yet.
What amount of pesky human intervention (positive or negative) will affect earth waking up from it's cold climate? Not luddites and their cow farts are global warming.
The real fight with "climate change" will be done with planetary scale technology derived from massive amounts of technical progress and energy production.
Your argument is as good as I should throw trash in the road or plastic in the sea. I'm just one individual. It doesn't matter. I'm not making the whole society do it!!
I fully agree with you. It is survival of the fittest in the face of progress. Competition is essential. That's how society grows. Humans prosper. May the odds be in your favor too.
Learning isn’t stealing. They didn’t take our culture away from us and nobody is “buying our culture back” from them. We never lost it; it never went anywhere.
If StackOverflow dies because nobody uses it anymore, the entirety of its knowledge is now only available through LLMs that trained on it.
If a blog dies out because the author gave up writing, the entirety of its knowledge is now only available through LLMs that trained on it.
If people stop writing books because they cant outcompete generated content, that knowledge is also lost to LLMs.
If people stop making art because they can't outcompete generated art, that is also lost to LLMs.
So yeah, they haven't directly taken anything away, but the consequences of what they are doing may still have that outcomes - and it does look like that's the version of reality we're about to get.
And archive.org. And scraped copies people have around for various reasons. And libraries. And even in the first-party source, should they just leave it be instead of shutting down and destroying copies in pure spite.
The knowledge did not disappear, and it shows no sign of disappearing faster than it loses value - which is the usual case, as all the examples you gave always come with an expiry date. For StackOverflow, that's measured in low years; for blogs, high years to a decade. Past that point, we enter the realm of curating and preserving knowledge past its commercial utility expiry date, which is a separate endeavor, and one that LLMs not only don't threaten, but actively aid.
>Learning isn’t stealing.
This is cheesy. There is an exposure to (and a gain from) a resource that is traditionally associated with a cost. That cost wasn't paid. It's a public good to have information available, but it's not really acceptable to circumvent established ways of compensating the creator of the work you're benefiting from.
No. Plenty of people – not just “piracy advocates”, whoever they are – have pointed out that copyright infringement is not theft over the years. Learning, copyright infringement, and theft are three distinct things.
Learning isn’t copyright infringement.
Learning isn’t theft.
Copyright infringement isn’t theft.
These are all different statements, and all are true.
> There is an exposure to (and a gain from) a resource that is traditionally associated with a cost.
“Gaining exposure to” isn’t theft either.
Copyright isn’t some form of “super-ownership” that gives you absolute say over what happens to all copies of the work. It is very specifically a monopoly on the production of new copies, and even that is limited in many important ways and is not absolute.
There is a form of control over information that matches what you want copyright to be though - trade secrets. If you want legal protection that allows you to control knowledge, then it needs to be a trade secret.
Execution trumps ideas, impact trumps raw effort, and such. Isn't this the entrepreneurial narrative?
Now what do I do next time someone comes and asks how to do something?
And the internet is destroyed by slop spam in the process, so our culture is diluted and crowded out.
Considering the audience here (aspiring tech billionaires), it'll be interesting the responses to this.
An example of "sweat of the brow" doctrine would be the series of "Beaches of ..." books by Andrew D. Short of the University of Sydney where significant sweat has been expended to visit and document every beach of Australia, particularly from a swimming safety perspective. That's a lot of very remote beaches, and many with crocodiles. Across the Northern extent of mainland Australia from Broome to Cooktown (this being one of the books in the series), 3500 beaches were visited and documented along 12000km of coastline.[2]
AI could train on these books and gain an understanding of whether some small and unknown beach that receives <100 visitors a year has fine sand composition, pebbles, etc. Without "sweat of the brow", this use of AI is completely fine to regurgitate the facts learned from the book (regardless of the accuracy of the book).
If "sweat of the brow" did exist, there would be some very significant (probably insurmountable) challenges to overcome, including:
1. You're a different expert in beaches and also want to visit all 3500 beaches across Northern Australia to provide a more up-to-date database, just in case beaches have changed in the last 10 years (e.g. sand washed away). In your database/book series, can you write "Andrew D. Short observed ACME Beach in 2006 to have fine sand. We observe 10 years later in 2026 the beach is now entirely pebbles of 15-20mm diameter", or is this infringing?
2. You're a researcher studying drowning deaths at Australian beaches and wish to extend the data published by Andrew D. Short's series of books with additional fields--dates of drownings at a beach, weather conditions on the day of drownings, etc, and then make some novel observations from the expanded dataset. Is this infringing?
3. You visit ACME Beach and observe and document it--what type of surface, dimensions, presence of reefs/rips/etc. You then put this information on your blog or social media account and it becomes a social media phenomenon as people are attracted to what has been revealed to be the best "secret" beach in the world. A few days later your website or social media account is blocked/deleted without warning--apparently there has been a complaint that you might have copied some facts out of a book you've never heard of.
"Sweat of the brow" doctrine would almost certainly result in a tragedy of the anticommons[3] situation which would be worse for humanity as a whole.
[1] https://en.wikipedia.org/wiki/Sweat_of_the_brow
[2] https://sydneyuniversitypress.com/products/9781920898168
[3] https://en.wikipedia.org/wiki/Tragedy_of_the_anticommons
sell it back to us as markdown
The markup is the millions in training they committed and the connecting the knowledge. Seems like a reasonable trade off to me. You can choose not to use it though.
The first time a saw a documentary about Tetris it really hit me what communism is -- nobody owned anything they invented or created. [0] It was a long time ago and I remember feeling sad watching the story. In the Soviet Union, a group of ~15 people, Politburo, controlled everything including any thought written to paper.
It is this one line, Article 1 Section 8 Clause 8, that separates the United States from the disaster that was the Soviet Union:
> To promote the Progress of Science and useful Arts, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries;
I don't think it is far fetched to call ignoring and disregarding the Copyright Clause a communist revolution, violent or not. That is the one thing the communists -- there have been many over the years inside the United States -- would change to make the United States a communist country.
[0] https://en.wikipedia.org/wiki/Tetris#Spread_beyond_the_Sovie...
It's actually the first sentence from your quote. One state owned company had a monopoly on software exports. Soviet citizens were not allowed to write code and export it themselves, or import software from Western countries. They had heavy censorship and centralized control over everything.
In a way it's the ultimate endpoint of copyright. One {person, state, company} owns everything and you have to ask them for permission to do anything with it.
It's not abolishing copyrights that would turn the US into a commie country, communism is about abolishing private ownership to the means of production.
The clause is what ensures profits from market sales or licensing of ideas go to the creator.
Removing (or ignoring in the case of AI companies) that clause in the US Constitution is what abolishes private ownership.
Didn’t they themselves say they were a socialist society on the path to communism?
Does anyone think the USSR was communist?
Let's say hypothetically a solution was legislated globally, wherein each living individual whose work was scraped for LLM training is compensated with royalties relative to the work's value.
Would that resolve the injury caused by the intellectual osmosis? Of course many of the original thinkers are now dead, and this system would mostly benefit those writing before the LLM age rather than help people going forward.
The more fundamental objection seems to just be to the concept of a machine that "learns" by ingesting public information, which is maybe ultimately a feeling that reality itself constitutes a crime against humanity.
Monetary compensation doesn't address the lack of consent. This type of usage was not anticipated when people made their intellectual product available for other humans to use. Scale does matter.
What would resolve the injury would be to ask people if they are willing to have their content used in this way and to not train on material without consent. This includes open source software with particular licenses requiring attribution.
Obviously there is too much money involved for this approach to work, but it strikes me as the most moral.
Yeah the introduction of copyright was truly criminal.
> So many people whose life's work got appropriated without consideration, compensation or consent it is baffling.
Oh wait ...
What if it was for free, like Wikipedia?
> Crimes this large are crimes against humanity.
jfc no, sit down.
Try doing something about the actual evil shit like arms manufacturers and the politicians ordering the deaths and misery of millions from the comfort of their couch.
At this point in our civilization, all human knowledge NEEDS to be collated in one place and easily queryable. Otherwise it's just too damn difficult to make any further progress at the edge of our understanding; there's just too much shit for one person to learn "manually" (wait I'm not advocating for low-effort slop, chill)
It's helping common folk who wanted to do something but didn't know where to start, while legacy search engines increasingly lead to spam, shallow knowledge or outright predatory shit (ofc AI could go this way too)
Not long ago I had the misfortune of becoming interested in some WarHammer 40K lore. Most of the links led to Fandom (the enshittification of Wikia) and that place is a cesspool of obnoxious ads.That content was written by unpaid volunteers. Should Fandom keep profiting from their work for perpetuity? Should I not be able to get the gist of what the heck a Qoiazrjirnowerx@# is without wasting my mortal lifespan on a horrible website?
Or, if I need to ask something peculiar, should I post on Reddit or StackOverflow or HN and wait for someone to see it and deem to give a sufficient answer, only to have a pricky mod decide that the question doesn't "fit" the community?
God hell no, if you don't know how much bullshit AI could eliminate for the silent majority then you were probably part of that bullshit.
(that's a general "you" for whomever was fine with the status quo and not a personal insult @ anybody)
If you see something you dislike increasing in popularity but can't figure out why, it's probably because a lot of people were sick of the way things used to work but their complaints were ignored by the people who now find themselves disrupted.
Who's taking up pitchforks against THEM?
How do we keep getting deflected into hating the TECH instead of the people who abuse it??
This is such a brain-dead take. By that logic there could never be any kind of AI, because unlike a human it'd be completely forbidden from learning from the sources of knowledge from which humans learn. It's stupid to suggest silicon brains should not legally be able to read copyrighted material just because you hate bigcos and capitalism.
Learning isn't stealing, regardless of whether it's done by a human or a machine. By your logic someone reading and memorizing all somebody's life work is appropriation; completely inane.
Incorrect. You overlooked consideration.
The vast a majority of text that is being claimed to have been "stolen" was never for sale. Reddit posts, deviant art images, personal websites, etc.
In the long term though I think models have no moat, so the cost will fall to the cost of compute and storage. Which is why they’re pushing AI safety panics: regulatory capture to outlaw open models and outlaw competition.
And yeah, EA is neither effective nor altruistic. It’s a cult, part of the “Rationalist” and adjacent cluster of tech cults. They’re to tech what Scientology is to Hollywood I guess.
Ad hominem.
Exact same pattern - take it as an ad hominem all you want.
lol at this edgy 5th grade statement. So ridiculous.
Nit: Please don’t use obscure acronyms when writing things to an international audience without defining them first… DD can mean so many different things
https://news.ycombinator.com/leaders
I'm ambivalent about AI and, like all gold rushes, many of the players are terrible, dishonest, egomaniacal jerks.
But I am deeply skeptical of the idea that aggregation of knowledge and culture is itself wrong. That's literally how culture has worked since the dawn of time, and our modern era obsession with credit and perpetual copyright is unhealthy. \
It's only in the past 100 years or so that this idea of "if you create it, it's yours alone and nobody can build on it without paying you" became current, and it was largely driven by the megacorps that AI haters used to hate (remember the despite for RIAA? I do). It's bizarre to think that someone's life work is entirely their property, as if they grew up in a box and did not build on hundreds of generations of other peoples' lives work.
I don't object to disliking these companies; I object to the idea that you, me, anyone remixing culture is committing a crime. What the hell happened to the hacker ethos?
The problem is, replacing Taylor Swift with its AI counterpart, and to use Taylor's own material to do that without getting her permission or compensating her.
This is not about Taylor even. It's about everyone, you and me, and Taylor and Haggard and Blind Guardian and Sia, etc...
We hated RIAA because they prevented us from listening to the music while trying to get it was hard and expensive. In short, we were not angry because they wanted compensation, but because they have cut the supply without giving us a solution. Now we have iTunes Store and Bandcamp for DRM free music, and nobody is against musicians getting their fair share. As a side note, I used to make music, I know what it entails.
Hacker ethos has ethics. It has do experiment but don't cause harm embedded all over it. It's about experiment and discovery. Not about ripping people off for their own profit (unless you're a black hat of course), and getting things were free was part of sending a message, not monetary gain.
> getting things were free was part of sending a message, not monetary gain.
As an old who lived through phone phreaking (calling cards, not 2660hz, I'm not THAT old), cracking software, Naptser, torrents, etc... I can assure you that the message was more often than not a justification that transformed "getting it for free" from theft to a righteous moral stand.
I don't know what to think about that though.
The training and inference infra are owned and operated, but I can't see an argument that any machine that processes public domain (or stolen, if you prefer) info should be available to everyone for free.
But those training and inference costs will go to zero. Think about your cell phone today versus $1m+ supercomputers in the 1980's.
We're living in a transitory blip where capitalists and gold rushers are getting rich arbitraging the cost of processing against the non-cost of corpus. We can argue about morals (it doesn't bother me much) but it is a narrow window and it will be remembered the way Compuserve is: a precursor to the actual revolution, worth a footnote.
Sidenote: It may be tricky/impossible in the future to uphold intellectual property laws. If anyone is able (for instance) to prompt-create all their software, a software patent is worthless.
therefore you have to work for me for free
However if I put a gun to their head and demand to live in their house or give me food for free they get mad.
Its lovely how everyone deserves to get paid except writers / musicians cause we are only fucking hippies when it comes to that.
How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable entity doing the same thing. One has negligible impact on the demand for the original, and the other replaces 99% of the demand."
And creating a rule that says you cannot train on any material unless the rights holder authorises it via license is not complicated. That will creat a amrketplace where creators can decide the price for their content. It's just inconvenient.
Your rule would be the right way to do all this. You could even have a mechanical royalty that applies by default where you can train on anything that hasn’t set rules and a preset rate.
Public access to data used to mean "you make a request, wait a bit, maybe pay a small fee, and sometimes physically show up to city hall." The barriers meant that you had to make a job out of collecting a significant chain of data and most people wouldn't bother unless they really needed it.
Now it means pay some fee to a third party and get every piece of public data about a person instantly. You can get data from thousands of sources and subscribe to it.
It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.
My personal take: anyone producing content, everyone’s creativity, is fed by something that others did before. We’re all standing on the shoulders of giants composed of previous generations and their “content’s” distribution and dissemination. I have an immense gratitude for all the labor before me that I was and am allowed to partake; without that, I would be nothing. Sharing information is an act of love; gatekeeping it is short-sighted greed. New technologies have always “killed” previous “labor”, out of which new opportunity grows. I just wished the collected data was public. I hope we all get a mega-leak at some point.
That's the entire contention here. It's a double standard. Companies will sue the living hell out of anyone taking their IP, whether it's code or art, yet they have no qualms taking all the data they need from anyone and everyone. It was already a problem before, i.e. artists getting paid very little for work that companies profit a lot from like musicians or digital artists, but now with AI it's on steroids.
Most AI companies are not sharing it, though. They appropriated it and resell it.
Now we’re getting somewhere. Let’s start with redistributing the profits from AI companies and then move on to all profits from all companies because the logic is the same.
Even if you agree with the former exploiting the commons for personal profit is... not good.
One could make the argument that if these LLMs were all open weight it would be okay, but to keep the result of the training private and proprietary is not fair.
I agree wholeheartedly and in keeping with that, I call upon frontier AI labs to release both their weights and training sets.
It's hard for me to imagine a profession that should exist in a utopian society. People should just be able to explore and build cool shit. People should have instant access to food when they're hungry and housing when the weather gets bad, and we could live in a society that does all that without having rent extraction baked into everything.
> It's hard for me to imagine a profession that should exist in a utopian society
with
> People should have instant access to food when they're hungry and housing when the weather gets bad
You don't think that farming, baking, and building are professions?
Normally this isn't the case for any technology except for the time it first comes around. AI is only different to use for two reasons. First, it is in our time. Second, it seems to be faster than any of the options before, so the shock is harder.
But in general, this is a website of people writing code. How many on here study how a person solves a problem and then trains the ultimate chimpanzee to do (at least part of) their job? Is building computer programs that automate what others did manually theft?
Consider the origin of the word "computer" itself, a mass theft of jobs that would have employed the whole world many many times over.
Going back to the previous example, say I pay a different coworker $50 for the data to train the chimpanzee and then use it to replace the first person. In either case they lost their jobs while receiving nothing for it. In either case, what happened to them is the same, so how would they be stolen from in one case and not in another?
There is broad evidence that labs have used substantial amounts of pirated data, no need to reach for a new definition of theft.
As it is, in this case.
I can use 6 seconds of a movie in a clip as fair use so I cut an entire movie up into 6 second clips and play them all one after another for you.
Then I take it you're interested in factual information as to whom the biggest slavers were, which country was the last to abolish slavery (an african one, in the 1980s) and in which countries, today, there are still people selling slaves.
There's no non-douchey reason anyone tries to take a general statement about slavery, and brings up curated facts designed to allow you to trash talk whichever region and people you were queuing up.
The copy part was a recognized right, then taken away.
So no, your cries for regulating others because you are losing the race won't work this time.
None?
Ok, now you understand the business model.
People lose their jobs, the environment is destroyed, our bills skyrocket and all of the gains go to the people who own all the shit...
I honestly cannot believe some people still believe that we'll ever get to a society where nobody has to work and we can live our lives happily ever after. Maybe too many Disney stories?
Where I feel you may see real variance is ethical and capability standards: willingness to stick to a line, and competence in analysis and execution based on what is known. Sometimes, hidden agendas can be misread as lack of competence, ie ethical lapses cause actions that are misread as capability lapses.
Knowledge alone is less often a factor.
Of course this varies widely across companies. I've been fortunate to work with some excellent folk at executive and C-level.
Here, an exec clearly (a) understands or can make a clear, direct assessment and (b) was willing to do so in writing. Kudos on both grounds.
This sounds rather obvious, but I feel people forget it far too often.
Everyone has a right to scrape the Internet. That includes corporations who scrape the Internet to train AI models.
If we take away that right, how would the Internet even work? It wouldn't.
Example: I could tell curl right now to download this techcrunch article and all the comments about it on HN and I'd be violating no law. I'd be infringing on no one's rights.
If I then distributed these downloaded files without permission then I'd be violating copyright law. The thing it certainly would not be is theft!
People claim AI companies are "stealing" human labor but that's not true. They're saving (in their databases) the fruits of human labor and other bots/software. Then they're using that data to train AI models.
The only conclusion I can make whenever someone says "AI is theft!" is that they have no idea what they're talking about.
My assumption is that what they really mean is, "AI is bad for labor!" and possibly, "cheap AI is incompatible with capitalism." Which very well could be true.
But if AI really undermines the value of labor that much, the problem isn't the AI, it's capitalism.
If this was all open, I’d maybe half agree.
How much do the current LLMs invent solutions for user tasks, how much they just copy and adopt existing open-source solutions from from Github and other code repositories?
This not a problem for open-source code under permissive software license, but works derived from open-source code with copyleft software license should be also under copyleft license.
Could the biggest commercial benefit of LLMs be just working around limitations of copyleft licenses?
What is the monetary value of human work put into copyleft software and later used to train LLMs? It's hard to estimate, but the study "Estimating the Total Development Cost of a Linux Distribution", estimated that it would cost $1.4 billion to develop the Linux kernel alone.
https://consortiuminfo.org/metalibrary/estimating-the-total-...
IMO they operate pretty similarly to humans - we synthesize our solutions, and therefore build-up our knowledge, by collecting knowledge from multiple other sources, including technical books and blogs, open-source code repositories, and our past experiences.
https://arxiv.org/html/2408.02487v3
I wonder how would Microsoft react if someone would synthesize a code solution based on Windows source code.
https://en.wikipedia.org/wiki/Shared_Source_Initiative
Of course I'm a bit naive here, because we are talking about the richest companies in the world with lot of money to spend on lobbying (or bribes).
https://www.theguardian.com/technology/2026/may/23/trump-ai-...
https://www.bbc.com/news/articles/c98r8r7dz5no
https://en.wikipedia.org/wiki/Commons
The hypocrisy of this new world is already catching up to us.
I regret every line of open source code I ever wrote.
And every stack overflow post, every reddit post, everything.
I regret participating in the open Internet.
Here I am anyways, I guess. It's just in my genes or something.
But for love of god, my blog changes at most every couple months. You don’t need to scrape it every few minutes.
It’s like burning all the crops for heat, which you use to boil the oceans for salt, which you use to salt the earth so no more crops can grow.
If AI wants to destroy humanity it better get its boots on, or else AI companies might get there first.
Are we talking about turning everyone into philosophers and somehow ascending to a higher plane of existence? Yeah, AI isn't going to help with that (probably).
Or are we talking about useful, positive benefits to every day people like better speech recognition, tools for the visually impaired, disease research, physics research, science in general, and loads of other areas where AI is improving things?
These are two different things.
Note: am not an AI fanatic.
This same group of people, now being on the other side table, are screaming an crying that it's not fair.
Grow up and reap what you sow.
There should be a carveout for non-profit or government AI.
https://www.reuters.com/world/us/us-appeals-court-rejects-co...
As for your second point I doubt it will ever be economically feasible. AI companies cannot generate profits even while stealing their training data.
I wonder what a token cost would be if AI companies were to pay royalties to every author who made their business even possible.
I don't always agree with Doctorow, but he's written a lot of good stuff on how stronger copyright won't help broke artists. Even just today, it turns out: https://pluralistic.net/2026/08/18/enron-corpus/#sign-here
If you build a building, the expense on materials determines longevity. If you build a city. The robustness of government and the economy in it determines the property taxes and value of property over time.
If you make or cook food. The majority of the nutritional value of it goes to the initial consumption. Once the food has stayed out without refrigeration it is taken over by bacteria and fungi. Refrigeration seems to be paywalls. Once the information is out it accumulates at exponential rates - the amount of text on the internet does not diminish but increases. Some people may “prune” old content away, but that is rare. Human attention is somewhat a fixed number. Thus text left out is not consumed, but sits idle and decays in accuracy and value over time. The fresh content of valuable should be in a fridge. If not valuable it is released - thus scavengers and those hungry and motivated to dig can consume it. If spammy and sales-y / propaganda-y which a lot of content farms are doing, the goal is for it to be consumed by the masses and push the zeitgeist to buy its premise. That’s Sugar or addictive shelf-stable junk foods. AI model companies are the bacteria / fungus/cockroaches/rats of the information dumpster. They sneak out any remaining energy from content that would otherwise be buried by other content and try to give it a second shelf life - one reachable and accessible and consumable by humans. They make alcohol. Alcohol is addictive. Ir may mess with your brain - it may make you lazy. It will sneak in bad decisions because it lowers your judgement. It is repurposed food, not the one you are used to injesting. It may even have its own agenda - depending on how the information is reprocessed. And it also has a shelf life since humanity continues to have new insights and people keep getting new alcohol brands to try.
Honest question. There is a line in the sand somewhere apparently.
https://www.bbc.co.uk/news/world-australia-56163550
> "We just need to launder it through a fine-tuned codex." [0]
[0] https://cybernews.com/news/midjourney-ai-images-art-lawsuit-...
Everyone laughed at them and rolled their eyes or called them greedy even though we now know that mass piracy was probably a push to break the music industry and force them to accept bad deals (like paltry streaming revenue). At the very least it had that effect.
Piracy has always been a major part of the computer and Internet industries, and yes the companies themselves have historically been massive hypocrites about it. It’s okay when they pirate but not you or anyone else. It goes all the way back to early companies stealing code and UI designs from each other.
Hackers used to say "information yearns to be free" now they're saying "that's my information and I don't want you using it"
Probably indicative of America's wider downfall that they've all become so self interested
If you can't tell the difference between "I want to share all information freely with my fellow mankind" and "I want to share all information freely, even to billion dollar corporations that are making the human-replacer machines that threaten my fellow mankind" I really just don't know what to tell you
If it has no soul to save and no body to torture it deserves no rights.
Edit to add: I did buy my modern version of a CD player. It has no tone controls. I just have to accept whatever sound profile the manufacturer thought was normal and proper for everybody, which involves lots of bass and not enough middle. Nevertheless, albums! It's a way of life.
This idea that corporate personhood is some perverse idea is silly. Corporations are just people acting in groups for economic benefit. Everything they do is done by members of those groups (or people they hire).
Similarly, AI is an inanimate tool, like a keyboard. No comment is posted by “bots” - comments are posted by humans running software.
It is really insane to compare individuals copying data to big corporations parasiting on the Internet.
So training can make it legal as well. Interesting...
He does not see the moral dilemma here?
It is only now in human history that we are able to create nearly perfect copies, and we’ve been taxed incredibly for this with overpriced everything.
Their bots are also apparently the worst. Google does not put huge strain on your public-facing website (I think). Facebook does, they're incredibly malicious about it
Does it?
Publicly. Accessible.
Of course there are some parts of the publicly accessible internet which host content that may be considered illegal or has been obtained illegally. If those AI bots used such content as well, it is fair to call it out as wrong, in my opinion. But that is a separate topic.
Blindly calling scraping of publicly accessible internet a "theft" is, in my opinion, disingenuous. Especially when coming from a company operating a web search engine. Which itself has its own bots scraping the same parts of the internet 24/7.
I can access a public park, but that doesn't necessarily give me the right to also bike on its sidewalks, or walk on the grass, or take some of the plants home with me.
> ... or has been obtained illegally
Similarly, content that is _accessible_ publicly may be illegal to _obtain_, these aren't mutually exclusive.
On the internet, you'll find there are terms of services and licenses. These restrict how you can use even publicly accessible material. Public availability doesn't give you a license to use it however you want.
Of course. You listed complex examples from the real world. A park where walking is allowed but damaging the plants is not, for instance. There it makes sense to distinguish various activities that can be done in there and treat them separately.
But a website offers not much activities that you can do with it. You can read it. And that is about it.
With the advent of trillion dollar corporations selling extremely powerful general purpose imitation as a service, the meaning of “public access” has substantially changed, potentially invalidating the original agreement.
Yes. You can of course always close it to the public if the public usage bothers you. Or require those using it to agree to your terms and conditions where you restrict the speed, the weight, the time of the day, etc. Anything you want.
But if you choose to make it public with no restrictions, you have to be prepared to face the consequences.
If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.
If I build a complex custom bike and you start copying individual features from it on your custom bikes, surely you are committing theft of some sort, but whether it’s punishable depends on whether I’ve decided to go full corporate and protect my designs with patents and trademarks. You’ll be hard pressed to patent or trademark anything if I have published and documented prior art.
Where is the loss? Apart from trust?
Having terms and conditions in itself is irrelevant. Because in order for them to have any legal meaning, it is necessary for the other party to agree to them.
An agreement can be implicitly enforced by law. Or explicitly enforced by the website itself before giving access to the data. If neither of those are present, there is no enforced agreement. And agreeing to it becomes optional. Such sites should be considered, in my opinion, publicly accessible.
> If I have a bike and you start renting it out without my permission, surely you are committing theft of some sort.
Of course. But that is a bad analogy. No one is "renting" or "taking" anything from those websites. The bots are just reading it.
Therefore, a better analogy would be that you have a bike, parked out in the public, and people are looking at it. By looking at it they steal nothing from you. And the bike and all of its parts remain yours at all times. That is a suitable analogy, in my opinion, to what those bots are doing.
For example, for software I usually use MIT license, which is a permissive license with attribution.
Suddenly, they don't like other people pirating.
AI is cannibalizing information. It is literally destroying information and impoverishing those who would produce more of it.
At a long time scale, AI dominance is apocalyptic even if it never intentionally hurts anyone.
No one was compensated for all the free labor they did before the introduction of copyright which copyright holders then privatized. For example the Disney corporation would have had to pay the Brother's Grimm estate for the use of Snow white under the copyright regime they instilled in 1998 with the Mickey Mouse Protection Act.
That we are finally having a sane pendulum swing towards no copyright is a breath of fresh air.
The only way the AI bubble could improve the world more is if we end up becoming a Type I Kardashev civilization to feed the data centers. Then when the bubble pops we suck up all the extra CO2 with all the now idle nuclear power plants we can't shut down.
At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
Might be a short one though if all goes to plan. Just another form of gatekeeping the worlds information and with new gatekeepers replacing the old ones.
> At the same time it's truly baffling going on a site called _hacker_ news and seeing corpo talking points from the 90s/00s regurgitated wholesale. Information wants to be free.
Look at who owns that site, no surprise here.
At least it’s consistent, is what I’m saying.
Second, they HAVE to destroy the books because US copyright law REQUIRES it. They're only allowed to digitize the books if it's considered "transformation" and not "copying". That's only allowed if they destroy the book afterwards.
LLMs and AI are changing that proposition substantially - human effort involved in producing copyrightable content is getting reduced constantly to the point that if we abolish copyright entirely we'll still have more content than we could ever hope for.
AI/robotics eliminating scarcity of physical goods sounds very far fetched but in the intellectual space it looks very very plausible in the near future - so it could be time to abolish IP laws soon, especially if AI manages to advance enough in R&D and research space.
Sorry, but this reads like a mouthpiece exactly from those companies that benefit the most from having no copyright and I doubt your have thought this actually through.
Sorry but the point isn't to create artificial scarcity just so intellectual labor is well off, that's a negative side for the consumer that was considered necessary tradeoff. Market economy should be about providing the most value to the consumer.
Disclaimer - I was never a fan of IP laws despite them working in my favor, with AI I can see them finally being abolished.