ai
We Built the Library. Someone Else Started Charging Admission.
"I am going to stop pretending this is not weird."
I am going to stop pretending this is not weird.
For roughly twenty-five years, the internet built a library. Not the metaphorical one, the actual one. Tutorials on how to configure BGP. Stack Overflow answers on why your Python import is failing. Long-form blog posts on niche debugging problems that maybe fifty people in the world will ever need but whose author took the time to write down anyway. Millions of GitHub repositories, MIT-licensed, with README files that explain how the code works and what it is for. Wikipedia. Reddit threads that solved a problem no manufacturer would document. Open source projects maintained by volunteers on evenings and weekends. Documentation for every language, framework, and API that anybody ever shipped.
None of it was paid for by the people who read it. All of it was produced by human beings, for free, on the theory that a shared knowledge commons was a good thing, that we all benefit when the person who figures out the hard thing writes down what they learned, and that the next person to hit the same problem should not have to solve it from scratch.
A few companies read all of that, then trained large language models on it, then packaged the resulting capability as a product, and then took that product to market with valuations in the hundreds of billions of dollars. And now, having done this, some of the more excited people in the industry are telling the very people who wrote the library that they are about to become obsolete. That their skills are commoditized. That their jobs will be automated away. That they should consider learning something new.
Read that back to yourself and tell me it does not sound bizarre.
I use these tools every day and that is exactly why I am writing this
I want to get the throat-clearing out of the way early, because otherwise it becomes the conversation.
I use Claude Code every day. I used it a few weeks ago to reverse engineer a 2004-era Wind River binary and stand up a modern replacement for a telecom platform I used to run. I have written about that. I built a small MCP server so I can talk to the light in my office in plain English. I have written about that too. I am not on the outside of this looking in. I know exactly which side of this transaction I benefit from.
That is the point. The people arguing that this is fine are, disproportionately, the people benefiting from it. The people whose work went into the training corpus, without their meaningful consent, without compensation, and without so much as an acknowledgement, are being told that the resulting product is going to eliminate their livelihoods. I am one of the people who is currently on the good side of that trade. That does not make it not weird. It makes it more incumbent on those of us who are benefiting to say clearly that the arrangement is not okay.
What the receipts actually look like
For those not familiar with the shape of the value transfer, a few concrete signals.
Stack Overflow, which was the reference library of professional software development for fifteen years, has seen its traffic collapse. The people who used to go there to ask questions now go to Claude, ChatGPT, or Copilot instead. The answers those tools provide were, in large part, trained on the years of Stack Overflow answers written by unpaid volunteers. The site's economic model was ads served against high-intent readers. That model is quietly dying, and the community that produced the content is watching it happen.
Reddit closed its API in 2023 and went public shortly afterwards, partly on the basis of licensing deals to sell its user-generated content to AI companies for training. The users whose posts and comments made Reddit valuable were not consulted. The moderators, most of whom work for free, were not consulted. The people whose specific writing became training data received nothing.
Cloudflare, whose CDN sits in front of a large fraction of the web, has begun offering by-default blocking of AI scrapers to its customers, on the basis that the scraping represents an uncompensated value extraction that the sites' owners never agreed to. That is a bellwether. The infrastructure layer has noticed.
Every major AI vendor is now the subject of at least one active lawsuit over training data provenance. The New York Times. The Authors Guild. Getty Images. A shifting portfolio of open-source project maintainers. The legal question of whether large-scale scraping and training constitutes fair use is genuinely unresolved and will be argued for years. That is a real legal question. But the economic question is separate from the legal one, and the economic question is not really being litigated at all.
The value flows one direction and it is always the same direction
You may have noticed a pattern in the last few things I have written on this blog. Sovereign AI is the new space race and Canada is bringing a sparkler. The US government took Fable 5 off the market with an export control directive. Eight days later Anthropic announced face scans via a third-party processor. And now, in this piece, the argument that the people who created the underlying training corpus have been quietly excluded from the resulting economic gains.
There is one thread through all of these. It is not "AI is bad." It is that the value produced by AI, and the control over who gets to use it, is flowing in one direction, toward a small number of very large US-domiciled companies, with everyone else in the picture (foreign countries, corporate customers, individual contributors, workers whose jobs are about to be automated) increasingly on the receiving end of the deal rather than on the shaping end of it.
The library scenario is just the sharpest version of that. Millions of people volunteered decades of their working knowledge into a shared commons. A handful of companies compiled that commons into a commercial product. The compiler gets the upside. The commons gets displaced.
So what can we do about this?
I have been thinking about this and I do not have a fully worked-out answer, but I have a starting point that I think is worth putting on the table.
If the productivity gains from AI are as large as the industry keeps promising they are, and if the labour-market disruption is going to be as significant as many of the same voices are also predicting, then there is a policy problem worth naming. The tax base of every advanced economy is built primarily on labour income. Corporations pay some tax. Capital pays some tax. But the bulk of what funds a modern state comes from the wages of people who are employed. If AI meaningfully reduces the wage bill (either by displacing workers outright, or by driving down what the remaining human labour is worth), the tax base contracts at exactly the moment the social safety net needs to expand.
Somebody has to close that gap. My proposal, offered as a starting point rather than a finished plan, is a token tax.
The idea is simple. Every token generated by a frontier commercial AI model above some threshold of scale carries a small levy. The revenue flows into a public fund. That fund pays for the transition. It funds retraining. It supports open-weight research. Over the long run, as the labour-market disruption really lands, it becomes the funding source for a universal basic income or its equivalent. The mechanism is the same as a carbon tax, or the Alaska Permanent Fund, or Norway's sovereign wealth fund. When a resource is being extracted at scale, and the extraction is imposing large public costs, the extractor should be compensating the public for those costs. In the AI case, the "resource" is the accumulated human knowledge that made the training possible, and the "public cost" is the displacement of the workers whose jobs the technology is designed to replace.
I want to be honest that this is complicated. Set the rate too high and you slow adoption. Set it too low and it does not offset the disruption. Design it badly and it becomes a regressive tax on individual users while enterprise customers negotiate their way around it. Who administers it, who receives the revenue, and how the money flows back into the economy are all real design questions that I am not going to solve in a blog post. But those are engineering questions, and engineering questions can be worked. The prior question, which is should the public capture some fraction of the value being generated by models trained on the public's shared knowledge, has a much simpler answer, and the answer is yes.
There will be a long argument about the mechanism. There is no serious argument about the principle.
The bottom line
The bottom line is this. For twenty-five years, human beings wrote things down and put them on the internet for anyone to read. That knowledge commons was the raw material of every large language model shipped by every major AI company in the world. The people who built that commons were not asked, were not paid, and are increasingly being told they should be worried about their own economic future. The companies that packaged the commons into a commercial product are, meanwhile, being valued at levels that would have been unthinkable a decade ago.
We built the library. Someone else started charging admission.
If we are going to do this at all, and it looks like we are, we should at least agree on one thing. The people who built the library get some of the ticket money back.
Topics: