LLMs in 2026 are still mostly plagiarism with better punctuation, and the reason is structural rather than moral. The training data is still mostly text written by humans. The loss function is still mostly next token prediction. The result is still mostly text that looks like the training data, rearranged, smoothed, and served with the confidence of a system that does not know it is rearranging. The plagiarism is statistical rather than literal, but statistical plagiarism runs as the kind the original author cannot detect, and the kind the original author is not compensated for.
The argument that the output is not plagiarism because the model was trained on the entire internet and the output is a synthesis rather than a copy is a real argument. The synthesis is a copy of the statistical patterns. The statistical patterns are the thing the original author created. The original author is not credited in the synthesis, the original author is not compensated for the synthesis, and the original author did not consent to the synthesis. The argument is technically correct. The argument is also morally bankrupt, and the law is going to catch up with the morality in the next few years.
What the training data actually is
The training data in 2026 becomes the public web, the books that have been scanned into the libraries, the academic papers that are available without a paywall, the news articles that are available without a subscription, the blog posts that have been crawled, the forum posts that have been archived, the code repositories that are public, and the documentation that has been indexed. The training data amounts to the work of the people who created the work, the people who are not credited in the output, the people who are not compensated for the output, and the people who did not consent to the output.
The specific composition of the training data is a trade secret for the major model vendors. The vendors do not publish the source list, the vendors do not publish the proportion of each source, and the vendors do not publish the mechanism for removing the work of the authors who have asked to be removed. The trade secret becomes the legal shield. The trade secret is also the reason the authors cannot verify whether the work is in the training data, and the work stands as the thing the authors created.
What the lawsuits have produced
The New York Times lawsuit against OpenAI, filed in December 2023, amounts to the test case. The Times is arguing that the training on the Times’ archive without permission is a copyright violation, and the Times is asking for statutory damages that would be large enough to matter to OpenAI’s bottom line. The case is still working its way through the courts, and the outcome is going to set the precedent for every other publisher that is in the same position.
The Authors Guild has organised a class action against several model vendors, and the class action is in the discovery phase. The class action covers the authors of the books that were used for training without permission, and the class action is asking for the kind of damages that would force the model vendors to pay for the training data they have been getting for free.
The music industry has been faster. The settlement between the major labels and the AI music generators in 2025 set a precedent for the model vendors: pay for the training data, or face the lawsuit. The model vendors that did not pay for the training data are facing the lawsuits now, and the lawsuits are expensive.
What the model vendors are doing about it
Some model vendors are negotiating licensing deals with the publishers. The New York Times has signed deals with several model vendors. The Associated Press has signed a deal with OpenAI. The deals are paying the publishers, and the deals are setting a market price for the training data.
Some model vendors are training on synthetic data. The synthetic data is generated by the model, filtered for quality, and used to train the next version. The synthetic data is a way around the licensing question, and the synthetic data is a way around the original author question. The synthetic data is also a way around the fact that the synthetic data is not as good as the real data, and the model performance on the synthetic data is not as good as the model performance on the real data.
Some model vendors are paying the authors through opt in programmes. The authors can submit their work, the model vendor can license the work, and the model vendor can pay the author. The opt in programmes are a step in the right direction, and the opt in programmes are a fraction of the work that has been used for training without opt in.
What the original authors should do
If you are a writer, a journalist, an academic, or a creator of any kind of work that has been used for training without permission, the right move is to join the class action, the right move is to file your own claim, and the right move is to demand that the model vendors pay for the work. The model vendors are making billions of dollars from the work. The original authors are getting nothing. The original authors are not getting nothing because the law says so; the original authors are getting nothing because the law has not caught up with the technology.
If you are a model vendor, the right move is to pay for the training data, the right move is to credit the original authors in the output when the output is a direct quote, and the right move is to negotiate the licensing deals before the lawsuits force the licensing deals. The lawsuits are going to happen. The lawsuits are going to be expensive. The licensing deals are the cheaper option, and the licensing deals are the option that lets the model vendors keep building the models that are going to be the future of the industry.

The bottom line
The patterns the post covers have been showing up in production for long enough that the patterns have names, the failures, the mitigations, the gaps. The work the security team and the engineering team and the operations team are quietly doing today sits as the work that decides whether the practice the post names sits as a tool the team uses or a liability the team is paying for.
Sources & Further Reading
All claims in this article are sourced from primary documentation, vendor advisories, and reputable security researchers.
Spotted an error? Email the editor. Corrections are issued with a visible correction note.
Editorial standards. Every article on humanrequired.org is reviewed by a human editor before publication. AI may assist with drafting or research; final editorial control is human. Read the full standards.


