The Wikipedia issue
Wikipedia turned 25 this year. It is also, almost by accident, the largest translation operation on the open web: 347 active language editions (363 were created over the years; 16 have since been closed), more than two million articles created by translating an existing one, all of it run by volunteers on infrastructure a mid-sized company would consider modest. Wikipedia is helpful to answer some big questions of the moment, too: Who trains on your content, and who pays? Where does AI translation help, and where does it pollute?
The whole story is remarkable. Thousands of companies are currently being told by a few big AI companies that in the future they will not manage without them. Wikipedia's translation operation, run by volunteers on a modest budget with no expensive technology, shows something else: there is always the option of finding your own solution, but you have to invest in know-how to find the path. Enjoy this issue, it is full of small details and insights.
Briefly: #19 is the last issue until mid August. I am going hiking, seeking sun, water and fresh air. I hope you do something similar. We meet again.
THE NUMBER
8%
In 2025 human pageviews on Wikipedia fell roughly 8% year on year, the Wikimedia Foundation reported after tightening its bot detection, while the bandwidth used for downloading its multimedia content rose about 50% (Marshall Miller, Diff, October 2025; Birgit Mueller, Chris Danis, Giuseppe Lavagetto, Diff, April 2025).
Why care? The encyclopedia is being read more than ever, just not by people. The relevance is there, but what is missing is a working model, which is beneficial not just for the AI companies, but for the organizations collecting relevant information. This is the same demand shift Babylon tracked for publishers in issues #12 and #13, hitting the one site that gives its content away on purpose.
THIS WEEK
Wikipedia starts sending AI companies a bill
16 July 2026 · Wikimedia Foundation
What's new: Wikimedia Enterprise, the foundation's paid data API, marked five years with a post on a shifting balance: if current trends persist, bot traffic on Wikimedia's sites has surpassed or will soon surpass human traffic. The foundation's answer is to move industrial-scale scrapers onto a paid, structured feed that helps sustain the public infrastructure (Lane Becker, Wikimedia Foundation).
Why care? The foundation announced Microsoft, Meta, Amazon, Mistral and Perplexity as signed Enterprise customers in January 2026, deals formalised over the preceding year; Google has paid since 2022 (Futurism, CNBC). The most-scraped free resource on the internet has concluded that free to read and free to industrialise are different things. Publishers arguing about crawler payment now have a nonprofit precedent to point at.
Current status: Individual deal financials are undisclosed, though Enterprise revenue appears in the foundation's annual reports. The measured basis for the trend: bots account for about 35% of pageviews on Wikimedia's sites but 65% of the most resource-intensive traffic (Diff, April 2025).
The plan to skip translation entirely is on trial
14 July 2026 · Meta-Wiki RfC
What’s new: Volunteers opened a formal request for comment on pausing or closing Abstract Wikipedia, the foundation’s six-year project to generate articles in any language from language-independent structured data (Meta-Wiki).
Why care? Abstract Wikipedia is the rival path to machine translation: instead of translating an English article into Malayalam, generate both from the same data. Six years in, not one article has shipped to any live Wikipedia. A random sample of 13 generated pages this July found four producing only error messages; the RfC’s author calls the output “broken nonsense at worst”. The foundation planned a first live article at Wikimania Paris, which started 21 July. Find out whether it happened this Saturday.
Current status: The RfC was open and unresolved at time of writing. A 2022 external evaluation already judged the project at substantial risk of failure; the foundation rejected it. Watch Wikimania this week.
Ten thousand translated articles, and the citations that were not there
April 2026 · 404 Media / Wikimedia Diff
What’s new: Editors on English Wikipedia restricted translators paid by the Open Knowledge Association, a small Swiss nonprofit (no relation to the Open Knowledge Foundation). The translators translated articles into English with ChatGPT and Gemini, which introduced errors and citations pointing to book pages that never mention the subject (Emanuel Maiberg, 404 Media).
Why care? OKA paid about $400 a month for full-time translation, mostly in the Global South, and published 10,000 articles in three years with 80 translators. The failure mode is the one this newsletter logs weekly: not invention, inheritance. The model translated fluently and carried the errors along.
Current status: OKA’s founder conceded specific failures in an unusually frank public post and tightened verification (Jonathan Zimmermann, Diff, 11 April 2026). New community rule: four documented errors and a translator can be banned. Wikipedia kept the programme. The lesson is governance, not abstinence.
The AI ban that spared translation
20 March 2026 · English Wikipedia RfC / Medianama
What’s new: English Wikipedia banned LLM-generated and LLM-rewritten article text by a 44 to 2 vote, with two exceptions: copyedits of your own writing, and machine translation of an existing article, both under mandatory human review (policy, Medianama).
Why care? The community that runs the internet’s most-cited source decided translation is the one job machines may start, as long as humans finish it. That is a sharper position on AI in the loop than most newsrooms have managed (see #15).
Zoom out: Wikipedia has been here before. In 2016 a configuration error let raw machine translations flood in, and English Wikipedia invented a one-off speedy-deletion rule, X2, just to clear them (Wikipedia). The 2026 ban is the third decade of the same argument.
TALK OF THE WEEK
The translation machine Wikipedia built for itself
Accountability is the variable, not the AI. The instinct after the OKA story above is to conclude that AI translation and Wikipedia do not mix. The record says something more precise: AI translation without accountability does not mix with anything. Because for three years, Wikipedia has been running its own machine translation, at a scale most companies would announce with a keynote, and it is working.
One service, five open model families. The system is called MinT, for Machine in Translation. The Wikimedia Foundation self-hosts open translation models on its own servers and plugs them into the Content Translation tool editors use to carry articles across language editions (Wikimedia Language Team, Diff, June 2023). Not one model: five families, selected per language pair. Meta's NLLB-200, the University of Helsinki's OpusMT, AI4Bharat's IndicTrans2 for the 22 scheduled languages of India, Google's MADLAD-400, and models from Softcatalà, a Catalan volunteer group whose translator covers ten languages to and from Catalan (full model registry). When it launched in 2023, 44 of its languages had never had machine translation from anyone, commercial or open.
It runs on ordinary processors, by choice. Here is the detail that should embarrass the industry: it runs on CPUs. Ordinary server processors, no Nvidia accelerator cards. Running the models on GPUs would have meant installing Nvidia's proprietary drivers, which Wikimedia's infrastructure team, committed to a fully open-source stack, judged not acceptable. So the models were compressed and optimised until they ran without them (Diff, June 2023).
That choice buys independence. Hold that against the rest of the field. The large AI companies are spending roughly thirteen dollars on infrastructure, most of it Nvidia hardware, for every dollar of revenue (covered in #8), and access to GPUs has become the choke point of the whole industry, down to export controls (#14). Wikipedia looked at that supply chain and opted out of it, serving 200+ languages on hardware most AI companies would call inadequate. The trade-off is real: translations arrive slower, and MinT's open models benchmark below DeepL or GPT-class systems. But nobody can raise Wikipedia's prices, revoke its drivers, or export-control its translation service. For a public-interest institution, that is not frugality. That is the whole point.
The workflow is a lock, not a suggestion box. The pipeline matters more than the models. Content Translation has produced more than two million articles since 2015 (Diff, May 2025), and it is not a suggestion box, it is a lock. Publication is blocked outright if 95% or more of a draft is unmodified machine output. Each wiki tunes its own threshold by ticket: Indonesian and Telugu require 70% human modification, Punjabi loosened its limit after MT quality improved (mediawiki.org). Translators whose work was deleted in the past 30 days face stricter limits still. English Wikipedia, bluntest of all, disables machine translation in the tool entirely; its community calls raw MT "worse than nothing" (en.wikipedia.org).
The suspicion pays off. And the result of all this suspicion? Articles started with the tool are less likely to be deleted than articles written from scratch (Diff, May 2025, foundation-measured). Supervised machine translation, wrapped in enforcement, outperforms unassisted humans on Wikipedia's own survival metric.
The OKA difference is the workflow, not the model. Set this against the OKA case and the difference is not the technology. OKA used frontier commercial models; MinT runs open weights that benchmark lower. The difference is the workflow: MinT's output cannot be published without documented human work, deletion statistics are public, and any community can switch the tool off. OKA's output entered the same encyclopedia through a side door, paid by the article, reviewed after the fact.
One shadow: Wikipedia trains the models that help write Wikipedia. Published translations flow back into open corpora that train the next generation of models (mediawiki.org). That loop is a gift when the humans did their job. When they did not, it can poison the well: an audit of major web-mined training sets found that in WikiMatrix, a corpus mined from Wikipedia, 19 of 20 audited languages had under 50% correct sentences (Kreutzer et al., TACL 2022).
A story from a few years ago is a cautionary tale: The Scots Wikipedia, half-written by an American teenager who did not speak Scots, likely poisoned every language detector’s idea of what Scots is (Wikipedia Signpost, 2020).
So, the 95% lock is not bureaucracy. It is what keeps the loop from eating itself. Big lesson here for other big projects betting on “full automation” to save costs.
GOOD TO KNOW
Wikipedia at 25: what the data tells us (Pew Research Center). The anniversary reference sheet. Take away the growth mechanics: several of the largest editions grew by bot, not by community.
Try MinT yourself (Wikimedia). The public test instance of Wikipedia's in-house translation service. Take away a benchmark: compare its output against DeepL or Google for a language pair you know, then remember it costs the foundation approximately nothing per article.
Translating MediaWiki (translatewiki.net). The interface itself, roughly 3,900 core strings, is localised by some 17,000 volunteers, and a language must clear a minimum translation threshold before developers will ship it. Take away that the buttons are a harder localisation problem than the articles.
ON THE CALENDAR
Wikimania 2026 · 21–25 July · Paris · wikimania.wikimedia.org · The movement's annual conference, underway as this issue goes out.
BEFORE YOU LEAVE
The second-largest Wikipedia by article count is Cebuano, a Philippine language: 99.12% of its article creations came from Lsjbot, an automated program built by one Swedish linguist (Cebuano Wikipedia, Vice). Meanwhile 90% of Wikipedia views from the Philippines go to the English edition (March 2021 figure, per Wikimedia traffic statistics cited in the same article). The case is the field's standard cautionary example: nearly all technology, almost no community, and local readers defaulting to English anyway. It points to a key challenge: how to save local languages in a world where English dominates so many fields. No easy answers here.
If you want to learn a bit more about Cebuano, spoken by over 22 million native speakers, go here: talkbisaya.com/cebuano
ABOUT & DISCLOSURE
I am Mirko Lorenz. I work on language technology projects at Deutsche Welle in Germany.
Three projects you will hear about in this newsletter:
plain X (plainx.com) — media localisation platform, DW Innovation / Priberam
ChatEurope (chateurope.eu) — AI chatbot network for 15 European news partners
Cleanfeed — content provenance and verification framework, DW Innovation with Fraunhofer FOKUS, castLabs and G&L (BMFTR-funded, March 2026–March 2029)
AI use: Full disclosure, I do use Claude (Anthropic) to research and edit this newsletter, with prompts I have refined many times. My goal is to find out where AI is reliable and where the hallucinations come in. Before publishing I check all facts and links.
For this issue I reached out to the Wikimedia press team to avoid publishing a "deep dive" full of false claims, facts or missing links. Big thanks to the team there who went deep into the full issue and provided awesome support.
Error log: I maintain an open Google Doc where I collect the small and big problems that a tool like Claude introduces to editorial work. The idea is to get a better understanding of where AI is good and where it is not. Read it here. Responsibility for stated facts, names, and links is entirely mine.
babylon-newsletter.com · 7,000 languages in the world, AI works for 20.

