Welcome back after the summer break. This issue of the Babylon newsletter has more personal updates than usual. Normal service resumes next week with a report from IBC in Amsterdam.
The good news on my side: Deutsche Welle and Priberam, our partner on plain X, are funded for a large EU project, working title "Europe Now". The call asked for a multilingual information offering for Europe. In our proposal we argued that there is an important step of preparation to be taken before publishing. Journalistic content first has to be labelled, carry its metadata, and be usable by agentic systems, so it can be cited and verified. We plan to build on using a number of standards that exist rather than invent another: ISCC for identifying content, C2PA for provenance, Really Simple Licensing (RSL) for permission and payment. Submitted end of May, funded last week, starts 2027.
THE NUMBER
2 December 2026
December 2 is when an important EU grace period ends: Article 50 of the EU AI Act, the transparency chapter, applied from 2 August. It requires four things: a system talking to a person must say it is AI; synthetic audio, image, video and text must be marked so a detector can read it; people under emotion recognition must be told; deepfakes and AI text published on matters of public interest must be disclosed (the article, Commission guidelines). Machine-readable marking got a transition for systems already on the market. That ends on 2 December (Cooley).
Why care? If you create synthetic voice, dubbed audio or AI-assisted text for EU audiences, you have twelve weeks to add a provenance layer most localisation stacks do not have. Fines can reach EUR 15m or 3% of global turnover. The duty follows the user, so a non-EU vendor serving EU viewers must comply. AI-assisted text can be published without labelling, but only where a human reviewed it and a named person holds editorial responsibility.
ONE THING THIS WEEK
Amazon and Meta are two-thirds of the crawl
3 September 2026 · Press Gazette / 51Degrees
What's new: AmazonBot and Meta-External-Agent produced 63% of all sessions from 72 AI bots across three billion visits in the year to 29 May (Press Gazette).
Why care? Many publishers are negotiating primarily with OpenAI or Perplexity. But your bandwidth is going to two companies that are not at the table. For now blocking works: scrapers honour robots.txt 95.6% of the time (Known Agents). It is a good idea to open the robots.txt file and do a thorough analysis.
TALK OF THE WEEK
Why I pick up trash on the trail and how that relates to AI
Three years ago I read in a brochure in South Tirol that at the end of every season, the villages organise volunteers to walk the trails and clear them of paper towels, plastic and cigarette butts. Since then, when I hike I carry a small bag and one rubber glove. I stop, I pick things up, I carry them to the next bin. The logic is simple: who should do it some day when I can do it now?
The activity costs four minutes across a six-hour day. I mention it because the habit came from somewhere, from something I read as information. I did not work this out myself. Someone reported it, someone published it, and three years later a stranger changes his behaviour in the mountains. I read something, and it changed my behaviour. That is the whole mechanism.
The arrival of AI is changing many things and there are costs: misinformation, falsifications, hallucinations are a real problem. We see "littering" of our shared communication space, and need to find ways to pick up the trash. We should not wait or hope that others will do that for us. We all can contribute, with small time investments. Active contributions by as many people as possible is the path towards more and better information for us humans.
Yes, undeniably there are problems, caused by AI. On 1 September 2026, Guardian Australia published what a program found when it pulled every reference from every submission to the current parliament's inquiries and checked them against academic databases. At least 39 submissions cite work that does not exist, and more than 100 documents carry ChatGPT tags in their reference links (Guardian). Real academics were credited with research they never did, and committee reports have gone on to cite some of these submissions.
And even worse, because this information can be found on an official government website it might then be scraped and presented as reference, for example by Google AI summaries. Google and ChatGPT will cite the parliamentary submission as the source. Fabrication acquires an institutional address. Committee chairs warned MPs two days later (Guardian, 3 September). The Guardian calls its own count conservative, since AI text with no references at all is invisible to the method.
On the other side of the ledger, a report published this summer tries to put numbers on what journalism returns, the positive effects of a healthy reporting layer. A study of 97 countries associates falling press freedom with a one to two percent drop in real GDP growth. A study of 152 countries associates free media access with fewer human rights abuses (DW Akademie).
So, back to what we can do. The criticism of journalism is not unearned. Much of what gets published is built to harvest attention or simply speculates but without a factual base or real new evidence. But the answer to bad reporting is not less reporting. Instead we need to review, refine and reform how to publish information which people will read through agentic AI interfaces.
A lot is changing with AI and the problems are real. But solutions tend to arrive the same way. When cars were introduced there were no seatbelts. As they got faster, road deaths climbed to numbers that would be unthinkable now. Germany peaked at over 21,000 road deaths in 1970. Last year it was around 2,800, with far more cars on the road, all of them faster. Nobody fixed that with a single decision. It took belts, then belt laws, crumple zones, limits, better roads, testing, decades of small enforced steps.
So: the bag and the glove. Four minutes of the day. I do it because I read it from someone who wrote about the problem. The trail is cleaner now. Whoever wrote the original will never know that. But something changed. Let us be optimistic we can achieve something similar for the coming AI world, in small steps, with concrete acts.
GOOD TO KNOW
Rip and replace throws away the asset · youtube Artur Nowakowski of Laniqo, at GenAI in Localization on 4 September: swapping your stack for a model discards decades of approved translations, terminology and brand style. What you get back is fluent, generic, and harder to error-check, so someone still has to fix it. More corrective work, not less.
SPOT leaves beta · findthatspot.io · code Describe a place in plain language, search OpenStreetMap for matches. The detail worth your attention: after two years fine-tuning a custom model, the team found open-weight models did it better and dropped their own. Free, AGPLv3, built by my DW colleagues.
Three investigative agent workflows · library Winners of Northwestern's agent challenge, published 27 August. The reusable patterns: citations carrying document ID, character range and hash; a cheap screening loop separated from an expensive sweep; a final tool that recomputes every total and shows the formula.
ON THE CALENDAR
IBC · 11 to 14 September · RAI Amsterdam · show.ibc.org The plain X team from DW and Priberam will be there, at booth 14.B56, Future Tech, Hall 14 (details). Watch how many localisation vendors now lead with agents rather than engines.
Computation + Journalism · 2 to 3 October · Northwestern, near Chicago · cj2026.northwestern.edu Includes the agent challenge winners' panel.
Languages & The Media · 4 to 6 November · Senate House, London · languages-media.com Audiovisual translation and media localisation. Theme: "Moving Images That Move Audiences: Localising with Intent."
BEFORE YOU LEAVE
Anthropic is changing how Claude writes so the output is easier to detect as machine-written, possibly at some cost to quality (announcement, Nieman Lab). Another feature will be invisible watermarks in the text, which generated some headscratching about how that could work in practice; Anthropic published a blog post describing the approach. John Gruber of Daring Fireball wrote a longer piece. His verdict is in the headline: watermarking by adulterating the text is a perversion of writing.
ABOUT & DISCLOSURE
I am Mirko Lorenz. I work on language technology projects at Deutsche Welle in Germany.
Three projects you will hear about in this newsletter:
plain X (plainx.com) — media localisation platform, DW Innovation / Priberam
ChatEurope (chateurope.eu) — AI chatbot network for 15 European news partners
Cleanfeed — content provenance and verification framework, DW Innovation with Fraunhofer FOKUS, castLabs and G&L (BMFTR-funded, March 2026–March 2029)
AI use: Full disclosure, I do use Claude (Anthropic) to research and edit this newsletter, with prompts I have refined many times. My goal is to find out where AI is reliable and where the hallucinations come in. Before publishing I check all facts and links.
Error log: I maintain an open Google Doc where I collect the small and big problems that a tool like Claude introduces to editorial work. The idea is to get a better understanding of where AI is good and where it is not. Read it here. Responsibility for stated facts, names, and links is entirely mine.
babylon-newsletter.com · 7,000 languages in the world, AI works for 20.

