Unsourced
Analysis

The EU AI Act activates enforcement on August 2. It still can't verify what happened on your server

On 2 August, the EU starts actively policing how general-purpose AI providers handle transparency, including whether they honour a publisher's request to keep their work out of AI training data. That is real progress. It is also not the same thing as proof. The one place this gets checked is not in Brussels. It is in your own server logs.

Published 28 July 2026 · Unsourced Research · All research

What actually changes on August 2

On 2 August, the European Commission's active supervision and enforcement powers over general-purpose AI providers switch on. It can request documentation, run evaluations, and issue fines of up to €15 million or 3% of global turnover, whichever is higher. The same day, the Act's Article 50 transparency duties, AI-content labelling, chatbot disclosure, deepfake marking, also take effect (the enforcement timeline is set out here). None of this is new law appearing overnight. GPAI transparency obligations have technically applied since August 2025. What changes on the 2nd is that the Commission can now actually check and act on them.

The clause publishers should care about

Buried inside those obligations is Article 53, and it is the one that matters if you publish anything online. It requires GPAI providers to publish a public summary of their training content, using a template the Commission adopted in July 2025, and to maintain a policy for honouring the text-and-data-mining opt-out that publishers already have under Article 4 of the EU's Copyright in the Digital Single Market Directive. In plain terms: you can already tell an AI crawler, via robots.txt or a similar signal, that you don't want your work used to train a model, and providers are now required to have and disclose a process for respecting that.

The gap: a disclosed process is not a checked one

Here is where the paperwork stops short. The template requires a provider to describe the procedure it says it follows to identify and honour opt-out requests. It does not include any mechanism for anyone outside the company, not a regulator, not a publisher, to independently verify whether that procedure was actually applied to a specific site (WilmerHale's summary of the template covers this directly). A provider can publish a compliant-reading policy and nobody outside the building can confirm whether your specific opt-out was honoured in practice. Disclosure is not proof.

We have seen this exact shape of problem before, on the inbound side. Every AI crawler announces its own identity with a User-Agent string, a line of text the visitor writes about itself, and for years that self-report was trusted at face value. When we verified 9,543 real crawler visits against operators' published IP ranges and reverse DNS for the AI Crawler Report, 5.4% were impostors wearing an AI brand's name. That is the error rate on identity where independent verification is actually possible. A legal compliance policy with no verification mechanism at all deserves at least as much scrutiny, not less.

Where the actual check happens

It happens on your own server, and nowhere else. The question isn't whether a provider's PDF policy sounds reasonable. It's whether the requests hitting your logs claiming to belong to that provider genuinely match its published IP ranges and pass forward-confirmed reverse DNS, and whether any of them fetched a page you had opted out of anyway. We are not claiming to audit a GPAI provider's internal training pipeline. Nobody outside that company can do that yet, template or no template. What we verify is narrower and more useful: what actually hit your specific site, with your specific opt-out signal, on a specific date. That is the one piece of this whole system a publisher can check without waiting on Brussels, and it is exactly the evidence a disclosure-only policy can't give you. Our full methodology sets out how we grade that evidence.

What to do this week

The EU is building real enforcement teeth this week. That is worth watching. It still doesn't hand you proof of what happened on your own site. That part was always yours to keep.

Common questions

What actually changes under the EU AI Act on 2 August 2026?

The European Commission's active supervision and enforcement powers over general-purpose AI (GPAI) providers switch on, including the ability to request documentation, run evaluations, and issue fines. Separately, the Act's Article 50 transparency duties, AI-content labelling, chatbot disclosure and deepfake marking, also take effect that day. The underlying GPAI transparency obligations themselves have technically applied since August 2025. 2 August 2026 is when enforcement gets teeth.

Does the EU AI Act make AI companies respect a publisher's opt-out from AI training?

It gives that opt-out real legal weight. Article 53 requires GPAI providers to publish a public summary of their training content and maintain a policy for honouring the text-and-data-mining opt-out rights publishers already have under Article 4 of the EU's Copyright in the Digital Single Market Directive. The Commission published the mandatory template for that summary in July 2025.

Can anyone independently verify an AI company actually honoured a specific site's opt-out?

Not through the Act's own paperwork. The template requires a provider to describe the process it says it follows to identify and honour opt-outs. It does not include a mechanism for anyone outside the company to check whether that process was actually applied to your site. A disclosed policy and a followed policy are not the same claim.

So what can a publisher actually do about it?

Check the one thing that happens on infrastructure you control: your own server logs. Whether a crawler claiming to belong to a specific AI operator genuinely matches that operator's published IP ranges, whether it fetched a page you had opted out of via robots.txt anyway, and whether you can produce dated evidence of both. That is the part of this system Brussels can't see and you can.

Related

Free AI crawler verifier

Paste a bot name or IP and check it against published ranges and forward-confirmed reverse DNS.

Read it →

Robots.txt can't stop AI scrapers

Well-behaved crawlers obey it; the rest ignore it. Why the rule needs teeth.

Read it →

The AI Crawler Report

9,543 verified crawler visits. 5.4% were impostors wearing an AI brand's name.

Read it →

See the evidence for your own site

Unsourced captures which AI assistants cite you, proves which crawlers really fetched your pages, and re-checks after you act — evidence, not a score.

© Unsourced — the evidence layer for AI search.