Blog institucional

Art. 4 Directive 2019/790: What a TDM Audit Actually Looks Like in Practice

Art. 4 Directive 2019/790: What a TDM Audit Actually Looks Like in Practice

Most teams treat Article 4 of Directive (EU) 2019/790 as a green light. It grants the right to perform Text and Data Mining on lawfully accessed content — without needing explicit authorisation from the rightsholder. That sounds straightforward. In practice, it is anything but.

The problem isn't understanding what Article 4 says. The problem is being able to demonstrate, at any given moment, that your operations actually qualify for it. There's a gap between having a legal right and being able to defend it — and that gap lives inside your data pipeline.

If your team has never walked through what an Article 4 compliance audit would look like, you are carrying a risk you probably haven't priced.


The Right Exists — The Burden of Proof Is Yours

Article 4 does not require you to ask permission. But it does require that the content you process was lawfully accessed. That condition seems obvious. It almost never is, once you start pulling the thread.

Lawful access means the source was publicly available, that no authentication barrier was bypassed, and that no contractual restriction explicitly prohibits machine processing. All three conditions need to hold — and they need to hold for every source in your pipeline, not just the ones you can name off the top of your head.

Most data operations start with a short list of trusted sources and expand over time. Expansion is rarely audited. Two years later, your pipeline may include sources whose terms of service have changed, sources that have moved behind a login, or sources that were added precisely because they were convenient — not because they were checked.

That's not a legal edge case. That's a standard condition for teams that haven't built compliance into their sourcing workflow.


Four Questions That Define TDM Compliance in Practice

When compliance counsel or a prospective enterprise client asks you to document your TDM operations under Article 4, they're typically asking four questions. How you answer them determines whether your operation stands or collapses.

1. Can you enumerate your sources? Not in general terms — specifically. A list of domains, types of sources, and the date each was first included. If you can't produce this, you can't prove lawful access.

2. Can you show access was not conditioned on authentication? Every source in your pipeline should have a documented status: fully public, partially public, or gated. Gated sources require a separate legal basis — Article 4 alone does not cover them.

3. Do you have a retention and deletion protocol? Article 4 TDM is not a permanent licence to hold raw content. The right applies to the processing act — the extraction of patterns, signals, and derived insights. Retaining reproductions of source content beyond what the analysis requires changes the legal nature of what you're doing.

4. Can you show what you retained vs. what you discarded? Derived analysis (frequencies, sentiment signals, entity mentions, trend indicators) looks very different from stored full-text. If what you're holding looks like a content archive rather than an analytical output, you're no longer clearly inside Article 4 territory.

These aren't theoretical concerns. They're the actual structure of the questions that arise in due diligence, regulatory review, and enterprise procurement.


Where Most Pipelines Break Down

The most common failure point isn't intentional non-compliance. It's operational drift — the slow divergence between what the team believes the pipeline does and what it actually does.

A source that was publicly accessible eighteen months ago may now sit behind a registration wall. A data feed that was added to cover a specific market may have been sourced from a provider who didn't have a clean legal basis for offering it. A retention policy written in a compliance document may not have been implemented in the actual storage layer.

Operational drift is invisible until it isn't. The moment it becomes visible tends to be a bad moment: a client audit, a regulatory inquiry, or a dispute with a rightsholder who has been watching.

The fix is not a one-time audit. The fix is building the audit into the pipeline itself — continuous source-status monitoring, documented access classification, and clear separation between raw processing and retained output.


What "Derived Analysis" Actually Means for Your Retention Architecture

The distinction between processing and retention is where Article 4 gets operationally specific.

TDM, as defined by the Directive, is the automated analytical technique used to analyse text and data to generate information such as patterns, trends, and correlations. The right covers the analysis. It does not create a blanket right to archive the content analysed.

This means your retention architecture should be designed around outputs, not inputs. What you keep long-term should be the signal, not the source material. Entity co-occurrence frequencies. Sentiment trajectories over time. Volume trends by topic cluster. These are analytical outputs — they have a different legal character than stored full-text from third-party sources.

Teams that conflate "we processed this under TDM" with "we can store everything we processed" are operating on a misreading of the Directive. The right to process does not inherit indefinitely into the right to retain.


Building a Defensible TDM Operation

Article 4 is a genuine legal asset. For organisations working with public internet data at scale — market intelligence, AI training datasets, competitive monitoring, research — it provides a solid foundation that did not exist before the Directive entered into force.

But it only works as an asset if the operation it covers is built to use it correctly. That means documented source classification, clean separation of processing from retention, regular source-status review, and the ability to produce a pipeline inventory on demand.

At TrawlingWeb, the entire data processing framework is built on this legal architecture — processing signals and derived insights from the public universe of the internet, without retaining reproductions of third-party content. The product delivers analysis. The legal foundation is what makes that analysis defensible.

Article 4 is not a clause to cite. It's a discipline to implement. The teams that treat it that way are the ones that can scale without carrying the compliance debt that eventually surfaces in the worst possible context.

If you haven't run that four-question check against your own pipeline yet, that's where to start.

← Volver al blog Hablar con el equipo