Blog institucional

Art. 4 Directive 2019/790: What Access Conditions Actually Determine in TDM

Art. 4 Directive 2019/790: What Access Conditions Actually Determine in TDM

Most discussions about Art. 4 of Directive (EU) 2019/790 stop at the same point: it creates a legal exception for Text and Data Mining on lawfully accessed content, and rights holders can opt out if they choose to. That framing is accurate. It is also, for anyone building real data pipelines, almost entirely beside the point.

The practical question is not whether the exception exists. The practical question is what qualifies for it — and what happens to your analysis when part of your source universe does not.


The Exception Is Not Automatic

Art. 4 permits TDM on any content that an organisation has lawful access to, provided that right has not been expressly reserved. The wording sounds permissive. In practice, it introduces a layered qualification process that most data teams underestimate.

Lawful access is not just about being able to reach a URL. It means the access itself — the terms under which you process that content — is consistent with the conditions attached to the source. A publicly visible page is not automatically a page you can process at scale for analytical purposes. The distinction matters because it affects which portions of the public internet can legitimately feed a TDM pipeline, and which cannot without additional legal groundwork.

For organisations operating under Art. 4 — not the research exemption of Art. 3, which carries no opt-out risk — this distinction shapes every decision about source selection.


What "Opt-Out" Looks Like at Scale

Rights holders can reserve their rights under Art. 4 through machine-readable means. In practice, this typically appears as directives embedded in robots exclusion protocols or explicit licensing overlays on publisher sites.

The problem is that these signals are inconsistent, sometimes contradictory, and frequently ambiguous. A site may carry a general robots.txt restriction that was not written with TDM in mind. Another may have terms of service that prohibit automated processing but carry no machine-readable signal at all. A third may have explicitly licensed content for one category of use while remaining silent on another.

At scale — across thousands of sources, across multiple languages and jurisdictions — this creates a qualification overhead that is not a one-time assessment. It is a continuous operational process. Sources change their terms. Publishers update their technical signals. Licensing conditions attached to aggregators shift. A source that was clean to process last quarter may carry a reservation today.

This is not a theoretical concern. It is the reason why the legal layer of a TDM pipeline cannot be treated as infrastructure that you configure once and ignore.


The Structural Consequence: Source Maps Are Legal Documents

For any organisation doing TDM at volume, the source universe is not just a technical asset — it is a legal position. Which sources are included, under what conditions they were qualified, and how that qualification is maintained over time determines whether the analytical output is defensible.

This reframes how source coverage decisions should be made. The question is not only "does this source give us better signal?" It is "can we process this source under Art. 4 without exposure, and do we have a mechanism to track when that status changes?"

The organisations that get this right treat source qualification as a living process, not a legal sign-off that happens at project launch. They maintain documented records of access conditions, monitor for changes in terms and machine-readable signals, and have defined procedures for removing or quarantining sources that shift outside the exception's scope.


What This Means for AI Training Pipelines

The pressure to build large, diverse corpora for AI training has accelerated interest in Art. 4 as a basis for data acquisition. The logic is sound — it is the broadest TDM exception available to commercial organisations under EU copyright law. But the framing that sometimes follows ("we can process anything publicly accessible") is not accurate.

Art. 4 covers lawfully accessed content where the right has not been reserved. It does not cover content where access itself is conditional on terms that prohibit downstream processing. It does not override contractual restrictions, even when the content is technically reachable. And it does not eliminate the need to assess each source individually, because the exception's scope is determined source by source, not at the level of "the internet."

For AI teams building training datasets from public sources, this means the legal audit of the corpus is not a bureaucratic formality. It is the step that determines whether the dataset can be used, how broadly, and under what conditions the model trained on it can be deployed.


The Operational Layer Nobody Talks About

What the directive establishes as a right, the operational environment often makes difficult to exercise cleanly. The gap between the legal exception and a working, compliant TDM pipeline is filled by infrastructure decisions: how sources are selected, how terms are monitored, how the provenance of each processed signal is documented, and how the pipeline responds when a source's status changes.

At TrawlingWeb, the approach to this gap is architectural. Processing the universe of public sources under Art. 4 and Art. 67 bis LPI requires a system that treats legal qualification not as a pre-flight check but as a continuous layer of the pipeline itself. Source maps are versioned. Access conditions are tracked. Analytical outputs carry provenance metadata that links back to source qualification status.

That is not overhead. That is what makes the analysis defensible — and what allows the pipeline to scale without accumulating legal exposure that compounds over time.


The Right Question to Ask Your Data Provider

If your organisation relies on a third-party data provider for TDM inputs, the question is not only "how many sources do you cover?" It is: how do you qualify those sources under Art. 4? How do you track opt-outs and changes in access conditions? What happens to data already processed if a source subsequently reserves its rights?

A provider that cannot answer these questions specifically is not managing the legal layer of the pipeline. They are externalising a risk that will eventually land on the organisation using the data — not the one supplying it.

Art. 4 is a real, usable exception. But using it correctly requires treating source qualification as an ongoing operational discipline, not a legal baseline you establish once and move on from.

← Volver al blog Hablar con el equipo