Art. 4 Directive 2019/790: What TDM Compliance Actually Looks Like in Operations
Most organisations that rely on public data have heard of Article 4 of Directive (EU) 2019/790. Fewer have actually mapped what it demands from their infrastructure, their workflows, and their data governance decisions.
That gap is where problems accumulate. Not legal problems, necessarily — but operational ones. When a team builds a Text and Data Mining (TDM) pipeline without internalising what the directive actually permits and where it draws the line, the result is a process that is neither fully defensible nor fully efficient.
This post is not a legal summary. It is an operational reading of what Art. 4 asks of the teams who use it.
The Scope of the Right — and What It Does Not Cover
Art. 4 grants any lawful user of publicly accessible content the right to perform Text and Data Mining for any purpose, including commercial. That is the baseline.
What it does not grant is unlimited retention, redistribution, or transformation into a derivative product that substitutes access to the original source. The right applies to the process of mining — not to the artefacts you produce if those artefacts effectively reconstitute the original content.
In practice, this means the analytical output of a TDM process is protected, but the intermediate copies made during processing must be subject to appropriate controls: retention limits, access restrictions, and documented purpose alignment. Teams that treat the directive as a blanket licence to store everything indefinitely are operating outside its logic, even if they stay technically within the text.
The "Lawful Access" Requirement Is an Active Condition, Not a One-Time Check
The directive grants TDM rights to parties who have lawful access to the content. That condition is not checked once at the start of a project. It is a continuous state.
If access conditions change — a source modifies its terms, a feed becomes paywalled, a public authority restricts automated access — the lawful access condition can lapse. A pipeline that continues running against that source after the condition changes is no longer operating under Art. 4 protection.
This is not a theoretical edge case. Sources change their access policies regularly. Organisations that do not monitor those changes and update their access documentation accordingly are building a compliance liability that compounds quietly over time.
The operational implication: lawful access must be auditable at any point in time, not just at project inception. Your infrastructure needs to log not only what was processed, but under what access conditions it was available when processing occurred.
Rights Reservation Under Art. 4(3): The Opt-Out Mechanism You Have to Track
Art. 4(3) allows rightholders to reserve their content from TDM by machine-readable means. When a reservation is validly expressed — through a robots.txt directive, a metadata flag, or a specific contractual clause — TDM activities against that content fall outside the Art. 4 safe harbour.
Tracking opt-outs is not a legal team responsibility. It is an infrastructure responsibility.
The teams building and maintaining data pipelines need mechanisms to detect, record, and respect these reservations in near real-time. A source that opts out today must be excluded from processing today — not at the next quarterly audit.
This is one of the most underestimated operational requirements in TDM workflows. Organisations that rely on periodic manual reviews of rights reservations are not compliant in any meaningful sense. They are compliant in retrospect, which is not the same thing.
Derived Analysis vs. Content Reproduction: The Distinction That Defines the Model
The practical boundary that Art. 4 draws — and that operational teams frequently blur — is between analysis derived from content and reproduction of content.
TDM produces insights: patterns, signals, frequency distributions, sentiment aggregations, trend indicators. These are the legitimate outputs of a TDM process. They are analytical derivatives, not copies.
What TDM does not authorise is building a system whose output is functionally equivalent to reading the original source — a feed that reproduces enough of the original text that a user could substitute it for access to the source itself.
This distinction has direct architectural consequences. It determines how outputs are structured, how much source text is retained in any given record, how excerpts are handled, and what gets delivered downstream. It is not a post-hoc legal judgement — it is a design decision that must be made before the pipeline is built.
At TrawlingWeb, the processing model is built around this distinction. What flows downstream is derived signal and analytical structure — not a replica of the source universe.
Documentation Is the Operationalisation of the Right
Art. 4 does not require registration, prior authorisation, or a licence. What it implicitly requires, if you want to rely on it in a dispute or an audit, is the ability to demonstrate that your process satisfied its conditions at every relevant point.
That means documentation. Not as a bureaucratic exercise — as an operational asset.
The minimum defensible record for a TDM operation includes:
- The sources processed, with timestamps and access condition status at time of processing.
- The purpose of the mining activity and how outputs align with that purpose.
- The retention and deletion policy for intermediate copies.
- A log of rights reservations detected and acted upon.
- The nature of the outputs produced and how they qualify as derived analysis rather than reproduction.
Teams that cannot produce this record are not in a better position because they operate under Art. 4. They are in a worse position, because they assumed a protection they cannot demonstrate.
The Real Cost of Treating Art. 4 as a Checkbox
The directive created a genuine right. It is broad, it is commercially applicable, and it matters for any organisation that uses public data at scale to produce analytical outputs.
But rights that are not operationalised do not protect you. They give you a legal argument that may or may not survive scrutiny, depending on whether your infrastructure, your documentation, and your architecture actually reflect what the directive requires.
The organisations that benefit most from Art. 4 are not the ones that know it exists. They are the ones that have built systems where compliance is a property of the process, not a claim they make about it.
That is the standard worth building toward.