This is the fifth and final post in my series on the EDPB’s Guidelines 02/2026 on Anonymisation and their consequences for the EU’s Digital Omnibus. The first post set out what the guidelines get right and where they remain operationally thin; the second defended contextual anonymisation against the legal panic that followed EDPS v SRB; the third examined the unequal consequences for research, SMEs and AI development; and the fourth develops the institutional settlement that the Omnibus should now pursue. This final post addresses the problem beneath all four: Europe repeatedly uses the expression “non-personal data” as though it described a coherent legal object, when in fact it is a residual umbrella covering data with very different origins, risk profiles, access conditions and evidential foundations. The argument developed here is that the Union now needs a cross-acquis taxonomy, accompanied by a practical Data Status Passport, that makes every claim of non-personal status answerable to the same structured questions.
European data law has a category problem that it has spent almost a decade describing as a market opportunity. The Free Flow of Non-Personal Data Regulation, the Data Governance Act, the Data Act and the AI Act all rely on the distinction between personal and non-personal data, while the European strategy for data assumes that vast common data spaces can be populated, exchanged and reused across sectors. Yet the formal definition generally offered for the second category remains strikingly empty: non-personal data means data other than personal data. That formula works tolerably well as a jurisdictional switch, because one can ask whether the GDPR applies, but it does almost nothing to explain what kind of data sits on the other side of the switch, how it arrived there, whether its status is general or recipient-specific, or what evidence should sustain the conclusion over time.2,3,4
That emptiness becomes visible as soon as several examples are placed beside one another: a wind-speed reading from a turbine may never have related to a natural person; a national unemployment rate was derived from individual records but no longer speaks about a particular worker; a research dataset may remain personal for the hospital that retains the key while being anonymous for an independent institute; a synthetic dataset may contain no copied records yet still memorise or reveal an individual; and a model, embedding or benchmark may sit even further downstream while preserving enough source-specific signal to permit membership inference or extraction. These artefacts can all be called non-personal in the right circumstances, but they are not the same legal or technical thing, and the habit of treating the umbrella as the taxonomy has allowed provenance, identifiability, access architecture and governance controls to be collapsed into one reassuring adjective.1,3

Non-personal data is the legal umbrella; native, statistical, anonymised, synthetic and derived data are different routes into it, while mixed or unresolved data belongs in quarantine rather than being relabelled by convenience.
The EDPB’s new anonymisation guidelines make that collapse harder to sustain because they accept, expressly and repeatedly, that anonymity may vary from one entity to another. Their core legal test asks whether information relates to a natural person and, if so, whether that person is identified or identifiable from the applicable perspective using means reasonably likely to be used; if either answer is negative, the data is anonymous from that perspective. The Board therefore recognises that information may be anonymous for some entities while remaining personal for others, and it describes the contextual approach as the full legal standard rather than a discretionary relaxation. That is the missing hinge for a serious taxonomy, although the guidelines do not themselves build one and occasionally use “anonymous” in a way that will trouble those who reserve that word for a more general or objective state. The correct response is not to retreat from contextuality, which the case law now requires, but to distinguish the scope of the status claim with far greater precision.1
How Europe arrived at a negative category
The historical starting point was not a market in non-personal data but a boundary around data protection. Directive 95/46 and later the GDPR defined personal data broadly because identifiability could arise indirectly and because information could relate to a person by content, purpose or effect, while Recital 26 of the GDPR excluded genuinely anonymous information from the Regulation’s reach. The underlying structure was binary, but the route across the boundary was always contextual: the recital asks about means reasonably likely to be used by the controller or another person, taking account of objective factors such as cost, time and available technology. What the legislation did not provide was a positive account of the material that remained outside the scope of personal-data law, since the category was legally important mainly as the absence of the condition that activated the data protection regime.1
For much of the twentieth century, the practical discussion took place inside statistical disclosure control, where the problem was how to publish useful information without revealing the people represented in the source records. Suppression, generalisation, sampling, perturbation and aggregation were developed as methods for controlling disclosure risk, and Latanya Sweeney’s work on k-anonymity gave the field a particularly influential formal model by requiring each released record to be indistinguishable, on specified quasi-identifiers, from at least k-1 others5. That work was important precisely because it moved the debate beyond the deletion of names and showed that combinations of ordinary attributes could identify. It also foreshadowed the limitation that now runs through the EDPB’s No Record Isolation and No Linkage criteria1: uniqueness within a table is relevant, but it is never the whole question because external information, inference and the context of access determine whether a person can actually be distinguished and treated differently.
The re-identification literature then broke the comforting assumption that a dataset became safely anonymous once obvious identifiers had been removed. Narayanan and Shmatikov’s 2008 work on the Netflix Prize data showed how high-dimensional preference data could be linked to sparse public information5, while Paul Ohm’s “Broken Promises of Privacy” translated the resulting technical anxiety into legal scholarship by arguing that re-identification science had undermined the faith placed in anonymisation6. The lasting contribution of that period was not the proposition that anonymisation is impossible, because that claim was always too absolute, but the recognition that release must be understood as an adversarial and dynamic process rather than as a one-off act of deletion. The less productive legacy was a zero-risk instinct under which any conceivable future linkage became a reason to keep data permanently inside the personal data category, even where the relevant actor lacked the information, capacity, access or lawful means required to identify anyone.
Legal scholarship responded in two directions that are often presented as opposites but should instead be read together. Rubinstein and Hartzog argued that anonymisation policy should focus on a disciplined process of reducing re-identification and attribute-disclosure risk rather than on a guarantee of perfection6, while Schwartz and Solove proposed a continuum attentive to the malleability of identifying information7. At the same time, Nadezhda Purtova warned that an ever-expanding conception of personal data could turn European data protection into the “law of everything”7, and Michèle Finck and Frank Pallas showed that the legal-computational boundary between personal and non-personal data remained deeply uncertain8. Inge Graef, Raphaël Gellert and Martin Husovec went further by criticising the very policy reliance on an illusive non-personal-data category, particularly where mixed datasets and changing technical capacities make the line unstable8. The sensible synthesis is neither the abolition of the boundary nor a promise of irreversible anonymity; it is a taxonomy that expresses degrees of scope, provenance and assurance without abandoning the legal question of whether a person is identifiable for the actor in question.
The Union nevertheless chose a negative legislative category when it adopted Regulation 2018/1807 on the free flow of non-personal data2. The Regulation defined its subject by opposition to GDPR personal data and concentrated on localisation restrictions, regulatory access and cloud portability rather than on the internal structure of the category. The Commission’s 2019 guidance attempted to give the category more substance by dividing it according to origin: data that never related to an identified or identifiable natural person, such as weather readings or industrial-machine maintenance data, and data that began as personal but was later made anonymous2. It also listed sufficiently aggregated statistics as a familiar example, while acknowledging that mixed datasets represent the majority of datasets in the data economy and that, where the personal and non-personal parts are inextricably linked, the GDPR applies to the operational whole.
Although that two-origin model was a useful beginning, the subsequent expansion of the digital acquis has made its limits increasingly obvious: the Data Governance Act3 and the Data Act3 repeated the residual definition, the European strategy for data made common sectoral data spaces a central economic project4, and the AI Act3 adopted the same umbrella while adding a revealing phrase in Article 59, under which personal data may be processed in an AI regulatory sandbox only where the relevant requirements cannot effectively be met by processing “anonymised, synthetic or other non-personal data”. The grammar matters because it treats anonymised and synthetic data as possible members of a broader non-personal category, yet it does not say that either label is self-executing. Synthetic data remains personal if it relates to an identified or identifiable person, just as aggregate data remains personal where an outlier, linkage or de-aggregation permits a specific and meaningful inference, which means that the legislation has accumulated examples faster than it has developed a coherent classificatory method.
Image 2 — Four Questions, One Bad Label

Legal status, provenance, access modality and assurance are distinct questions; confusing them is how synthetic, aggregate, pseudonymised or clean-room data acquires a legal status that the underlying analysis has not earned.
The EDPB’s Contextual Turn Solves One Problem While Exposing Another
The relational analysis that the legislation left implicit has been supplied gradually by the Court of Justice, whose case law now makes it difficult to defend either a wholly intrinsic or a wholly subjective account of personal data. Breyer9 established that a dynamic IP address could be personal for an online media service where that service had legal means reasonably likely to be used to obtain identifying information from an internet provider; Scania9 and IAB Europe10 confirmed that identifiability can depend on information or methods distributed across actors; and OC v Commission10 required the relevant audience to include investigative journalists where a press release was designed for public dissemination. EDPS v SRB10 then confirmed two propositions that must be held together:
pseudonymised information is not necessarily personal data for every person in every circumstance, but the applicable perspective can be fixed by the particular GDPR obligation, so the SRB’s transparency duty had to be assessed from the controller’s perspective at collection rather than Deloitte’s perspective after transfer.10
Guidelines 02/2026 convert that jurisprudence into a structured framework by beginning with the question “for whom is the data intended to be anonymous?”, identifying relevant entities and applicable perspectives, and then testing No Record Isolation, No Linkage and No Inference against the means reasonably likely to be used in that context. The analysis remains objective because the actor’s preferences do not determine the outcome; cost, time, access, auxiliary information, legal powers, technology, foreseeable developments and the surrounding relationships do. Contextuality, therefore, does not mean subjectivity, and a dataset does not become anonymous merely because a recipient promises not to look; rather, its legal status is indexed to an evidenced relationship between an artefact and an actor, which is a familiar mode of legal analysis even if data lawyers have often spoken as though anonymity must be an intrinsic property of the file.1
The terminology nevertheless needs repair because there is a legitimate intuition that “anonymous information” should describe information capable of circulation without recalculating status for every recipient, whereas “non-personal for a particular entity” describes a narrower, recipient-specific conclusion. The EDPB now uses the former expression for both, provided the applicable perspective has been correctly identified, and that usage is defensible under the mirror structure of Article 4(1) and Recital 26 because data that is not personal for an entity is, from that entity’s legally relevant perspective, anonymous. Nevertheless, using one word for both a broadly distributable state and a recipient-bounded state invites exactly the confusion that the Article 6(11) DMA debate has exposed. The solution is to retain the legal conclusion while attaching a scope label: general-scope anonymised data where the assessment covers all relevant entities within the intended circulation environment, and recipient-scope anonymised data where the conclusion is confined to identified independent recipients or a defined recipient class.1
The limits built into the guidelines demonstrate why this is not a licence to launder personal data. A processor does not obtain its own independent anonymity perspective where it acts on behalf of a controller; the controller’s perspective follows the processing, because otherwise outsourcing would become a way of evading the GDPR. Likewise, the Board states that if transfer to another entity is reasonably likely and it cannot be ruled out that the recipient has means reasonably likely to identify, the data must be treated as personal for the transfer and subsequent processing. Contractual prohibitions may alter the factual assessment but are not prohibitions by law and must complement technical measures rather than substitute for them. Those rules make recipient-scope anonymity narrower than its critics suggest and also show why it should never be represented as a free-floating property that automatically survives onward transfer.1
The Board’s No Inference analysis contributes another distinction that any serious taxonomy must preserve, because it separates a specific and meaningful inference about a person represented in the original data from population-level knowledge that can be applied to someone who was never part of the source dataset. Its bank-loan example explains that an anonymised historical dataset may reveal a general association between attributes and default risk, and that applying the association to a new applicant creates personal data about the applicant without converting the historical dataset back into personal data. This is a crucial conceptual separation between individual connection and general learning, although the Board is equally clear that aggregates, synthetic data and AI models can fail where membership, extraction, prompting or reconstruction reveals source-specific information.1
Four Questions that Europe Keeps Collapsing Into One
A workable taxonomy should, therefore, begin by separating four questions that European policy repeatedly collapses into a single label. The first concerns legal status: does the artefact relate to an identified or identifiable natural person from the applicable perspective? The second concerns provenance or form: was it observed directly, transformed from personal records, aggregated into statistics, generated synthetically, or derived as a model, feature, embedding or output? The third concerns circulation: will it be published, exported, queried through an API, held inside a trusted research environment, or exposed only through privacy-tested outputs? The fourth concerns assurance: which technical tests, access controls, contractual duties, audits, monitoring systems and reassessment triggers support the status claim? Only the first question determines whether the GDPR applies to the artefact for the actor under examination, while the other three provide evidence, context and risk management without becoming status by themselves.
Separating those questions makes several persistent errors easier to diagnose because each of the familiar privacy terms performs a different legal or technical function. Pseudonymisation is a safeguard applied to personal data and does not become non-personal simply because the recipient lacks the key; encryption controls intelligibility and access but remains designed to be reversible; aggregation describes form but can still disclose outliers or enable de-aggregation; synthetic describes how an artefact was produced but says nothing conclusive about memorisation or individual-level inference; and a clean room describes where processing occurs rather than what the underlying data is. Contracts, audit and purpose restrictions can make a chain of identification less reasonably likely and can be indispensable to a responsible access regime, but they do not perform the technical transformation that the word anonymisation implies. These are valuable legal and engineering tools, yet they become false friends when converted into status labels.
The distinction also prevents the opposite error, in which every downstream artefact is presumed personal because personal data appeared somewhere in its lineage. A national statistic is not personal merely because tax records helped produce it; a model is not necessarily personal merely because it was trained on personal data; and a synthetic dataset is not inevitably personal merely because its generator learned from individual records. Lineage creates a reason to test, document and monitor, but the legal inquiry still asks whether the artefact itself relates to an identified or identifiable person from the applicable perspective. Treating lineage as permanently contaminating would reproduce Purtova’s law-of-everything claim and would make the AI Act’s preference for anonymised, synthetic or other non-personal alternatives largely meaningless.
The resulting structure is better understood as a status stack than as a binary sticker: at the top sits the legal conclusion, personal or non-personal for the specified actor and operation; beneath it sits the route by which the conclusion was reached; beneath that sits the circulation environment that defines access and relevant adversaries; and beneath that sits the assurance record that makes the claim auditable and time-bounded. An organisation should therefore be able to say considerably more than “this dataset is anonymous”, explaining instead that a particular version is recipient-scope anonymised for an accredited research institute, under a defined access modality, after specified linkage and inference testing, with reassessment triggered by a new auxiliary dataset, material system change or security incident.
The Six Routes and the Quarantine

Six different routes can lead into the non-personal-data umbrella, while mixed or unresolved data remains in a quarantine category until separation or a defensible status assessment is possible.
Six Routes into the Umbrella, with a Quarantine Category Beside It
The first route, N0, covers native non-personal data: information that never related to an identified or identifiable natural person by content, purpose or effect. Weather readings from remote sensors, machine-vibration measurements and many forms of industrial telemetry will often fall here, although the label should not be applied mechanically because seemingly technical data can become personal where it is used to evaluate a worker, driver, household or device user. The category is, therefore, defined by the absence of the personal nexus rather than by the fact that the source is a machine, retaining the Commission’s 2019 origin distinction while making the content-purpose-effect analysis explicit.2
The second route, N1, covers statistical or population-level non-personal data, including aggregates, indicators, coefficients and generalised findings in which the connection to individual records has been left behind and no specific and meaningful inference can be made about a person represented in the source material. It includes an average, rate or distribution only where outliers, small cells, overlapping tables and query patterns do not permit reconstruction, and it must therefore be tested rather than assumed. The category is especially important because statistics can remain socially consequential, commercially valuable or discriminatory without being personal data, which is why leaving the GDPR does not mean leaving law, ethics or fundamental-rights scrutiny.1
The third route, N2, covers general-scope anonymised data: information that began as personal data but has been transformed so that the likelihood of identification is insignificant in reality for every relevant entity within the intended circulation environment. Where the intended environment is public release, the set of relevant actors and auxiliary information is broad; where circulation is restricted to a defined professional network, the scope is narrower but still extends beyond a named recipient. N2 is the closest practical analogue to the conventional idea of “anonymous data” as a distributable object, although its status remains evidenced and revisable rather than metaphysically irreversible.1
The fourth route, N3, covers recipient-scope anonymised data, where the artefact remains personal for one actor, commonly the source controller, but is anonymous for a specified independent recipient or defined recipient class because that recipient lacks means reasonably likely to identify and is not acting on the controller’s behalf. The hospital-and-independent-research-institute example in the guidelines sits here, as can some data-space and DMA Article 6(11) access arrangements if the recipient-specific technical and institutional conditions genuinely make identification insignificant. N3 must always carry the identity or class of the recipient, the applicable role analysis, the access environment and the onward-transfer boundary; without those qualifiers, a narrow conclusion is liable to be mistaken for N2 and circulated beyond the conditions that made it true.1,11
The fifth route, N4, covers synthetic non-personal data, although the word synthetic describes generated rather than directly observed records and therefore establishes provenance rather than legal status. The non-personal conclusion follows only after testing whether the artefact reproduces source records, reveals membership, memorises rare sequences or permits specific and meaningful inference about a person. A synthetic sample that merely illustrates schema and distributions may qualify, while a high-fidelity generator trained on sparse health or search data may not. The category consequently requires method transparency, privacy and memorisation testing, and a statement of the intended recipient context, so that the adjective synthetic is treated as provenance until non-personal status has been earned.1,3
The sixth route, N5, covers derived non-personal artefacts such as models, embeddings, features, benchmarks, rules, coefficients and privacy-tested outputs where the artefact no longer supports source-specific identification, membership or meaningful individual inference. This category gives the taxonomy a place for the objects through which modern AI and analytics actually circulate, instead of forcing every debate back into the language of row-level datasets. It also prevents the category from becoming a model-washing device, because the same output must be tested for extraction, regurgitation, reconstruction and inference, with the status claim attached to the particular version and access interface rather than to the model family in the abstract.1
Alongside those six routes, the taxonomy requires an X category that quarantines mixed, inseparable or unresolved data rather than allowing uncertainty to be translated into a convenient non-personal label. A bundle containing personal and non-personal components should not be represented as wholly non-personal merely because most fields are technical or aggregate, and the Commission’s 2019 guidance correctly observed that the GDPR applies across an inextricably linked mixed dataset where the parts cannot be processed separately. The EDPB applies a similar operational rule where a dataset contains both anonymous and personal records: unless the parts are treated separately, the dataset must be handled as containing personal data. Quarantine is not a permanent legal category but an instruction to separate, suppress, transform or complete the assessment before wider circulation.1,2
From Taxonomy to Infrastructure: the TRACE Data Status Passport
A taxonomy becomes useful only when organisations can attach it to real artefacts, which is why the classificatory proposal should be paired with a Data Status Passport. I call the proposed structure TRACE: Type and lineage; Relation to natural persons; Actor and applicable perspective; Circulation modality; and Evidence, expiry and reassessment. TRACE does not replace a legal opinion or technical test, but it forces the conclusion into a form that can be inspected, compared and reused across regulators, data spaces, research environments and AI supply chains. Its central discipline is that every non-personal-data claim must state what the artefact is, for whom the status is asserted, how it may circulate, and what evidence keeps the assertion valid.
The Type field records whether the artefact is native, transformed, aggregate, synthetic or derived, together with its source lineage and version. Relation records whose circumstances, conduct or characteristics the information concerns by content, purpose or effect, including whether a model or statistic merely supports population-level knowledge or preserves a connection to source individuals. Actor identifies the holder, recipients, processors, possible onward recipients and relevant adversaries, and it states the applicable perspective under the governing legal obligation. Circulation records whether the artefact is public, exportable, API-mediated, held in a trusted environment or output-only, since the accessibility of both the given data and auxiliary information changes the means reasonably likely to be used. Evidence and expiry records the tests, threat model, controls, audit, safety margins, known auxiliary datasets and triggers that require the status to be revisited.
Because the passport must travel with the artefact across technical and institutional boundaries, it should be readable both by people and by machines. A short label might state: “N3 recipient-scope anonymised; accredited online-search-engine recipients; controlled API; no local copy; linkage, membership and extraction tests passed on version 3.2; review after ninety days or on material change.” Behind that label would sit the field map, transformation record, role analysis, red-team results, utility assessment and incident history. This is not bureaucratic decoration, because it allows a receiving organisation to understand the scope of the claim without reverse-engineering the source controller’s entire assessment, while giving regulators a common object for certification, challenge and reassessment.
TRACE also supplies a disciplined answer to temporal instability by converting the EDPB’s correct observation that data can move back into the personal category into an operational set of expiry rules. The Board’s instruction to reassess “where possible and appropriate” remains too open-ended for infrastructure, whereas a passport can specify sunset dates and event-based triggers, including a security incident, a new public dataset, a material model update, a new recipient class, a change in processor architecture, evidence of successful extraction or a major reduction in the cost of relevant attacks. The status claim then becomes durable enough to support investment and reuse without being fictionalised as eternal.1

TRACE turns “anonymous” from a sticker into an auditable claim by recording type, relation, actor, circulation and the evidence and expiry conditions supporting the status.
Article 6(11) DMA is the Test Case Europe Needed
Article 6(11) of the Digital Markets Act reveals why this distinction is more than academic. The provision requires a gatekeeper operating an online search engine to provide rival search engines with access on fair, reasonable and non-discriminatory terms to ranking, query, click and view data generated by end users, while stating that any query, click and view data that constitutes personal data shall be anonymised. Recital 61 adds the contestability objective and instructs the gatekeeper to protect personal data against re-identification by appropriate means such as anonymisation without substantially degrading quality or usefulness. The legal sequence is important: utility should be preserved after the anonymisation condition has been met, rather than being used to invent a weaker DMA-specific meaning of the word.11
My submission to the Commission on the Google Search specification proceeding argued that controlled access cannot substitute for anonymisation.12 Where common-tail search data can be technically anonymised for ordinary export, it should be exported; where an already anonymised artefact presents marginally elevated operational or combination risk under ordinary download, controlled access can provide an additional layer; and where the underlying data remains identifiable, moving it into a clean room does not convert it into non-personal data. The correct alternatives are suppression, aggregation, regulator escrow, or privacy-tested model and output access in which the object leaving the environment is itself tested against memorisation, extraction, reconstruction and re-identification.
The proposed taxonomy makes those routes explicit by assigning different Article 6(11) artefacts to different status paths rather than pretending that every useful signal should emerge as the same kind of dataset. A privacy-preserving aggregate API can produce N1 statistical data; an exportable transformed common-tail dataset may qualify as N2 or N3 depending on the relevant recipient universe; a synthetic sample is N4 only after privacy and memorisation testing; and a benchmark, feature or model output may qualify as N5 if source-specific leakage has been excluded. A clean room containing identifiable rare-query or sequence data remains governed personal data, even where its access architecture is exemplary, while a mixed dataset containing personal query text and non-personal ranking signals belongs in quarantine until fields are separated or transformed. The access modality may reduce risk and support compliance, but it should never be allowed to silently rewrite the status field.11,12
The same structure explains the role of the Anonymisation and Utility Impact Assessment proposed in my submission. An AUIA would record the data inventory, actor analysis, threat model, technical parameters, utility tests, recipient controls, red-team evidence and reassessment triggers for each access tier. Within TRACE, that assessment populates the Evidence field and shows whether the claimed route has been achieved before the Commission asks whether contestability value has been preserved. It also supports the principle that Article 6(11) should deliver contestability value rather than raw-form parity, because different artefacts can serve different search functions through the least risky modality that preserves genuine competitive utility.12
Search data demonstrates particularly clearly why non-personal status cannot be treated as the end of the rights inquiry, since a query can reveal illness, debt, fear, vulnerability, political anxiety or another person’s condition, while a generalised search signal may shape ranking, visibility and market access even after no individual can be identified. My broader work on relational rights argues that digital law is too often organised around the sovereign user’s possessive grammar of “my data” and “my choice”, when information and technical decisions structure the conditions under which other people exercise rights.13 The taxonomy should therefore stop GDPR category error without pretending that non-personal data is normatively empty; competition law, the AI Act, confidentiality, trade-secret law, consumer protection, non-discrimination and public-law review continue to govern what the artefact can do.

Article 6(11) requires different routes for different artefacts; aggregate, transformed, synthetic and output layers may become non-personal, while identifiable data remains personal even inside a clean room.
What the Omnibus Should Now Do
The Digital Omnibus offers a rare opportunity because the Commission already proposes to repeal most of the Free Flow of Non-Personal Data Regulation and consolidate its surviving localisation principle and other data-acquis rules into the Data Act, yet structural consolidation without conceptual consolidation would squander much of the value of that exercise. The new framework should retain the broad umbrella definition but add an annexed taxonomy stating that non-personal data may be native, statistical, general-scope anonymised, recipient-scope anonymised, synthetic or derived, while mixed and unresolved data remains subject to the personal-data regime to the extent that the personal component cannot be separated.14
The Omnibus should also require scope labels that prevent a recipient-specific conclusion from masquerading as an unrestricted claim about the artefact. An organisation relying on recipient-scope anonymity should not describe an artefact simply as “anonymous” in documentation, licences or marketing, because that wording invites onward circulation beyond the assessed context. The label should identify the recipient or class, role, access mode and reassessment boundary, while general-scope release should state the circulation environment that formed part of the assessment. This would preserve the EDPB’s contextual legal standard while answering the legitimate concern that anonymous information should not be confused with a narrow conclusion of non-personality in one bilateral relationship.1
A third reform should create recognised assurance pathways and rebuttable presumptions around TRACE rather than around particular techniques. Certification should not declare that k-anonymity, differential privacy, synthetic generation or a clean room is always sufficient; it should confirm that the organisation has correctly identified the route, mapped relevant actors and auxiliary data, applied appropriate tests, documented the circulation modality and adopted meaningful reassessment triggers. Sectoral profiles could then specify stronger evidence for genomic, location, financial, communications or search data, while SMEs and research bodies could reuse standard methods instead of commissioning a fresh legal-engineering project every time.
A fourth reform should make mixed data and derived artefacts first-class objects of the acquis, since the current legal architecture repeatedly assumes a dataset while AI systems increasingly circulate embeddings, weights, features, evaluation sets, retrieval indices and outputs whose status cannot be inferred from their source alone. The Omnibus should require output-specific leakage testing where a non-personal conclusion is claimed, clarify that controlled access to personal data remains personal-data processing, and support privacy-preserving compute-to-data models without allowing infrastructure labels to perform the work of anonymisation. The AI Act’s own distinction between anonymised, synthetic and other non-personal data provides the legislative foothold for this broader treatment.3
The taxonomy should finally be tied to institutional accountability, because my Innovation Mandate’s central claim is that Europe has built a digital supervisory state whose guidance, remedies, priorities and timelines operate as economic infrastructure, and that the answer is not deregulation but a stronger discipline of public justification.15 A cross-acquis data taxonomy would make that discipline concrete: the EDPB, Commission, AI Office, data-space authorities and national regulators would have to state which route they recognise, what evidence they require, how status changes across actors, and why the chosen burden is proportionate to the rights risk. That is a better settlement than either maximalist ambiguity, which keeps almost everything personal by caution, or permissive relabelling, which calls a controlled dataset anonymous because the policy objective requires it.
Europe does not need a third category sitting vaguely between personal and non-personal data, because the legal test can remain binary at the point of application. What it needs is a richer account of how non-personal status is reached, scoped, evidenced and maintained. The six routes, the quarantine category and the TRACE passport supply that account without weakening the GDPR’s rights floor, while also giving researchers, AI developers, data-space operators, competition authorities and courts a common language for artefacts that are currently discussed past one another. The EDPB has supplied the doctrinal hinge by recognising that anonymity is contextual but objective; the next Omnibus should turn that hinge into an architecture that others can actually use.1,15
Sources and references
1. European Data Protection Board, Guidelines 02/2026 on Anonymisation, version 1.0, adopted for public consultation on 7 July 2026, especially pp. 2-18, 21-29 and Annex 1 at p. 33.
2. Regulation (EU) 2018/1807 on a framework for the free flow of non-personal data, especially Recitals 8-10 and Articles 1-3; European Commission, COM(2019) 250 final, Guidance on the Regulation on a framework for the free flow of non-personal data in the European Union, especially sections 2.1-2.2.
3. Regulation (EU) 2022/868 (Data Governance Act), Article 2(4); Regulation (EU) 2023/2854 (Data Act), Article 2(4); Regulation (EU) 2024/1689 (AI Act), Articles 3(51) and 59(1)(b).
4. European Commission, A European strategy for data, COM(2020) 66 final; European Parliament, The emergence of non-personal data markets, PE 740.098, October 2023.
5. Latanya Sweeney, “k-Anonymity: A Model for Protecting Privacy” (2002) 10 International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems 557; Arvind Narayanan and Vitaly Shmatikov, “Robust De-anonymization of Large Sparse Datasets” (2008) IEEE Symposium on Security and Privacy 111.
6. Paul Ohm, “Broken Promises of Privacy: Responding to the Surprising Failure of Anonymization” (2010) 57 UCLA Law Review 1701; Ira Rubinstein and Woodrow Hartzog, “Anonymization and Risk” (2016) 91 Washington Law Review 703.
7. Paul M Schwartz and Daniel J Solove, “The PII Problem: Privacy and a New Concept of Personally Identifiable Information” (2011) 86 NYU Law Review 1814; Nadezhda Purtova, “The Law of Everything: Broad Concept of Personal Data and Future of EU Data Protection Law” (2018) 10 Law, Innovation and Technology 40.
8. Michèle Finck and Frank Pallas, “They Who Must Not Be Identified: Distinguishing Personal from Non-Personal Data under the GDPR” (2020) 10 International Data Privacy Law 11; Inge Graef, Raphaël Gellert and Martin Husovec, “Towards a Holistic Regulatory Approach for the European Data Economy” (2019) 44 European Law Review 605.
9. Case C-582/14 Breyer v Bundesrepublik Deutschland EU:C:2016:779; Case C-434/16 Nowak EU:C:2017:994; Case C-319/22 Gesamtverband Autoteile-Handel v Scania EU:C:2023:837.
10. Case C-604/22 IAB Europe EU:C:2024:214; Case C-479/22 P OC v Commission EU:C:2024:215; Case C-413/23 P EDPS v SRB EU:C:2025:645.
11. Regulation (EU) 2022/1925 (Digital Markets Act), Article 6(11) and Recital 61.
12. Dr M.R. Leiser (Mark), Making Access to Google Search Data Work Without Inventing a Privacy Fiction: Submission to the European Commission, Case DMA.100209 - SP - Alphabet - Article 6(11) DMA (DigiData Consulting Ltd, April 2026), especially pp. 2-10 and the proposed Anonymisation and Utility Impact Assessment.
13. Dr M.R. Leiser (Mark), “Against the Sovereign User: Fundamental Rights, Technology and Relational Rights” (work in progress, DigiData Consulting Ltd, 2026), including the Article 6(11) search-data case study.
14. European Commission, Proposal for a Digital Omnibus, COM(2025) 837 final, especially the explanatory memorandum and proposed consolidation of the Free Flow of Non-Personal Data Regulation, Data Governance Act and Open Data Directive into the Data Act.
15. Dr M.R. Leiser (Mark), The Innovation Mandate lecture series (DigiData Consulting Ltd, 2026), especially slides 3-6 and 12-16, on digital regulation as economic infrastructure, supervisory discretion, proportionality, assurance and evidence.
Publication note. This DigiData version reproduces an essay originally published through LawBhoy on Substack. The Substack page is identified as the canonical source for the existing post.