Big Data and Data Protection: A GDPR Guide

·

Big data and data protection are not at odds, provided a business first defines why it needs the information, which legal basis allows it to use that information, and what risks the processing creates for individuals. Volume does not lower GDPR's bar — if anything, it raises the stakes for governance, minimisation, security and the ability to demonstrate every decision.

Big data and data protection: start with purpose and legal basis

An analytics project should not start by hoovering up every available dataset "just in case it turns out useful." The starting point is a specific, legitimate and understandable purpose — detecting fraud, forecasting demand, reducing churn, segmenting communications or improving a process. Only then should the business work out which data is genuinely needed and which basis under Article 6(1) GDPR legitimises each processing activity.

The most common legal bases are:

Where special categories of data are involved — health, biometrics used for identification, political opinions or other data under Art. 9(1) — a valid exception under Art. 9(2) is required in addition to an Art. 6 basis. The use of data on criminal convictions and offences is governed separately by Art. 10.

The information given to individuals must satisfy Art. 13 or Art. 14 GDPR, depending on whether the data is collected directly or indirectly. In complex settings, layered notices can be used, but the first layer must clearly explain the main purposes, the relevant logic behind any profiling, and how to exercise rights.

Before integrating sources, our data governance service helps build an inventory of datasets, purposes, owners, quality, provenance, retention periods and permissions. That traceability keeps a data lake from turning into an uncontrolled repository.

Minimisation, purpose limitation and data quality

Article 5(1)(b) GDPR requires data to be collected for specified, explicit and legitimate purposes and not further processed in an incompatible manner. Art. 5(1)(c) imposes minimisation: data must be adequate, relevant and limited to what is necessary. Art. 5(1)(d) adds accuracy, and Art. 5(1)(e) limits retention.

In practice, an SME can translate these principles into concrete controls:

  1. Define the business question and drop variables that do not deliver a demonstrable improvement.
  2. Separate direct identifiers from the analytical dataset and restrict access to the mapping table.
  3. Set rules for quality, provenance and refresh cycles — a massive but biased dataset produces massive, biased decisions.
  4. Set retention periods per purpose and automate review, restriction or erasure.
  5. Avoid incompatible reuse. If a new purpose emerges, its compatibility must be assessed under Art. 6(4) GDPR, or a separate valid basis must be found.
  6. Work with samples, aggregates or synthetic data whenever they achieve the objective with less impact.

Data protection by design and by default under Art. 25 requires these decisions to be built in before the platform is procured or the model is trained. By default, access should not extend to the whole organisation, nor should everything be retained indefinitely. Art. 32 requires security measures appropriate to the risk: access control, encryption, activity logging, backups, testing and incident response, among others.

When big data requires a DPIA

A Data Protection Impact Assessment (DPIA) is not a box-ticking document. Under Art. 35(1) GDPR, it must be carried out before processing begins whenever the processing is likely to result in a high risk to individuals' rights and freedoms, taking into account its nature, scope, context and purposes.

Art. 35(3) expressly mentions, among other cases, systematic and extensive evaluation of personal aspects based on automated processing — including profiling — used for decisions with legal or similarly significant effects, and large-scale processing of special categories or criminal-offence data.

There is no universal headcount that defines "large scale." The number or proportion of individuals affected, the volume and variety of data, the duration and the geographical extent must all be weighed together. The AEPD (Spanish DPA) and the European Data Protection Board (EDPB) also point to further criteria: systematic monitoring, combining datasets, vulnerable individuals, innovative technology, highly personal data, or anything that prevents someone from exercising a right or accessing a service.

As a prudent rule, if several high-risk criteria are present, the decision should be documented and a DPIA should normally be carried out. Its minimum content, under Art. 35(7), includes:

A DPIA is a living process. It must be reviewed whenever the data, purpose, algorithm, recipients, infrastructure or risk level change. If, after mitigation, a high residual risk remains, prior consultation with the supervisory authority is required under Art. 36. Where a DPO has been appointed, they must provide advice, per Arts. 35(2) and 39(1)(c).

Anonymisation versus pseudonymisation

The two are not equivalent. Pseudonymisation, defined in Art. 4(5) GDPR, means data can no longer be attributed to a person without additional information, provided that information is kept separate and protected. It reduces risk, but the data remains personal data and the GDPR keeps applying.

Anonymisation aims to make a person no longer identifiable by any means reasonably likely to be used, taking into account cost, time, technology and auxiliary sources. If it is genuinely irreversible, the information stops being personal data. Replacing names with codes, removing ID numbers or applying a plain hash is usually pseudonymisation, not anonymisation.

In big data, the risk of re-identification by combining seemingly harmless variables grows. That is why singling out, linkability and inference should be assessed; generalisation, suppression, aggregation or perturbation applied; and the risk checked periodically. The AEPD's anonymisation guidance stresses minimising re-identification without destroying utility, but no method should be presented as an absolute guarantee regardless of context.

Who holds the key and what external sources exist also matter. The controller must document the technique, the assumptions, the tests performed and any reuse restrictions. Publishing an "anonymised" dataset demands a far more demanding assessment than using it in a controlled internal environment.

Profiling and automated decisions

Profiling is automated processing used to evaluate personal aspects, as defined in Art. 4(4) GDPR. It can analyse performance, economic situation, preferences, reliability, behaviour, location or movements. Not every profile is prohibited, but each one needs a purpose, a legal basis, transparency, minimisation and bias controls.

Art. 22(1) recognises the right not to be subject to a decision based solely on automated processing that produces legal effects or similarly significantly affects the individual. Potential examples include automatically refusing an essential service, excluding someone from an opportunity, or making a significant employment decision without genuine human involvement.

The exceptions under Art. 22(2) are limited: necessity for entering into or performing a contract, legal authorisation with safeguards, or explicit consent. In the contractual and consent scenarios, individuals must at least have the right to obtain human intervention, to express their point of view and to contest the decision — Art. 22(3). Decisions must not be based on special categories of data, except under the strict conditions of Art. 22(4).

A human review cannot be a rubber stamp. The reviewer must have the competence, information, time and authority to change the outcome. In addition, Arts. 13(2)(f), 14(2)(g) and 15(1)(h) require meaningful information about the logic involved, and the significance and envisaged consequences, where applicable.

International transfers and vendors

Using a cloud service, an analytics API or remote support can involve access from outside the European Economic Area. Arts. 44 to 49 GDPR require that the transfer does not lower the level of protection.

The practical order is to check whether an Art. 45 adequacy decision exists. If not, the appropriate safeguards under Art. 46 can be used, such as standard contractual clauses, accompanied by a transfer impact assessment and supplementary measures where necessary. The Art. 49 derogations are meant for specific situations, not as a structural solution for routine data flows.

The contract with the processor must satisfy Art. 28: subject matter, duration, instructions, confidentiality, security, sub-processors, assistance with rights requests, return or deletion of data, and audit rights. The business needs to know hosting regions, backups, telemetry, support arrangements, sub-processors and whether data is used to improve the vendor's own services. A contract is no substitute for technical verification.

Accountability: demonstrate, review and improve

Art. 5(2) GDPR requires compliance with the principles and the ability to demonstrate it. Art. 24 requires appropriate measures that can be reviewed. In big data, accountability rests on evidence:

Infringing these principles or rights can trigger the fines under Art. 83(5) GDPR: up to €20 million or, for an undertaking, up to 4% of total worldwide annual turnover for the preceding financial year, whichever is higher. The point of sound governance is not to work out of fear of sanctions, but to reduce harm, improve decisions and preserve trust.

Summum Consultoría can support an SME through the inventory, the legal basis, the DPIA, the contracts and ongoing oversight. Where processing calls for continuous advice, our outsourced DPO service provides an independent, risk-oriented function without taking over the company's own responsibility.

Frequently asked questions

Does the GDPR apply if we use publicly available data?

Yes, if the data can identify individuals. The fact that data is accessible does not authorise any use of it — a purpose, a legal basis, transparency and respect for individuals' rights are still required.

Is consent needed for every big data project?

No. Another basis under Art. 6(1) GDPR may apply, but it must fit the purpose. Legitimate interest requires necessity and a balancing test; a contract only covers processing that is objectively necessary.

Does every large-scale processing activity require a DPIA?

Art. 35 requires one when a high risk is likely, and expressly for large-scale special categories or criminal-offence data. Profiling, monitoring, combined data sources, vulnerability and innovative technology must also be considered.

Are pseudonymised data outside the scope of the GDPR?

No. They remain personal data because they can still be attributed with additional information. Pseudonymisation is a security and design safeguard, not an exemption.

Can we use an analytics vendor located outside the EU?

It can be possible if GDPR Chapter V is applied: adequacy or appropriate safeguards, an assessment of the context, and supplementary measures where needed. The Art. 28 processing agreement must also be put in place.