themes

Ethics assessment of technology

In a data-driven society, healthcare is no exception. The datafication of the healthcare system can generate value for the individual patients themselves, known as the “primary use of health data”, when data is used for direct patient care and healthcare operations; but it can also create benefits at a societal level, boosting scientific research, innovation, policymaking, AI training, and regulatory activities – known and “secondary use of health data” or “further processing”. The secondary use is seen as crucial for medical progress, as it helps "expand society’s collective knowledge" and enhances healthcare quality and effectiveness. Scientific research leverages this data, especially large datasets vital for AI development and studying rare diseases and childhood cancers, where individual institutions often lack sufficient data to produce significant results. Another example was the COVID-19 pandemic, as it highlighted the urgent need for authorities to access data for evidence-based policy-making and public health tracking. Ethically, this is often justified by the "Veil of Ignorance" argument: if we were unaware of our own health status, we would logically support sharing data to boost collective medical knowledge for the common good. This article explores the legal landscape that involves the secondary use of health data in the EU, in particular the General Data Protection Regulation (GDPR) and the European Health Data Space (EHDS).

What does the law say

The EU legal framework for secondary use of health data is essentially structured around the General Data Protection Regulation (GDPR) and, increasingly, the European Health Data Space (EHDS). Together, these instruments reflect the EU’s attempt to reconcile two potentially conflicting objectives: the protection of individual privacy and informational self-determination, on the one hand, and the facilitation of data-driven innovation, scientific research, and public health governance, on the other. Under the GDPR, health data is classified as a “special category” of personal data according to Article 9(1), meaning that its processing is considered more sensitive and warrants stronger protection. In principle, processing of special categories of personal data is prohibited unless one of the exceptions listed in Article 9(2) applies, including, for example, explicit consent of the data subject, for scientific research purposes, or reasons of public interest in public health.  The GDPR does not prohibit all forms of secondary use of personal data. Although Article 5(1)(b) establishes the principle of purpose limitation (i.e., personal data should be collected for specified, explicit, and legitimate purposes and not reused for incompatible purposes), the GDPR is more flexible regarding research and public-interest activities. In particular, the GDPR provides that further processing for scientific research purposes, statistical purposes, or archiving in the public interest is generally presumed to be compatible with the original purpose for which the data was collected. As a result, health data initially gathered in the context of clinical care may, under certain conditions, subsequently be reused for research or public health purposes without automatically violating the purpose limitation principle.  However, such secondary use remains subject to identifying a legal basis under Articles 6 and 9(2), complying with the broader principles of proportionality, necessity, transparency, data minimization, and storage limitation, as well as strict safeguards, including technical and organizational measures designed to protect the rights and freedoms of data subjects, as required under Article 89 GDPR. Furthermore, Member States retain significant discretion in regulating research-related derogations, particularly under Articles 9(4) and 89 GDPR. Therefore, the practical conditions for secondary use continue to vary substantially across the Union, especially concerning ethics approvals, biobank governance, public-sector data access, and national health registries. This fragmentation has long been criticized as an important obstacle to cross-border health research and large-scale data-driven innovation in Europe. In light of this scenario, the European Health Data Space was introduced as part of the EU data strategy to address these structural limitations through a dedicated sector-specific framework for health data governance. The EHDS establishes a harmonized infrastructure for both the primary and secondary use of electronic health data across Member States.  The EHDS does not come to replace the GDPR, which remains the foundational legal framework; instead, the EHDS operationalizes and institutionalizes access mechanisms for secondary use through a system of Health Data Access Bodies (HDABs), standardized permit procedures, interoperability requirements, and secure processing environments. The objective is to facilitate the lawful reuse of health data for purposes such as scientific research, innovation, public health, healthcare planning, policymaking, patient safety, and the development of artificial intelligence systems. The EHDS innovates in its governance-oriented approach. It detaches the necessity around individual consent to reuse the health data, and rather builds institutional oversight, controlled access environments, pseudonymization requirements, and purpose-based restrictions. In this way, access to the health data by researchers, public authorities, and certain private actors may be obtained without the patient’s individual consent but through permits issued by national access bodies based on legal obligation (Article 6(c) GDPR) or public interest, subject to compliance with specified safeguards and limitations.  The EHDS establishes a protection strategy as well: the regulated infrastructure prohibits the use of electronic health data for purposes such as advertising, discriminatory decision-making, harmful profiling, or decisions related to employment and insurance that may adversely affect individuals.

What is at stake

Despite robust legal frameworks, the boundary between "anonymous" and "identifiable" data is increasingly fragile. The vast availability of data and advances on AI and data analytics render the reliance on anonymization and pseudonymization as safeguards to reuse data reuse with privacy protection a hope more than a reality. Sophisticated AI tools can now discover complex patterns and correlations to de-anonymize datasets once thought to be safe. This is what happened, for example, to the UK Biobank. Researchers with access to the biobank health data inadvertently posted, in theory, anonymized patient files on public platforms like GitHub. The investigations revealed that confidential health records of over 400,000 participants were exposed online. The dataset did not contain patients’ names, addresses, birth dates, or any other identifier; still, the Guardian journalists were able to trace back the data to the person based on exposed online information, for example, data of surgery published on social media. Data that was deemed anonymized allowed identification. Later, reports confirmed that UK Biobank records had been listed for sale on Chinese websites, highlighting that data can be leaked through either accident or dishonesty, even within highly secure environments.  The rise of AI also exposes the limitations of newer privacy-preserving techniques such as synthetic data. Although synthetic datasets are designed to replicate the statistical properties of real patient records without directly reproducing identifiable information, they may still be vulnerable to attacks. Malicious actors may infer whether a particular individual’s data formed part of the original training dataset, thereby exposing sensitive information regarding that person’s health status or participation in research. These developments demonstrate that privacy risks may emerge not only through direct disclosure of data, but also through inferences generated by AI systems themselves. These developments also raise broader concerns regarding autonomy and public trust. Emerging frameworks such as the EHDS increasingly rely on institutional governance and administrative authorization rather than individualized consent. While such models may facilitate large-scale research and innovation, they also risk weakening individuals’ sense of control over their most sensitive information. In healthcare settings, patients often possess limited bargaining power and may effectively face “take-it-or-leave-it” conditions regarding the reuse of their data. If individuals perceive that privacy protection is subordinate to institutional or commercial interests, public confidence in health-data governance may deteriorate. Such a loss of trust could ultimately discourage participation in research and undermine the very scientific and public-health objectives that secondary-use regimes are intended to advance.

Between innovation and privacy

Ultimately, the regulation of the secondary use of health data in the European Union reveals the difficulty of reconciling technological innovation with the protection of individual rights. On the one hand, large-scale access to health data is increasingly seen as essential for scientific research, public health policy, and the development of AI-driven medical technologies. On the other hand, the growing ability of modern technologies to infer, connect, and re-identify information challenges traditional assumptions about anonymity and data security.  The GDPR and the EHDS attempt to respond to these tensions by creating frameworks that allow data sharing under stricter forms of oversight and governance. Yet legal safeguards and technical measures alone are unlikely to be sufficient if individuals lose trust in how their data is handled. The long-term success of the European model will therefore depend not only on facilitating research and innovation, but also on ensuring that patients continue to feel that their privacy, autonomy, and reasonable expectations remain genuinely respected. The future development of secondary-use frameworks in the European Union will depend on whether institutions can reconcile the growing demand for data-driven innovation with meaningful protection of individual rights and public confidence in the governance of health data.

References:

  • Becker, R., & Dove, E. S. (2026). The EU GDPR and secondary use of health and genetic data for research support purposes. International Data Privacy Law, Vol. 00, No. 00, 1–16. doi:10.1093/idpl/ipag001.
  • Lalova, T., Kindt, E., Chelioudakis, E., Verhenneman, G., Delforge, A., & Herverg, J. (2020). Study on the secondary use of personal data in the context of scientific research. Final Report (Contract No. EDPS/2019/02-04), prepared by Milieu for the European Data Protection Board (EDPB).
  • Shabani, M., & Yilmaz, S. (2022). Lawfulness in secondary use of health data: Interplay between three regulatory frameworks of GDPR, DGA & EHDS. Technology and Regulation, 128–134. doi:10.26116/techreg.2022.013.
  • Vallevik, V. B., Befring, A. K. C., Elvatun, S., & Nygård, J. F. (2026). Processing of synthetic data in AI development for healthcare and the definition of personal data in EU law. International Journal of Law and Information Technology, 34, eaag002. doi:10.1093/ijlit/eaag002.
  • European Data Protection Board (EDPB). (2021). EDPB Document on response to the request from the European Commission for clarifications on the consistent application of the GDPR, focusing on health research. Adopted on 2 February 2021.
  • Booth, R. (2026, May 11). Palantir’s access to identifiable NHS England patient data is ‘dangerous’, MPs say. The Guardian.
  • Devlin, H., & Burgis, T. (2026, March 14). Confidential health records from UK BioBank project exposed online. The Guardian.
  • Guardian Staff. (2026). UK Biobank health data listed for sale in China, government confirms. The Guardian.
  • Kolstoe, S. (2026, April 28). UK Biobank records listed for sale in China: why open data might be the answer. The Conversation.
 
Read in 5 mins, 10 mins, 20 mins

News in Ethics assessment of technology

Articles

CyberEthics Lab. Newsletter

CyberEthics Lab. Newsletter was born with the idea of sharing every month, and precisely on the first Monday, some in-depth articles or essays on sensitive or important topics for our society, from the point of view of ethics, technology, philosophy, human rights and social reflection. We want to offer our readers and anyone who is […]

Where the information went

Intellectual property comes in several kinds, and public argument about AI has settled on one of them. Copyright covers expression, which is where the fight over training data has taken place. Patents cover inventions, granting twenty years from filing in exchange for disclosing how the invention works. Trademarks cover the signs that tell a customer […]