Malin, Bradley A.; Yan, Chao; Bonomi, Luca. (2026). . Annual Review of Biomedical Data Science, 9(1), 47–67.
Health data are increasingly collected and shared across healthcare, research, and artificial intelligence (AI) applications, creating new opportunities for medical discovery and clinical decision-making. At the same time, these uses raise growing concerns about privacy, security, and trust. This review examines risks and protections throughout the health data life cycle, from their initial collection and use in healthcare to their later use in research and AI development. We discuss how regulations, organizational practices, and technology influence data protection, as well as emerging risks such as unintended disclosure of sensitive information through AI tools. We review methods for protecting health data, including controlling and monitoring access, assessing the risk that individuals could be re-identified, using statistical privacy techniques such as differential privacy, and generating synthetic data that resemble real data without directly representing individual patients. We also discuss approaches that allow organizations to collaborate without directly sharing sensitive data, including federated learning and secure cryptographic methods. These approaches involve important trade-offs between protecting privacy and preserving the usefulness of data. Overall, greater standardization, transparency, and practical guidance are needed to strengthen privacy and trust while supporting responsible use of health data in healthcare, research, and AI.

Figure 1 Simplified representation of different stages in the health data flow with potential risks and mitigation strategies. In the data-generation stage, health data are generated by individuals (e.g., patients). In the primary use stage (pink), data are collected for care by healthcare organizations or third-party systems. In the secondary use stage (yellow), data are shared and analyzed to accelerate knowledge discovery.