JOURNAL ARTICLE

Semantic-based Privacy-preserving Record Linkage.

Lu Yang

Year: 2022 Journal:   International Journal for Population Data Science Vol: 7 (3)   Publisher: Swansea University

Abstract

IntroductionSharing aggregated electronic health records (EHRs) for integrated health care and public health studies is increasingly demanded. Patient privacy demands that anonymisation procedures are in place for data sharing. ObjectiveTraditional methods such as k-anonymity and its derivations are often overgeneralising resulting in lower data accuracy. To tackle this issue, we proposed the Semantic Linkage K-Anonymity (SLKA) approach to balance the privacy and utility preservation through detecting risky combinations hidden in the record linkage releases. ApproachK-anonymity processing quasi-identifiers of data may lead to ‘over generalisation’ when dealing with linkage data sets. As most linkage cases do not include all local patients and thus not all modifying data for privacy-preserving purposes needs to be used, we proposed the linkage k-anonymity (LKA) by which only obfuscated individuals in a released linkage set are required to be indistinguishable from at least k-1 other individuals in the local dataset. Considering the inference disclosure issue, we further designed the semantic-based linkage k-anonymity (SLKA) method through extending with a semantic-rule base for automatic detection of (and ruling out) risky associations from previous linked data releases. Specially, associations identified from the “previous releases” of the linkage dataset can become the input of semantic reasoning for the “next release”. ResultsThe approach is evaluated based on a linkage scenario where researchers apply to link data from an Australia-wide national type-1 diabetes platform with survey results from 25,000+ Victorians about their health and wellbeing. In comparing the information loss of three methods, we find that extra cost can be incurred in SLKA for dealing with risky individuals, e.g., 13.7% vs 5.9% (LKA, k=4) however it performs much better than k-anonymity, which can cause 24% information loss (k=4). Besides, the k values can affect the level of distortion in SLKA, such as 11.5% (k=2) vs 12.9% (k=3). ConclusionThe SLKA framework provides dynamic protection for repeated linkage releases while preserving data utility by avoiding unnecessary generalisation as typified by k-anonymity.

Keywords:
Linkage (software) Record linkage Computer science Linked data Anonymity Information retrieval Data mining Identifier Set (abstract data type) Unique identifier Data set Data science Computer security Artificial intelligence Semantic Web Medicine Population

Metrics

0
Cited By
0.00
FWCI (Field Weighted Citation Impact)
0
Refs
0.12
Citation Normalized Percentile
Is in top 1%
Is in top 10%

Topics

Data Quality and Management
Social Sciences →  Decision Sciences →  Management Science and Operations Research
Privacy-Preserving Technologies in Data
Physical Sciences →  Computer Science →  Artificial Intelligence
Ethics in Clinical Research
Health Sciences →  Medicine →  Public Health, Environmental and Occupational Health

Related Documents

JOURNAL ARTICLE

Privacy‐preserving record linkage

Vassilios S. VerykiosPeter Christen

Journal:   Wiley Interdisciplinary Reviews Data Mining and Knowledge Discovery Year: 2013 Vol: 3 (5)Pages: 321-332
BOOK-CHAPTER

Privacy-Preserving Record Linkage

Dinusha VatsalanDimitrios KarapiperisVassilios S. Verykios

Encyclopedia of Big Data Technologies Year: 2018 Pages: 1-8
BOOK-CHAPTER

Privacy-Preserving Record Linkage

Dinusha VatsalanDimitrios KarapiperisVassilios S. Verykios

Encyclopedia of Big Data Technologies Year: 2019 Pages: 1300-1307
BOOK-CHAPTER

Privacy-Preserving Record Linkage

Rob HallStephen E. Fienberg

Lecture notes in computer science Year: 2010 Pages: 269-283
BOOK-CHAPTER

Privacy-Preserving Record Linkage

Dinusha VatsalanDimitrios KarapiperisVassilios S. Verykios

Encyclopedia of Big Data Technologies Year: 2022 Pages: 1-10
© 2026 ScienceGate Book Chapters — All rights reserved.