Anonymization starts with the question your retained prospect data must answer. Remove every route back to a person that the output no longer needs. A clean-looking record can still fail: if someone could re-identify an individual through reasonably available means, the data is pseudonymized rather than effectively anonymized.1 Build the reporting copy in layers, test the combinations that remain, and keep any reconnectable version under a separate control path.
Define the retained copy
Start by writing down what the retained record must support. This gives you a reason to keep each field and a basis for removing the rest.
If you do not need to identify individuals, anonymize the data so identification is no longer possible.2 Anonymization permanently removes or alters personally identifiable information so individuals cannot be identified directly or indirectly.3 The result should be non-identifiable and irreversible.4
Write the reporting question in plain language, then review every field against it. Keep a field when it helps answer that question. Remove, generalize, or aggregate a field when it adds detail without improving the report. For analytics, anonymize or aggregate data wherever possible.5
Remove direct identifiers
Make the first editing pass mechanical. Search for values that identify one person on their own, remove them from the retained copy, and record what you removed in your process notes.
Find and redact direct identifiers first.6 A person's name, address, or telephone numbers that specifically identify them are direct identifiers.7 Redact these identifiers from the data.8
Do not preserve a direct identifier in a comment, free-text field, exported note, or attachment while deleting it from the main record. Apply the same treatment across every field you plan to keep. Move on when a reviewer can inspect the copy without finding a direct identifier that the report does not need.
Inspect indirect identifiers
The second pass requires judgment. A field can look harmless on its own and identify someone when combined with other details available to the people who will use the report.
The second step is to find, mark, and consider indirect identifiers in the data.9 Indirect identifiers can reveal an individual when combined with other information, including occupation, salary, age, and location.10 Keep only the indirect identifiers that are essential for understanding the data and leave out the rest.11
Review combinations as well as individual columns. Ask whether a person could be singled out by joining a location with an age or an occupation with a salary. One unusual detail can also identify someone when combined with another retained field. When that risk appears, generalize or remove the detail, or change the report so the combination is no longer needed.
Choose the right output
Decide whether the work needs an anonymized dataset, an aggregate report, or a reconnectable working copy. This determines how you handle codes and lookup files.
Aggregation can protect confidentiality before data is shared, but it limits utility.12 Use an aggregate output when the report can answer its question without preserving a row for each prospect. Keep the grouping broad enough to avoid recreating an individual from a small or unusual combination of values.
Pseudonymization replaces identifiers with coded values and keeps a separate key that can reconnect them to the individual.13 If the team needs to reconnect records for an operational reason, keep that version in a separate workflow and do not release it as anonymized data.
Keep the substitution list in a separate, secure file when names or other labels have been changed. The list should record each changed person, place, organization, company, or product name and its substitute.14 Access controls help prevent sensitive attributes from being linked to specific participants and locations.15
Test before retaining or sharing
Treat the final review as an attempt to break the anonymization. Give the copy to someone who understands its intended use and ask what could be inferred from the fields that remain.
Check whether a person could realistically be re-identified directly or indirectly. If that route exists, the data is not truly anonymized and may still count as personal information.16 Test the combinations available to the people who will receive the report, including information they could reasonably obtain outside the retained copy.
Pay extra attention to fields whose disclosure could cause harm, legal issues, or damage to a person's reputation. Sensitive data includes information with those risks whether or not it is legally protected.17 Remove details that the report does not need, and tighten access to any working copy that still permits linkage.
If the retained output does not pass this review, keep it under the appropriate controlled workflow or delete what you no longer need.
What not to do
Keep these failure modes beside the workflow while someone prepares the retained copy.
- Do not leave unneeded personal data identifiable after deciding to retain the record. Erasing or anonymizing it supports data minimization and accuracy and reduces the risk of using it in error.18
- Do not treat anonymization as a label applied after editing. Use it as an alternative to deleting data only when it has been done properly.19
- Do not approve a release because direct identifiers are gone. A realistic route to direct or indirect re-identification means the data may still be personal information.16
- Do not use sensitive fields simply because they are available. Disclosure of identifiable, sensitive, or legally protected data can harm participants.20