Outbound Wiki

Prospect data anonymization

Replacing identifying prospect details with non-identifying values when records must be retained for reporting or analysis.

Anonymization starts with the question your retained prospect data must answer. Remove every route back to a person that the output no longer needs. A clean-looking record can still fail: if someone could re-identify an individual through reasonably available means, the data is pseudonymized rather than effectively anonymized.1 Build the reporting copy in layers, test the combinations that remain, and keep any reconnectable version under a separate control path.

Define the retained copy

Start by writing down what the retained record must support. This gives you a reason to keep each field and a basis for removing the rest.

If you do not need to identify individuals, anonymize the data so identification is no longer possible.2 Anonymization permanently removes or alters personally identifiable information so individuals cannot be identified directly or indirectly.3 The result should be non-identifiable and irreversible.4

Write the reporting question in plain language, then review every field against it. Keep a field when it helps answer that question. Remove, generalize, or aggregate a field when it adds detail without improving the report. For analytics, anonymize or aggregate data wherever possible.5

Remove direct identifiers

Make the first editing pass mechanical. Search for values that identify one person on their own, remove them from the retained copy, and record what you removed in your process notes.

Find and redact direct identifiers first.6 A person's name, address, or telephone numbers that specifically identify them are direct identifiers.7 Redact these identifiers from the data.8

Do not preserve a direct identifier in a comment, free-text field, exported note, or attachment while deleting it from the main record. Apply the same treatment across every field you plan to keep. Move on when a reviewer can inspect the copy without finding a direct identifier that the report does not need.

Inspect indirect identifiers

The second pass requires judgment. A field can look harmless on its own and identify someone when combined with other details available to the people who will use the report.

The second step is to find, mark, and consider indirect identifiers in the data.9 Indirect identifiers can reveal an individual when combined with other information, including occupation, salary, age, and location.10 Keep only the indirect identifiers that are essential for understanding the data and leave out the rest.11

Review combinations as well as individual columns. Ask whether a person could be singled out by joining a location with an age or an occupation with a salary. One unusual detail can also identify someone when combined with another retained field. When that risk appears, generalize or remove the detail, or change the report so the combination is no longer needed.

Choose the right output

Decide whether the work needs an anonymized dataset, an aggregate report, or a reconnectable working copy. This determines how you handle codes and lookup files.

Aggregation can protect confidentiality before data is shared, but it limits utility.12 Use an aggregate output when the report can answer its question without preserving a row for each prospect. Keep the grouping broad enough to avoid recreating an individual from a small or unusual combination of values.

Pseudonymization replaces identifiers with coded values and keeps a separate key that can reconnect them to the individual.13 If the team needs to reconnect records for an operational reason, keep that version in a separate workflow and do not release it as anonymized data.

Keep the substitution list in a separate, secure file when names or other labels have been changed. The list should record each changed person, place, organization, company, or product name and its substitute.14 Access controls help prevent sensitive attributes from being linked to specific participants and locations.15

Test before retaining or sharing

Treat the final review as an attempt to break the anonymization. Give the copy to someone who understands its intended use and ask what could be inferred from the fields that remain.

Check whether a person could realistically be re-identified directly or indirectly. If that route exists, the data is not truly anonymized and may still count as personal information.16 Test the combinations available to the people who will receive the report, including information they could reasonably obtain outside the retained copy.

Pay extra attention to fields whose disclosure could cause harm, legal issues, or damage to a person's reputation. Sensitive data includes information with those risks whether or not it is legally protected.17 Remove details that the report does not need, and tighten access to any working copy that still permits linkage.

If the retained output does not pass this review, keep it under the appropriate controlled workflow or delete what you no longer need.

What not to do

Keep these failure modes beside the workflow while someone prepares the retained copy.

  • Do not leave unneeded personal data identifiable after deciding to retain the record. Erasing or anonymizing it supports data minimization and accuracy and reduces the risk of using it in error.18
  • Do not treat anonymization as a label applied after editing. Use it as an alternative to deleting data only when it has been done properly.19
  • Do not approve a release because direct identifiers are gone. A realistic route to direct or indirect re-identification means the data may still be personal information.16
  • Do not use sensitive fields simply because they are available. Disclosure of identifiable, sensitive, or legally protected data can harm participants.20

Sources

  1. 1
    “if you could at any point use any reasonably available means to re-identify the individuals to which the data refers, that data will not have been effectively anonymised but will have merely been pseudonymised.”
  2. 2
    “If you do not need to identify individuals, you should anonymise the data so that identification is no longer possible.”
  3. 3
    “Anonymization is the process of permanently removing or altering personally identifiable information (PII) from a dataset so that individuals cannot be identified, either directly or indirectly.”
  4. 4
    “Once anonymized, the data should be non-identifiable and irreversible.”
  5. 5
    “Analytics data should be anonymised or aggregated wherever possible.”
  6. 6
    “Step 1: Find and redact direct identifiers in your data”
  7. 7
    “Direct identifiers are ones like the participant's name, address, or telephone numbers that specifically identify them.”
  8. 8
    “These should always be redacted from the data.”
  9. 9
    “Step 2: Find, highlight, and consider the indirect identifiers in your data”
  10. 10
    “Indirect identifiers are ones that, if placed with other information, could reveal an individual (e.g., by cross-referencing occupation, salary, age, and location).”
  11. 11
    “Consider which indirect identifiers are essential for understanding the data and which ones to leave out.”
  12. 12
    “To protect this confidentiality, data are often aggregated to a geographic level such as census tracts or transportation analysis zones (TAZs) before being publicly shared.”
  13. 13
    “This differs from pseudonymization, where identifiers are replaced with coded values but can still be reconnected to the individual using a separate key.”
  14. 14
    “Create a listing, in a separate file, kept somewhere safe, of all the names – people, places, organizations, companies, products – that you have changed and what you have substituted for them.”
    Anonymizing Qualitative Data

    researchmethodscommunity.sagepub.comBack to the text

  15. 15
    “Implementing de-identification strategies and adopting access control measures are vital for preventing these risks, as they safeguard against the linkage of sensitive attributes with specific participants and locations.”
  16. 16
    “If there's any realistic way to re-identify a person (directly or indirectly), then the data isn't truly anonymized and may still count as personal information.”
  17. 17
    “Sensitive data encompasses any information that, if disclosed, could lead to harm, legal issues, or damage to a subject's reputation, whether or not it is legally protected.”
  18. 18
    “Apart from helping you to comply with the data minimisation and accuracy principles, this also reduces the risk that you will use such data in error – to the detriment of all concerned.”
  19. 19
    “Anonymization is a valid alternative, but only if done properly.”
  20. 20
    “The disclosure of identifiable, sensitive, or legally protected data can harm participants.”