Skip to content
    August 23, 2026| Top Floor Team| 12 min read

    De-Identifying PHI: Safe Harbor vs Expert Determination

    There are two lawful ways to de-identify protected health information and no third one. 45 CFR 164.514(b) says a covered entity may determine that health information is not individually identifiable "only if" either a qualified person applies statistical and scientific principles, determines the re-identification risk is "very small," and documents the methods and results, or the eighteen listed categories of identifiers are removed and the entity has no actual knowledge that what remains could still identify someone. Get either right and 45 CFR 164.502(d)(2) takes the data out of the Privacy Rule entirely: "The requirements of this subpart do not apply to information that has been de-identified in accordance with the applicable requirements of § 164.514." The contrarian point is which method to reach for. Safe harbor is the cheap, self-service one, and for analytics it is usually the wrong one, because the price of the safe harbor is dates and geography, which is most of what makes health data analytically valuable.

    What follows is what each method actually demands, the middle path the regulation offers and almost nobody uses, and the four things de-identification does not buy you.

    Key takeaways

    • Only two methods exist. Hashing a medical record number, dropping names, or "anonymizing" by internal convention are none of them, and the resulting data is still PHI.
    • Safe harbor removes eighteen categories, including all date elements finer than year, ages over 89, and geography below the state level except a three-digit zip in populous areas. It is free, mechanical, and analytically expensive.
    • Expert determination has no named credential and no numeric risk threshold in the regulation. What it does have is a documentation requirement, and that documentation is the deliverable you are buying.
    • The limited data set is the underused middle: sixteen direct identifiers removed, dates and full geography retained, three permitted purposes, and a data use agreement. It is not de-identified data, but it is often the right answer.
    • De-identification under HIPAA does not clear you under state law, the FTC, your own contracts, or your customers' expectations about AI training. Those are four separate questions.

    Safe harbor: the eighteen, read carefully

    The list at 164.514(b)(2)(i) runs from (A) to (R), and it applies to identifiers of the individual "or of relatives, employers, or household members of the individual." Names; geographic subdivisions smaller than a state; all date elements except year for dates directly related to an individual; telephone numbers; fax numbers; email addresses; social security numbers; medical record numbers; health plan beneficiary numbers; account numbers; certificate and license numbers; vehicle identifiers and serial numbers including plates; device identifiers and serial numbers; URLs; IP addresses; biometric identifiers including finger and voice prints; full face photographic images and comparable images; and finally, the one people forget, "any other unique identifying number, characteristic, or code."

    Three of these deserve individual attention because they are where safe harbor projects actually fail.

    Dates. The rule removes "all elements of dates (except year) for dates directly related to an individual, including birth date, admission date, discharge date, date of death; and all ages over 89 and all elements of dates (including year) indicative of such age," with ages 90 and over permitted to be aggregated into one bucket. Year granularity destroys any longitudinal analysis: length of stay, time to readmission, treatment sequencing, seasonality. If your model depends on knowing that the second visit was eleven days after the first, safe harbor is not your method.

    Geography. Anything smaller than a state comes out, with one exception: the first three digits of a zip code may stay if, per current publicly available Census data, all zip codes sharing those three digits contain more than 20,000 people, and the three-digit prefixes for units of 20,000 or fewer must be changed to 000. Note that this is a live dependency on Census data, not a fixed list you can hardcode once.

    The catch-all at (R). "Any other unique identifying number, characteristic, or code" is what defeats the two most common shortcuts. A hashed MRN is a unique code. A pseudonymous patient ID that persists across records is a unique code. A free-text clinical note containing "the patient is the mayor of a town of 900 people" contains a unique characteristic. The catch-all is why safe harbor on structured columns is easy and safe harbor on clinical text is a research problem.

    There is a second condition that sits outside the list and is easy to skip: 164.514(b)(2)(ii) requires that the covered entity "does not have actual knowledge that the information could be used alone or in combination with other information to identify an individual." Removing all eighteen and knowing perfectly well that the remaining combination fingerprints someone is not compliance.

    Expert determination: what you are actually buying

    The regulation describes the expert as "a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods for rendering information not individually identifiable." That is the whole qualification standard. No certification is named, no licensing body is designated, no degree is specified. Anyone telling you a particular credential is required by HIPAA is describing a market convention, not the rule.

    The standard the expert has to reach is that "the risk is very small that the information could be used, alone or in combination with other reasonably available information, by an anticipated recipient to identify an individual." Two phrases carry weight. "Reasonably available information" means the analysis has to consider what an adversary could join your data against, not just what is in your file. "Anticipated recipient" means the determination is scoped to a recipient: a determination for a closed research consortium does not automatically hold for a public release.

    The regulation attaches no number to "very small." No percentage, no k value, no re-identification probability threshold appears in the text. That is a feature for a competent expert and a hazard for a buyer, because it means the rigor of the analysis is not visible from the outside.

    Which is why the deliverable is the second requirement, not the first. 164.514(b)(1)(ii) requires that the expert "documents the methods and results of the analysis that justify such determination." The documentation is what you hand a customer's security reviewer, what you produce if OCR asks, and what makes the determination something other than an assertion. When you commission an expert determination, buy the report, and check that it names the recipient, the data set version, the assumptions about available auxiliary data, and the date. A determination without a scope statement ages into a liability the moment your data set changes.

    The limited data set, which is the answer more often than people think

    Sitting between the two methods is a construct most teams have never used. Under 164.514(e), a limited data set excludes sixteen direct identifiers: names, postal address information other than town or city, state, and zip; telephone and fax numbers; email addresses; social security numbers; medical record numbers; health plan beneficiary numbers; account numbers; certificate and license numbers; vehicle identifiers; device identifiers; URLs; IP addresses; biometric identifiers; and full face images.

    Compare that list with the eighteen. Dates survive. Town, city, state, and full zip survive. The catch-all does not appear. For most analytic work, that difference is the entire ballgame.

    The price is that a limited data set is still protected health information, so three conditions attach. It may be used or disclosed only for research, public health, or health care operations. It requires a data use agreement that establishes permitted uses, names who may use or receive it, and binds the recipient. And because it is still PHI, the Security Rule and the Breach Notification Rule still apply to it.

    For a health tech company doing internal analytics, or sharing with a research partner, or standing up a customer-facing benchmarking product, the limited data set plus a data use agreement is frequently cheaper, faster, and analytically stronger than either de-identification route. It is worth pricing before you commission an expert.

    What de-identification does not buy you

    It does not exit other regimes. Part 2 records cross-reference the HIPAA standard rather than replacing it: 42 CFR 2.16 requires policies for "rendering patient identifying information de-identified in accordance with the requirements of 45 CFR 164.514(b)," so the same two methods govern. State health privacy statutes, general consumer privacy laws, and non-US regimes each define identifiability their own way, and several are stricter than the safe harbor.

    It does not exit the FTC's rules for non-HIPAA products. If your product is consumer-facing rather than sold to covered entities, the relevant definition may be the FTC's, where 16 CFR 318.2 turns on whether there is "a reasonable basis to believe that the information can be used to identify the individual." That is a different sentence from the safe harbor list, and satisfying one does not mechanically satisfy the other. Whether HIPAA applies to your app at all is the prior question.

    It does not override your contracts. Business associate agreements and enterprise customer agreements routinely restrict use of derived and de-identified data more tightly than HIPAA does, and some prohibit model training on it outright. The regulation permits a covered entity to use PHI to create de-identified information, and to disclose PHI to a business associate for that purpose. Your agreement may not.

    It does not settle the AI question. Training on properly de-identified data is lawful under the Privacy Rule because the data is no longer PHI. Whether it is contractually permitted, whether your customers were told, and whether the trained model can be induced to emit something identifying are three separate questions, and only the first one is answered by 164.514. If your training corpus is clinical free text, notice that the safe harbor catch-all applies to characteristics, not just numbers, and that free text is exactly where unique characteristics hide.

    Re-identification codes, done lawfully

    One provision is worth knowing because it prevents an unnecessary architectural compromise. 164.514(c) lets a covered entity assign a code to de-identified records so it can re-identify them later, on two conditions: the code "is not derived from or related to information about the individual and is not otherwise capable of being translated so as to identify the individual," and the entity "does not use or disclose the code or other means of record identification for any other purpose, and does not disclose the mechanism for re-identification."

    The first condition is why a hash of the MRN does not qualify. It is derived from information about the individual. A random surrogate key held in a mapping table you control does qualify. And per 164.502(d)(2)(i), disclosing the code itself is a disclosure of PHI, so the mapping table lives on the PHI side of your boundary, not with the de-identified extract.

    The honest caveat

    If your use case survives on year-level dates and state-level geography, do safe harbor yourself. It is a list of eighteen categories, your engineers can implement it against structured data in a sprint, and the only outside help worth buying is a second pair of eyes on the free-text fields and the catch-all. Paying a consultancy to run a checklist you can read in ten minutes is not a good trade, and we will say so on the call.

    Expert determination is genuinely specialist work, and we will be straight about where we sit. We do the scoping, the data-flow analysis, the contractual review, and the documentation and evidence layer around it. The statistical determination itself is signed by someone whose profession that is. Any firm that offers to both design your de-identification pipeline and certify that its output is de-identified is holding two roles that should not sit in one place; ask them how they separate the work, and treat a vague answer as an answer.

    Where Top Floor fits

    We map where PHI actually flows, decide with you between safe harbor, expert determination, and a limited data set, and build the evidence trail under our HIPAA practice. Ongoing controls, vendor terms, and the customer-facing explanation of all of it sit under Compliance as a Service, and independent assessment and reporting under audit and assurance.

    How to decide this week

    Write down the two analytic questions your data set exists to answer. If either depends on intervals shorter than a year or on geography finer than a state, safe harbor is off the table and you are choosing between expert determination and a limited data set. That single check resolves most of these decisions in an afternoon.

    Then grep your pipeline for persistent identifiers. Any column that is stable across records for the same person, including hashes, is a unique code under the catch-all, and its presence means your "de-identified" extract is still PHI.

    Finally, read the derived-data clause in your three largest customer agreements. HIPAA is not the binding constraint in most of these conversations; the contract is, and it is the one people discover last.

    Frequently asked questions

    Is removing names and social security numbers enough to de-identify PHI?

    No. HIPAA recognizes only two methods, and neither is a partial removal. Safe harbor under 45 CFR 164.514(b)(2) requires removing eighteen categories of identifiers, including all date elements finer than year, ages over 89, geography below state level with a narrow three-digit zip exception, device and vehicle identifiers, URLs, IP addresses, biometrics, full face images, and any other unique identifying number, characteristic, or code. It also requires that you have no actual knowledge the remainder could still identify someone. The alternative is expert determination under 164.514(b)(1). Anything else leaves the data as protected health information.

    Does HIPAA require a certified expert for expert determination?

    No credential is named anywhere in the regulation. 45 CFR 164.514(b)(1) describes "a person with appropriate knowledge of and experience with generally accepted statistical and scientific principles and methods for rendering information not individually identifiable," and it sets no licensing, certification, or degree requirement. What the rule does require is substantive: the expert must determine that the re-identification risk is very small for an anticipated recipient given other reasonably available information, and must document the methods and results justifying that determination. Buy the documentation, and check that it names the recipient, the data set, and the date.

    What is the difference between a limited data set and de-identified data?

    A limited data set removes sixteen direct identifiers but keeps dates and full geography, which makes it far more useful analytically. The trade is that it remains protected health information: it may only be used or disclosed for research, public health, or health care operations, it requires a data use agreement with the recipient, and the Security Rule and Breach Notification Rule still apply to it. De-identified data under 164.514(a) and (b) is outside the Privacy Rule entirely. For internal analytics and research partnerships, the limited data set is often the better answer and is routinely overlooked.

    Can we train an AI model on de-identified patient data?

    Under the Privacy Rule, properly de-identified data is not protected health information, so the Privacy Rule does not restrict its use, including for model training. That answers only one of four questions. Your business associate agreements and customer contracts frequently restrict derived and de-identified data more tightly than HIPAA does. State and non-US privacy law may define identifiability differently. And if your corpus is clinical free text, the safe harbor catch-all covers unique characteristics as well as numbers, which is exactly where free text fails. Answer the contract question before the regulatory one; it is the one that usually bites.

    Share Share on LinkedIn

    Need help with your compliance program?

    Our team of senior practitioners can help you navigate complex compliance requirements and build a security program that holds up under scrutiny.

    Schedule a Free Consultation

    Get insights like this in your inbox

    Practical compliance and security guidance for teams preparing for their next audit. No spam, unsubscribe anytime.

    Ask to be added to our mailing list for practical compliance and security guidance. We add you by hand, we confirm before sending anything, and we never share your address.