What Is PII? A
Term With No Single Definition
The Garlick Group | July 2026
Ask a
compliance officer to define personally identifiable information and the honest
answer starts with a question of their own: under which law? No single federal
statute defines PII for the United States as a whole, because no single federal
privacy statute governs the United States as a whole. What exists instead is a
patchwork built sector by sector and, increasingly, state by state, and the
term shifts meaning depending on which patch a given piece of data happens to
fall into.
The most
commonly cited baseline comes from the National Institute of Standards and
Technology, whose Special Publication 800-122 defines PII as any information
about an individual maintained by an agency, including anything that can be
used to distinguish or trace a person's identity, such as name, Social Security
number, or biometric records, whether alone or combined with other linkable
information. 1 That definition draws a distinction worth sitting
with: some data identifies a person on its own, and some data identifies a
person only in combination. A Social Security number does the former. A birth
date does the latter, since a birth date attached to nothing else describes
millions of people and a birth date attached to a ZIP code and a name describes
one.
Direct
identifiers are the easy category: full name, passport number, driver's license
number, biometric data, financial account numbers. Indirect identifiers are
harder to reason about, because their identifying power depends entirely on
context and combination. A first name alone identifies almost no one. A first
name, an employer, and a job title, posted together on a public forum, can
identify someone within minutes. Latanya Sweeney's now-famous finding that ZIP
code, birth date, and sex uniquely identify the overwhelming majority of the
U.S. population is the standard illustration of why regulators stopped treating
“indirectly identifying” as a synonym for “low risk.”
Outside the
United States, and increasingly inside it, the operative term is “personal
data” rather than PII, and the difference is not merely stylistic. The GDPR
defines personal data as any information relating to an identified or
identifiable natural person, a formulation broader than most American PII
definitions because it captures opinions and inferences about a person, not
only facts that name or trace them. 2 Nineteen U.S. states now have
comprehensive privacy statutes on the books, and their definitions of “personal
data” largely track the GDPR's structure: information linked or reasonably
linkable to an identified or identifiable individual, rather than a fixed list
of data types. 3
The California
Consumer Privacy Act illustrates how far the newer definitions have moved from
the old checklist model. Personal information under the CCPA includes anything
that identifies, relates to, describes, or could reasonably be linked, directly
or indirectly, with a particular consumer or household, a definition that
reaches purchase history, browsing behavior, and inferences drawn from other
data, alongside the traditional identifiers.
A second axis
matters as much as the identify/non-identify distinction: sensitivity. Most
frameworks now separate ordinary PII from sensitive PII, a narrower category
that triggers stricter handling because exposure carries greater potential for
harm. Social Security numbers, financial account credentials, precise
geolocation, health information, biometric data, and characteristics such as
race, religion, or sexual orientation typically fall into this tier, and most
state privacy laws require opt-in consent before that category can be processed
at all, rather than the opt-out standard that governs ordinary personal data. 4
Protected health information sits inside this sensitive tier as its own
regulated subset, defined not by the content of the data alone but by who holds
it: health information becomes PHI specifically when a HIPAA-covered entity or
its business associate collects it in connection with treatment or payment. The
same lab result is PHI in a hospital's records and ordinary sensitive data if a
person posts it to a health forum themselves. 5
This is the
detail that trips up organizations building a compliance program from a single
template: the same data element can carry different labels, different
obligations, and different remedies depending on who collected it, why, and
under which state's law the individual resides. A business operating in
California, Colorado, and Virginia is not managing one PII policy. It is
managing three overlapping ones, each with its own thresholds, its own
exemptions, and its own enforcement mechanism, and the count keeps growing as
more states pass comprehensive statutes.
Rhode Island's
new privacy law, effective January 1, 2026, is a useful reminder that the trend
line still runs toward more state activity rather than less, and that even
newly enacted statutes do not converge on a shared vocabulary; Rhode Island's
statute conspicuously omits any defined term for “personally identifiable
information” at all, relying instead on “personal data.” 6 Termly's
2026 guide goes further and argues that PII itself, as a term, is being phased
out of the newer legislative vocabulary in favor of “personal data” or
“personal information,” even though the older term persists in federal usage,
in security literature, and in ordinary conversation. 7
None of this
converges toward a tidy answer, and that absence of convergence is itself the
answer worth taking away. Treating PII as a fixed checklist, of the kind that
circulated in privacy policies a decade ago, understates what current law
actually covers. The safer operating assumption is functional rather than
definitional: any data point that could, alone or in combination, trace back to
a specific person deserves handling as if it were regulated, because under some
applicable law, in some jurisdiction where a business has customers, it very
likely is.
www.garlickgroup.com
No comments:
Post a Comment