Skip to content

Most identifiers have a shelf life: a cookie goes stale (or expires); a device ID resets when a user resets their advertising ID in their phone’s settings; a partner ID stops working when the partnership ends. Email addresses, by contrast, endure: people change jobs, phones, and browsers far more frequently than they update their primary email address.

That durability explains why email has become the foundation of identity resolution in advertising. Email isn’t a wildly high-tech workaround for cookies or anything, but more importantly, it is stable, and most people will voluntarily provide it in exchange for a discount or other perk. You don’t collect it passively in the background, and therein lies the distinction: the value of something given versus collected.

What Is an Identifier?

In short, an identifier is any value used to recognize a device, browser, app instance, or person across events.

A caveat: Not all identifiers are the same thing, but we talk about them interchangeably. Technically, hashed emails, mobile advertising identifiers, cookie identifiers, and partner identifiers are all types of identifiers. Each lets a system recognize the same user multiple times. However, they differ in what they’re derived from and how long they last.

  • Browsers assign cookie identifiers that expire when the user resets or clears the browser’s storage.
  • A device’s operating system assigns mobile advertising identifiers, which users can reset whenever they want.
  • Hashed emails are generated from an email address a person provides, and they last as long as the email address is in use.

That hierarchy determines an identifier’s use and value. Short-lived identifiers, such as cookies, work well for in-session frequency capping. Persistent identifiers, such as mobile advertising identifiers, make cross-device and cross-platform matching possible. One identifier currently underpins most identity resolution work: the hashed email.

What Is a Hashed Email?

A hashed email is an email address that’s been converted through a one-way cryptographic function into a fixed-length string. This conversion matches a person across systems without transmitting the actual address. Two platforms holding the same email address and running it through the same hashing process will independently arrive at an identical hash. Now, you have a match without either party sharing raw personal contact information.

Obtaining that identical hash depends on normalization. In practice, this normalization process is where hashed-email systems most often break. Hashing requires a standardized email address converted to lowercase and trimmed of stray whitespace. If you skip that step, John.Doe@Company.com and john.doe@company.com — the same person’s email captured by two different systems with two different formatting habits — produce two completely different hashes.

The problem? The systems have no way to know they’re looking at the same person, so the match fails. That’s why you need normalization for hashing to work as a matching mechanism.

Most often, the ad tech industry uses the SHA-256 hashing function for this purpose, and although older implementations using MD5 or SHA-1 still show up in the wild on occasion, SHA-256 is the modern industry standard because it provides a significantly more secure, collision-resistant one-way guarantee. All three take an input and produce a fixed-length output, which you can’t revert back to the original address.

How Hashing Works, and What It Doesn’t Do

One-way By Design

A cryptographic hash function is deterministic and one-directional. The same input — that is, the same normalized email — always produces the same output, on any system running the same algorithm. But, the reverse doesn’t hold: given only the output, there isn’t a computation that recovers the original input. That asymmetry is the point of using a hash and not the raw email. It lets two parties confirm they’re looking at the same person without either seeing what data the other party has.

Why Normalization Comes First

What if two publishers both know the same reader by email? Publisher A stores the address as entered at signup: “First.Last@Gmail.com.” Publisher B stores it as the reader entered it later during a purchase: “first.last@gmail.com ” (with a trailing space). If each hashes the address exactly as stored, the outputs won’t match because the inputs aren’t identical — not because the hashing failed.

Normalizing both addresses to a consistent lowercase, whitespace-trimmed format before hashing resolves this. This step is the most common source of match-rate problems in production identity systems, more so than any flaw in the hashing algorithm itself.

Why Hashing is Not Anonymization

A hash is a proxy for one person. You need that consistency for effective matching; this is why it qualifies as personal, not anonymous, data. If a system can tie a hash back to an individual — and the process relies on that capability — the hash is pseudonymized, not anonymized.

Under GDPR’s Article 4(5), pseudonymized data remains personal data because it can still be attributed to a specific person, even if indirectly. Another important caveat is that a hash of a known, common email address can be looked up against precomputed hash tables. A hash isn’t secret in the same way a password was designed to be.

What Does Hashing Protect Against?

Hashing does offer benefits, e.g., it eliminates the need for raw email addresses to travel between parties to enable matching. A publisher and advertiser can confirm they’re describing the same customer without either party transmitting an address through cyberspace, which reduces exposure.

What Is an External ID?

An external ID is an identifier assigned by a party outside the platform (such as an advertiser’s or partner’s internal customer identifier), then passed inside the platform so the party’s and platform’s records can be matched.

For example: An advertiser’s CRM assigns every customer a unique internal ID when they enter the system. When that advertiser shares data with an ad platform, they can pass the CRM ID with a hashed email so that both systems can agree, going forward, that they’re describing the same customer. The advertiser’s record and the platform’s record become linked.

A platform-assigned ID originates inside the platform, and it’s meaningless outside the platform. An external ID originates with the advertiser or partner and is passed in. The platform hasn’t created it, but agrees to recognize it.

What Is a MAID?

A mobile operating system assigns device-level mobile advertising identifiers (MAIDs) for advertising purposes. The two current implementations in use are Apple’s IDFA and Google’s GAID. Unlike hashed emails, MAIDs don’t derive from something a person voluntarily hands over. The device itself generates the MAID — and, more importantly, it’s resettable. A user can reset their MAID in their device settings whenever. After reset, the device presents as a new, unconnected identity to anything that previously tracked it by that value.

A note: Consent frameworks introduced in recent years, like Apple’s App Tracking Transparency prompt, mean that a good chunk of app traffic doesn’t have a usable MAID because:

  1. The user declined tracking or;
  2. The app never requested it.

Published figures on opt-in rates vary by source, app category, and date, and a number that was accurate a year, or even six months ago, is likely out of date. What is stable is the mechanism. MAIDs are conditional on consent in a way that hashed emails, which rely on a direct exchange with a person, typically are not.

Identifier Types Compared

Durability matters most across the various identifier types — that is, how long a given value keeps working before it resets, expires, or becomes unavailable. Everything else about an identifier’s usefulness follows from that durability.

Identifier type What it’s derived from How long it persists Where it works Main limitation
Hashed email A normalized email address run through a one-way hash function. As long as the person keeps using that email address. Works across web, app, offline environments wherever an email was collected. Only covers people who gave you an address; one person often has several.
MAID Assigned by the device’s mobile OS. Resettable by the user at any time from device settings. Works within mobile app environments on the device to which it was assigned. Absent or blocked for a large and growing share of app traffic thanks to consent requirements.
Third-party cookie ID Assigned by a domain other than the one the user is visiting. Cleared whenever the browser’s storage is cleared; increasingly blocked by default. Works only in browsers that still permit third-party cookies. Being phased out across major browsers; unavailable in many environments already.
External/partner ID Assigned by an advertiser or partner outside the platform. Persists as long as the external party maintains that record. Works wherever the external party shares it and a platform agrees to map it. Only useful once matched to a platform-native or shared identifier.
Authenticated first-party ID Assigned by a publisher or platform when a user logs in. Persists as long as the account remains active and the user stays logged in. Works only in logged-in, first-party environments. Limited to sessions where the user authenticates.

Why Email-Based Identity Supports Owned Data

An email address comes from a direct relationship with a person. They offer it willingly at signup, checkout, or when subscribing to something in exchange for a discount, service, or content they wanted. You can point to that exchange and the consent attached to it.

That asset category is different from a licensed audience segment, where a data provider has aggregated behavioral signals from people with whom you don’t have a direct relationship. In that case, you’re buying access to a version of these users that’s been shaped by someone else’s methodology.

Licensed segments can be useful, but they’re rented. You’re paying for reach into a group you otherwise wouldn’t have a connection to, and as soon as you stop paying for it, that reach disappears. A customer-provided email address works differently; it’s yours to connect to other systems and build on for as long as your relationship with that person continues.

Email-based identity offers a practical bridge between the data a business owns and the reach it needs beyond its properties. What’s specific to owned data is the relationship under the address — the fact that the person chose to share their email for a reason you can name.

The Shortcomings of Email-Based Identity

Email-based identity has real benefits, but it has limits, too:

Coverage Is Bounded by Collection

Email-based identity only covers those who provide an address. If the bulk of your buyers browse anonymously and don’t sign up, log in, or check out with an email, they stay outside your system’s reach; any hashing/matching sophistication won’t change that fact.

One Person Often Has Multiple Email Addresses

Email-based identity was meant to solve fragmentation, but if someone uses a personal email for shopping, work email for professional subscriptions, and a third for anything else, you’re dealing with a smaller version of the same problem. Each address hashes to a different value, so unless something else links them, that person can appear in your systems as three unconnected identities.

Shared Devices and Shared Accounts Create False Matches

A household email used for a shared streaming account — or a shared device logged into a family member’s inbox — can result in a platform treating several different people as one. The result? Distorted targeting and measurement.

You can’t automatically use an address provided to you for order confirmations for advertising matching, too. Treating it the same way creates a trust problem with the customer and, depending on jurisdiction, a compliance one.

How Hashed Emails Relate to Cookieless Identifiers

Several identity solutions — built to operate without third-party cookies (e.g., industry frameworks that clean rooms and DSPs used for matching) — are derived from, or resolve back to, email-based identity at some point in the pipeline. These newer frameworks don’t replace the underlying hashed-email mechanism because they’re a standardized layer built on top of it. They add governance, encryption, or refresh logic you can’t get from a raw hash.

Because most identifiers built to succeed cookies trace to the same source, you should improve your hashed-email hygiene (and see the benefits across multiple downstream systems) with:

  • Accurate normalization.
  • Clean collection.
  • Clear consent.

Key Takeaways

“Identifier” is the umbrella term under which hashed emails, MAIDs, cookie IDs, and external IDs fall. Each has different durability and different origins. Hashing enables matching without moving raw addresses between systems; it doesn’t anonymize the resulting data, which remains personal data because it’s still attributable to a specific person.

Email-based identity supports owned data because it originates in a direct relationship with a person (versus because the matching technique is exclusive to one company). The main constraint when using email-based identity is coverage: the number of customers who provide their email and how many separate addresses each uses.

Frequently Asked Questions (FAQs)

Is hashed email personal data?

Yes. A hashed email address can be directly or indirectly attributed to a specific person. As a result, GDPR and other frameworks classify it as personal data. Consult your legal counsel for how it applies to your specific use case.

Can you reverse a hashed email?

No, not through computation, i.e., you can’t mathematically invert a properly implemented hash back into the original address. But, you can check it against a known list; if someone has a candidate email and hashes it the same way, you can compare the result to see if it matches.

What’s the difference between a hashed email and a UID2?

You feed a raw, hashed email (like SHA-256) into the UID2 infrastructure. UID2 processes that hash to generate a dynamic, encrypted payload: the UID2 token. UID2 adds rotation, consumer opt-out compliance, and governance.

Consent requirements for using hashed emails in advertising depend on your jurisdiction, how you originally collected the address, and for what purpose. Talk to your legal or compliance counsel about what applies to your specific collection flow.

Why do my match rates look low?

These factors can cause low match rates:

  • Differences in trimming, lowercasing, or character formatting before hashing, which produce different SHA-256 outputs.
  • Comparing offline customer records from your CRM against active online/logged-in users creates a coverage gap.
  • Using different hashing algorithms (e.g., SHA-256 versus MD5) or applying custom salt values (which prevent cross-platform matching unless both parties share the exact same salt key) will break deterministic matches. Additionally, failing to strip email sub-addresses (such as Gmail ‘+’ aliases) during the normalization step prevents matching against the user’s primary registered hash.

Written by

Jodi Ireland

16 articles

You may also like

Ready to Realize Your Brand’s Potential?

Start Campaign