Table of Contents
- What Is an Identity Graph?
- A Graph Is What a Platform Has; Resolution Is What It Does
- Why Two Platforms See the Same Buyer Differently
- What’s Inside a Graph, and Why It Explains the Difference
- How a Graph Accumulates
- What Is a Data Spine?
- Why a Bigger Graph Isn’t Automatically a Better One
- How to Evaluate a Platform’s Graph
- Key Takeaways
- Frequently Asked Questions (FAQs)
Imagine that a buyer interacts with two advertising platforms over the course of a week. They research a product on their phone, continue on their laptop, make a purchase while logged in, and return later from another device.
Platform A recognizes those interactions as belonging to the same person. Platform B sees three separate users.
Neither platform is necessarily broken; the difference comes down to what each platform has been able to observe. An identity graph can only contain relationships its owner has had an opportunity to see, and that opportunity can vary significantly between platforms.
As such, two platforms using similar identity resolution techniques can produce different views of the same audience, and one may recognize connections the other simply hasn’t observed. For performance marketers, those differences can affect reported reach, frequency, targeting, suppression, and measurement, even when campaign settings look nearly identical.
To understand identity graphs, you need to look beyond how identity resolution works and ask an important question: What has the platform actually observed?
What Is an Identity Graph?
An identity graph is a database that maps multiple identifiers to a single real person, enabling consistent recognition across devices and sessions.
You can think of the graph as a map showing how different identifiers relate to one another. In graph terminology, identifiers are nodes. These might represent browsers, devices, hashed emails (HEMs), mobile advertising IDs (MAIDs), or other identifiers available to the platform.
The relationships between those nodes are called edges. An edge represents the platform’s belief that two identifiers belong to the same person. The graph can also retain information about how confident the platform is about that relationship and how recently it was established or confirmed. An identity graph is also sometimes called a user graph.
Importantly, a graph isn’t simply a list of identifiers; its value comes from the relationships between them. Those relationships allow a platform to recognize that activity occurring across different devices, browsers, sessions, and environments may represent the same person.
Those relationships don’t appear automatically. Each edge exists because the platform was in a position to establish a connection based on something it observed.
A Graph Is What a Platform Has; Resolution Is What It Does
Identity graphs and identity resolution are closely related, but they aren’t interchangeable. Perhaps the best way to look at it is that identity resolution is the process, while an identity graph is the structure it reads from and writes to.
Identity resolution determines when multiple identifiers are believed to represent the same person and connects them. The identity graph stores and maintains those relationships. A graph without resolution would essentially be a static collection of information, whereas a resolution without a graph would have nowhere to sustain the relationships it discovers.
A platform can improve its matching techniques, update its models, or introduce new resolution capabilities, but it can’t simply switch on years of observations it never made. Its graph contains the relationships the platform has had the opportunity to establish over time. In other words, two platforms can use similarly sophisticated resolution techniques and still recognize very different populations, because their graphs were built from different observations.
That’s why comparing resolution capabilities alone only gives you part of the story. You also need to understand where a platform has visibility, what kinds of relationships it has established, and how long it has been accumulating them.
Why Two Platforms See the Same Buyer Differently
As mentioned, identity graphs reflect what their owners have observed, which explains why two platforms can encounter the same buyer, yet develop very different views of that person.
A Graph Can Only Hold What Was Observable
No identity resolution technique can create a reliable relationship from observations a platform never had. If a buyer moves through an environment where Platform A has visibility but Platform B doesn’t, only Platform A can incorporate that activity into its graph. That is a structural constraint, not necessarily a difference in matching quality.
Breadth of Properties
One of the main differences between two platforms is how many websites, apps, and other sources provide data to the graph. Consider the scenario I touched on earlier of a buyer researching a purchase: They might read an article on their phone in the morning, compare products on their laptop at lunch, use an app later that day, and return to another website that evening.
A platform with visibility across several of those environments has more opportunities to observe connections. A platform that sees only one or two encounters receives fewer opportunities.
This is why footprint matters so much: Resolution can only work with the observations available to it. Greater breadth of relevance creates more opportunities to recognize that seemingly separate interactions belong to the same person.
Authenticated Share
Authenticated activity provides another important source of identity information. When someone identifies themselves in an environment, such as by logging in, that interaction can provide a high-confidence anchor for connecting identifiers. Other observations can then potentially extend outward from that anchor. A platform with relatively few authenticated moments has fewer of these high-confidence opportunities and may have to rely more heavily on inferred relationships.
Time Depth
A platform that has been observing relevant interactions for many years has had more opportunities to establish and maintain relationships than one with a much shorter history. Time alone doesn’t guarantee graph quality, though; in fact, stale relationships can actually reduce it. Still, a longer period of relevant observation creates opportunities that can’t be reproduced instantly.
| What Happened | What Platform A Recorded | What Platform B Recorded | Why They Differ |
| Buyer visits mobile web on a site both platforms see. | Platform A records a new mobile identifier. | Platform B records the same mobile interaction. | Both have visibility into this property. |
| Buyer visits desktop on a site only Platform A sees. | Platform A observes an additional desktop identifier. | Platform B has no observation to record. | Property breadth gives Platform A additional visibility. |
| Buyer makes a purchase while logged in. | Platform A connects activity to an authenticated identifier. | Platform B records only the interaction it can observe. | Authentication provides a stronger identity anchor. |
| Buyer later starts an in-app session. | Platform A connects the session to recent observations. | Platform B treats the session as another user. | Timing affects which observations can support a connection. |
| Buyer returns on the original mobile device. | Platform A reinforces its existing user record. | Platform B retains separate user records. | Maintained connections persist while weaker ones may decay. |
What’s Inside a Graph, and Why It Explains the Difference
Here’s a closer look at the basic components of an identity graph, which help explain why platforms can develop different views of the same person.
Nodes: The Identifiers
Nodes are the individual identifiers stored within the graph. One buyer may generate several identifiers across devices, browsers, apps, and other environments. A graph doesn’t assume that the nodes represent separate people; it stores them so that relationships can be evaluated and maintained. Different platforms observe different identifiers, so they start with different sets of nodes.
Edges: The Observed Relationships
Edges connect nodes that the platform believes belong to the same person. For example, a graph might contain one node representing a mobile identifier and another representing a desktop browser. If the platform establishes evidence that both belong to the same person, an edge can connect them.
This is where footprint becomes critical again. A relationship can only be added to the graph if the platform has the opportunity to observe evidence supporting it. Different observations therefore produce different edges.
Confidence
Not every relationship carries the same level of certainty. A usable graph needs to account for the strength of the available evidence supporting each connection. Some edges can carry greater confidence than others, based on the method and evidence used to establish them. Two platforms may therefore contain similar-looking connections, but have different levels of confidence in those relationships.
Recency and Decay
Identity relationships aren’t permanent. Devices change hands, browsers are replaced, identifiers expire, and people’s digital behavior changes. A relationship that accurately represented a person six months ago may no longer be reliable today. Graphs therefore need to account for when connections were observed and whether they remain useful. Two graphs of similar size can differ significantly if one contains fresher, better-maintained relationships.
| Component | What It represents | Example | What It Means When Platforms Differ |
| Node | A node represents an identifier stored in the graph. | The buyer’s mobile browser creates one identifiable node. | Platforms observing different identifiers start with different graphs. |
| Edge | An edge represents a believed relationship between identifiers. | Mobile and desktop identifiers are connected to the buyer. | Different observations create different relationships between nodes. |
| Confidence Score | The score reflects certainty in an identity connection. | The buyer’s authenticated interaction strengthens a connection. | Platforms may assign different confidence to similar relationships. |
| Timestamp/Recency | Recency records when a relationship was last supported. | The buyer’s return visit refreshes an existing relationship. | Fresher observations can make one graph more current. |
How a Graph Accumulates
You should never build an identity graph once and then leave it alone. It needs to be developed through a continuous cycle of observation, connection, reinforcement, and removal. Here are some key points to remember:
Authenticated events can provide anchors. When a person identifies themselves, that event can establish a high-confidence relationship around which other identifiers may be connected.
Observed co-occurrence can extend the graph. Repeated patterns across environments can provide additional evidence that identifiers are related, extending the record beyond its original anchors.
New observations continuously update it. As the platform observes additional interactions, existing edges can be reinforced, and new relationships can be added.
Stale or contradicted edges need to be pruned. A graph that only adds relationships will eventually accumulate outdated connections. Removing or weakening stale edges is therefore just as important as creating new ones.
This continuous accumulation explains why graph quality can’t necessarily be replicated quickly. Two platforms might use similar architecture and identity resolution techniques, but they haven’t necessarily accumulated the same assets.
The difficult part isn’t just building the structure, then — you need to accumulate and maintain useful relationships over time.
What Is a Data Spine?
A data spine, also known as an identity spine, is the persistent, unified identifier that serves as the backbone for an identity graph. It acts as the central connective tissue that stitches together fragmented, anonymous data points (like cookies, device IDs, and IP addresses) into a single, comprehensive view of a consumer.
Without a shared spine, different systems can maintain separate representations of the same person. A targeting system might recognize one set of identifiers, for example, while a measurement system works from another.
The data spine provides a consistent reference point. Incoming signals can be associated with resolved identities, rather than remaining attached only to individual devices, browsers, or interactions.
Signals and identifiers perform different jobs. Identifiers help establish which observations belong to the same person, while signals describe behavior, interests, context, and other information associated with that user. Connecting both through a persistent identity layer allows systems to work from a more consistent user record.
Realize’s Implementation
Realize uses Realize ID as its unified ID, clustering connected identifiers into a persistent view of a user. Those clusters are part of the broader Realize Identity Graph, enabling signals from fragmented touchpoints to contribute to the same user profile.
Why a Bigger Graph Isn’t Automatically a Better One
It can be tempting to assume the largest identity graph must also be the best, but that’s too simplistic. Consider these three points:
- Scale without maintenance can become a liability. A graph that continually adds connections without pruning stale or contradicted edges may look impressive by size, while becoming less reliable.
- Breadth only matters when it’s relevant. Extensive visibility into environments your customers rarely use may add little practical value to your campaigns.
- Total graph size isn’t the same as audience coverage. An advertiser doesn’t need a platform to recognize every person on the internet. They need the platform to recognize a meaningful share of the people relevant to its business.
That leads to a more useful question than simply asking, “How big is your identity graph?” Instead, you should ask, “How much of my audience does your graph recognize, and how confidently?” The answer will provide more information about how useful that accumulated identity asset may be for your campaigns.
How to Evaluate a Platform’s Graph
When comparing identity graphs, focus on the observations and relationships behind the numbers. Here are five questions that can help:
- Across what properties does the graph observe? Look at the sites, apps, and environments where the platform can establish identity relationships.
- What share of that activity is authenticated? Authenticated moments can provide high-confidence anchors for identity connections.
- How long has the graph been accumulating? Relevant time depth can create more opportunities to establish and reinforce relationships.
- How are stale relationships removed? A graph should maintain and prune connections, rather than simply accumulate them indefinitely.
- How much of your own audience does the graph recognize? Relevant coverage matters more than the size of the headline graph.
Key Takeaways
An identity graph isn’t simply a feature a platform can turn on, it’s an accumulated asset built from the identity relationships a platform has observed. That’s why two platforms can use similar resolution techniques, but recognize the same buyer differently. Breadth of properties, authenticated activity, and time depth can all influence what enters the graph.
Size alone doesn’t determine quality, either. Relationships need to remain current and relevant to the advertiser’s audience.
When evaluating identity graphs, look beyond the technology and headline numbers. Ask what the platform has actually observed, how well those relationships are maintained, and how much of your audience the graph can confidently recognize.
Frequently Asked Questions (FAQs)
What’s the difference between an identity graph and identity resolution?
Identity resolution is the process; an identity graph is the structure it reads from and writes to. Resolution determines when identifiers are believed to belong to the same person. The graph stores and maintains those relationships, including information about their confidence and recency. For more identity resolution terms, please see our Realize ID glossary.
Why do two platforms report different reach for the same audience?
Platforms can recognize different populations because their identity graphs contain different observations and relationships. Differences in property breadth, authenticated activity, time depth, and graph maintenance can cause the same person to appear as a single resolved user on one platform, and as multiple users on another.
How do identity graphs recognize the same person across devices?
Identity resolution evaluates available identifiers and observations to determine whether different devices are likely to belong to the same person. When a connection is established, the graph records a relationship between those identifiers and can associate them with the same persistent user record.
Can you build your own identity graph?
Yes. Advertisers and agencies can build identity graphs from their own customer data and observed identity relationships. However, what the graph can recognize will still depend on the data and environments available to its owner. An advertiser’s existing graph can also serve as an input to another platform’s identity resolution process.
Do identity graphs store personal information?
It depends on how a particular graph is designed and the identifiers and data it uses. Identity graphs may contain or link identifiers such as hashed email addresses, mobile advertising IDs, publisher IDs, and other identity references. The exact data stored varies by platform and implementation.