WebsiteCategorizationAPI
Home
Demo Tools - Categorization
Website Categorization Text Classification URL Database Taxonomy Mapper Carbon Footprint
Demo Tools - Website Intel
Technology Detector Quality Score Competitor Finder
Demo Tools - Brand Safety
Brand Safety Checker Brand Suitability Quality Checker
Demo Tools - Content
Sentiment Analyzer Context Aware Ads
Resources
API Documentation Pricing Login
Try Categorization
Cookieless Audience Intelligence · Data Types

Purchase Intent Data: In‑Market Signals Without Tracking Users

Purchase intent data is any signal indicating that someone — or an audience — is actively considering a purchase in a category right now: comparing models, checking prices, reading reviews. It has classically been harvested by tracking individuals. This guide explains the traditional sources, and a newer alternative: inferring the in-market state of a page's readership from the page's content itself — 283 intent segments in 34 groups, no cookies, no IDs, no tracking.

34Intent groups (PI.*)
283In-market segments
102MDomains pre-scored
0Users tracked
Definition

What is purchase intent data?

Purchase intent data captures the difference between someone who likes a category and someone who is about to buy in it. A person can be interested in cars for thirty years; they are in-market for a car for perhaps ninety days of that. Intent signals are the observable traces of that in-market window: the searches, comparisons, calculator sessions, spec-sheet downloads and review-reading binges that cluster tightly around a purchase decision and then stop.

That time-bounded quality is what makes intent data the most valuable targeting input in advertising. Interest data describes a durable affinity and is useful for reach and brand planning; intent data describes a transient state and is useful for conversion. An ad for mortgage rates shown to someone actively comparing lenders performs on a different order from the same ad shown to a general finance audience — which is why in-market segments have historically commanded the highest CPMs of any third-party data category, and why B2B intent platforms became a standard line item in enterprise marketing budgets.

The catch has always been sourcing. Most intent data was built by observing individuals — their queries, their browsing, their transactions — which means it inherited every problem of user-level tracking: cookie dependence, consent requirements, coverage that collapses on Safari, Firefox and iOS (roughly 40%+ of traffic carries no usable third-party cookie today, and third-party cookies persisting in Chrome does not change that), and growing regulatory exposure under GDPR, CCPA and their successors. This page is part of our hub on cookieless audience segmentation, and it looks at intent through that lens: what the classic sources actually measure, what a content-inferred alternative measures instead, and how the two compare on privacy, coverage, latency and granularity.

One definitional point worth fixing early: intent exists at different resolutions. It can be claimed about an individual ("this user is in-market for an SUV"), an account ("this company is researching data warehouses"), or an audience ("the readership of this page skews heavily toward people shopping for a new vehicle"). The three are often conflated in vendor marketing. They should not be — they have different evidentiary bases, different privacy footprints and different appropriate uses, as the rest of this guide makes explicit.

The classic sources

Where purchase intent data traditionally comes from

Five source families account for nearly all commercial intent data. Each observes a different behavior, at a different resolution, with different blind spots — understanding them is the fastest way to evaluate any intent product's claims.

Search-query data

Queries are the purest intent signal that exists: "best 7-seater EV 2026 price" is a person announcing their in-market state in their own words. Search engines monetize this directly through search ads, and query-derived audiences power much of the in-market segmentation inside the large ad platforms' walled gardens.

Limit: the raw signal stays inside the platforms that captured it; outside them, buyers get pre-packaged segments with little transparency into recency or construction.

Site-behavior data

Product views, cart adds, configurator sessions, pricing-page visits — behavioral events collected on a marketer's own properties (first-party) or historically across the web via third-party cookies and SDKs. This is the signal behind retargeting and most "in-market" segments sold on data marketplaces.

Limit: cross-site collection depends on tracking infrastructure that is blocked by default on Safari, Firefox and iOS, and consent-gated everywhere else.

Panel data

Opted-in panels of users who share their browsing, search or purchase activity in exchange for compensation. Panels observe deep, longitudinal behavior for a small consented population, which is then statistically projected onto the broader market. Widely used for measurement, category sizing and calibrating other datasets.

Limit: panel sizes are small relative to the web, projection introduces modeling error, and niche categories may have too few in-panel buyers to read.

Transaction data

Card networks, retailers, receipt-scanning apps and e-commerce platforms see actual purchases — the ground truth every other source approximates. Transaction-derived segments ("bought baby products in the last 90 days") are strong predictors of adjacent and repeat purchases.

Limit: it is inherently backward-looking — it tells you what was bought, not what is being considered — and it is among the most sensitive personal data categories under privacy law.

B2B intent platforms

B2B intent vendors watch for content-consumption spikes at the account level: when an unusual number of readers resolved to one company start consuming content about, say, container orchestration — across publisher co-op networks, bidstream data or the vendor's own media — that account is flagged as "surging" on the topic and pushed to sales and ABM teams.

Limit: IP-to-company resolution is noisy (VPNs, remote work), topic taxonomies are vendor-proprietary, and the observation network covers only a slice of relevant reading.

Content-inferred intent

The newest family, and the subject of the rest of this page: instead of observing people, analyze the page. What a URL is about reveals the likely in-market state of its readership — no individual is observed at all. It is the intent analogue of contextual targeting, upgraded from "what is this page about" to "what is this page's audience shopping for".

Limit: it describes audiences, not individuals — a constraint we treat honestly in the limitations section below.
The cookieless alternative

Content-inferred intent: reading the in-market state of a page's audience

Consider a mortgage-calculator page. Nobody reads it for fun. Its audience is, almost by definition, people in-market for a mortgage — and that is knowable from the page alone, without observing a single visitor. The same logic holds across the web: a "best CRM for small business 2026" comparison is read by CRM buyers; a hotel-review roundup for Lisbon is read by people planning travel; an EV charging-cost explainer is read by people considering an electric vehicle. The content selects the audience.

Content-inferred intent systematizes this. A model reads the page — its topic, angle, funnel position, price signals, comparison structure — and outputs the purchase-intent segments its readership plausibly occupies, each mapped to a controlled vocabulary and carried with a banded confidence (low / medium / high). The unit of analysis is the page or domain, never the person. That single design choice is what makes the approach privacy-safe by construction: there is no personal data anywhere in the pipeline, so there is nothing to consent, store, hash or delete.

In our implementation, intent is one of five signal families (alongside demographics, life stage, B2B firmographics and personas) returned by the audience segmentation API for any URL in real time, and pre-computed across a 102M-domain dataset for planning and curation work at scale. Intent is only asserted where content actually supports it — a mortgage calculator earns a high-confidence real-estate finance signal; a general news homepage earns none rather than a guess.

  • Page-level resolution: per-URL API calls capture that a site's EV reviews and its motorsport news carry different intent, even on the same domain.
  • Domain-level scale: the 102M-domain dataset supports planning, enrichment and inventory curation without any crawling on your side.
  • Deterministic where possible: personas derive from a fixed IAB-category-to-persona mapping; inferred attributes like intent carry explicit confidence bands instead of false precision.
  • Versioned vocabulary: every intent value comes from the controlled PI.* code list (v1.0), aligned with the IAB Audience Taxonomy 1.1 — not free-text model output.

The intent vocabulary at a glance

34purchase-intent groups (PI.* tier 1)
283in-market segments (PI.* tier 2)
v1.0versioned controlled vocabulary
IAB 1.1aligned with IAB Audience Taxonomy
3 bandsconfidence: low / medium / high
102Mdomains with pre-computed profiles

Intent codes are stable identifiers (PI.travel.hotels_and_resorts), so segments survive model updates, joins and multi-vendor pipelines. The full branch is browsable on the audience segmentation taxonomy page.

Side by side

Panel & behavioral intent vs content-inferred intent

Neither approach dominates the other on every axis — they answer different questions. The honest comparison looks like this.

DimensionPanel / behavioral intent (user-observed)Content-inferred intent (page-observed)
Data basis Observed actions of individuals or accounts: queries, browsing events, transactions, panel activity, IP-resolved content consumption. The page itself: topic, funnel position, comparison structure, price signals. The audience is inferred from what the content selects for.
Privacy Processes personal data; requires consent management, contracts and deletion workflows; exposure under GDPR/CCPA and sensitive-category rules. No personal data processed at any stage; privacy-safe by construction. No consent dependency, no data-subject requests, no browser-policy risk.
Coverage Bounded by the tracking footprint: collapses where third-party cookies are blocked (Safari, Firefox, iOS), thins with consent opt-outs, and panels cover a small projected sample. Works on 100% of pages and traffic, including cookieless environments, new visitors and unconsented sessions — the signal lives in the content, not the browser.
Latency Segments are built from accumulated observations; individual membership often lags behavior by hours to weeks, and stale membership decays silently. Evaluated at request time per URL (or from the pre-computed domain dataset); a new page can carry intent signals the moment it is published.
Granularity Individual or account level — genuinely per-person when the data is good, which is its core strength and its core privacy cost. Audience level: the aggregate in-market skew of a page's readership. Precise about pages, deliberately silent about persons.
Best use Retargeting, closed-loop measurement, sales triggers on named accounts, suppression of existing customers. Prospecting reach, in-market contextual targeting, inventory curation, seller-defined audiences, account scoring by what a company's buyers read.

In practice sophisticated buyers run both: user-observed intent where consented first-party relationships exist, and content-inferred intent to extend in-market reach across the growing share of traffic where user-level signals are unavailable. The deeper comparison of the two philosophies — observation versus inference — is covered in our guide to contextual vs behavioral targeting.

The taxonomy

The intent categories: 34 groups, 283 segments, PI.* codes

Intent data is only as useful as the category system behind it. Free-text model output ("seems interested in cars?") cannot be traded, joined or audited — controlled codes can.

Every intent value we return comes from the Purchase Intent branch of our controlled vocabulary: 34 tier-1 groups containing 283 tier-2 in-market segments, each with a stable machine-readable code in the PI.* namespace and a human-readable label. The branch covers the complete IAB Purchase Intent structure — the vocabulary is versioned (v1.0) and aligned with the IAB Tech Lab Audience Taxonomy 1.1, which means segments built on it slot directly into industry pipelines that speak IAB, from data-marketplace listings to Seller Defined Audiences. How that alignment works across all our attribute families — and why standard codes matter for interoperability — is the subject of our companion guide to the IAB Audience Taxonomy.

Groups span consumer and business purchasing: Automotive (ownership, products and services), Travel and Tourism, Finance and Insurance, Software, Consumer Electronics, Real Estate, Education and Careers, Business and Industrial, and thirty more. Segments are where activation happens — not "Travel" but Cruise Travel, Hotels and Resorts, Travel Insurance; not "Finance" but Mortgage Lenders and Brokers, Retirement Planning, Credit Cards. The complete browsable structure, with every code and label, lives on the audience segmentation taxonomy page.

PI.auto_ownership Automotive Ownership PI.travel Travel and Tourism PI.finance_insurance Finance and Insurance PI.software Software PI.consumer_electronics Consumer Electronics PI.real_estate Real Estate PI.education_careers Education and Careers PI.business_industrial Business and Industrial PI.health_medical Health and Medical Services PI.home_garden_services Home and Garden Services PI.family_parenting Family and Parenting PI.clothing_accessories Clothing and Accessories + 22 more groups · 283 segments total
Activation

How teams activate purchase intent segments

The same PI.* signals feed very different workflows depending on who is holding them.

In-market audience targeting

Buy the pages whose readership is in-market rather than the users a cookie once claimed were. A campaign for travel insurance runs across every page scoring Travel Insurance or Cruise Travel intent at medium-plus confidence — reaching in-market readers on 100% of traffic, including Safari and iOS where behavioral segments cannot follow.

B2B account scoring

Score accounts by what their buying committees actually read: enrich a target-account list with the intent profiles of the trade content, comparison pages and documentation their teams consume, and weight Software or Business and Industrial intent into your ABM prioritization — a content-side complement to surge-based intent vendors.

Curation and Deal IDs

Curation platforms and SSPs package inventory into Deal IDs defined by intent: "auto in-market, high confidence, brand-safe" becomes a curated PMP assembled from the 102M-domain dataset plus per-URL scoring of fresh inventory — a sell-side data product with a derivation trail buyers can audit.

Seller-defined audience segments

Publishers attach PI.* signals to their own inventory and pass them in the bid request as Seller Defined Audiences — IAB-aligned intent codes make the segments legible to any DSP that speaks the standard, turning editorial content about mortgages or EVs into packaged in-market supply.

Sales prospecting

Sales and partnerships teams use intent-scored domain lists as prospecting filters: every domain whose audience shows Logistics and Delivery intent is a lead list for a freight platform; every site with Retirement Planning intent is one for a wealth-management tool selling ad placements or integrations.

Persona-enriched planning

Intent tells you what an audience is shopping for; personas tell you who they are. Combining PI.* segments with our deterministic 1,667-persona layer — covered in the companion guide to audience personas in advertising — yields plans like "First-Time Homebuyer personas on pages with Mortgage Lenders and Brokers intent".

Honest limits

What content-inferred intent can and cannot tell you

Intent data has a long history of overclaiming. These are the constraints of the content-inferred approach, stated plainly — because segments you can defend to a buyer, a lawyer or an auditor are worth more than segments you cannot.

It describes audiences, not individuals

A high-confidence New Vehicles signal on a page means its readership skews toward car shoppers — not that any specific visitor is one. Some readers of an EV comparison are journalists, competitors or the merely curious. For media buying this is the right resolution (media is bought in aggregate impressions), but it cannot replace user-level data for tasks that genuinely require it, like suppressing recent purchasers.

Confidence is banded, not spuriously precise

Inferred intent carries a low / medium / high confidence band, not a decimal probability — because the underlying evidence (how strongly content selects an in-market readership) does not support four significant figures. A mortgage calculator is high-band; a general personal-finance column is low-band or carries no intent at all. Thresholds are yours to set per use case: strict for guaranteed deals, permissive for prospecting.

It is not a purchase prediction

No intent product — behavioral or contextual — predicts that a purchase will occur, and content-inferred intent does not claim to. It identifies where in-market attention concentrates. Conversion depends on the offer, the creative, the price and the moment; intent data improves the odds of being present when consideration happens, nothing more.

Inference quality tracks content quality

Thin, ambiguous or misleading pages yield weak or absent signals — by design, the system returns nothing rather than guessing. And because signals derive from content, a page that changes topic changes profile: per-URL scoring handles this in real time, while domain-level profiles refresh on the dataset's update cycle rather than instantly.

Worked example

From an auto-review page to tradable intent segments

A single API call against an electric-SUV comparison review on an automotive publisher. Coded PI.* values on the left; the same segments rendered as labels for planning and packaging on the right.

POST /api/audience/segment.php
{ "query": "https://autoreviews.example/2026-electric-suv-comparison" }

// response — purchase_intent family (other families omitted)
{
  "purchase_intent": [
    { "code": "PI.auto_ownership.new_vehicles",
      "label": "New Vehicles",
      "confidence": "high" },
    { "code": "PI.finance_insurance.insurance",
      "label": "Insurance",
      "confidence": "medium" },
    { "code": "PI.auto_products.automotive_parts_and_accessories",
      "label": "Automotive Parts and Accessories",
      "confidence": "low" }
  ],
  "vocab_version": "1.0"
}

Rendered for a media plan

In-market — high confidence
New Vehicles
In-market — medium confidence
Insurance
In-market — low confidence
Automotive Parts and Accessories

The review compares purchase prices and trims (high-band New Vehicles), discusses EV insurance costs in one section (medium-band Insurance), and mentions accessories in passing (low-band). Nothing else is asserted — no demographic guess is smuggled in as intent, and the same response carries demographics, life stage and personas as separate families with their own confidence bands. An auto OEM buys the high band; an insurer tests the medium band; an accessories retailer probably passes.

FAQ

Purchase intent data, frequently asked

What is purchase intent data in marketing?
Purchase intent data is any signal indicating that a person, account or audience is actively considering a purchase in a specific category — in-market rather than merely interested. Classic signals include search queries, product-page and pricing-page behavior, panel-observed research activity, transactions, and (in B2B) spikes in a company's content consumption on a topic. It differs from interest data in being time-bounded: interest is a durable affinity, intent is a shopping window.
How is purchase intent data collected?
Traditionally by observing individuals: search engines log queries, sites and SDKs record browsing and cart events, opted-in panels share their activity, retailers and card networks capture transactions, and B2B intent platforms resolve content consumption to companies via IP data. A newer, privacy-safe alternative observes the page instead of the person: content-inferred intent analyzes what a URL is about to determine the likely in-market state of its readership — a mortgage-calculator page is read by mortgage shoppers — without processing any personal data.
What is intent data without cookies?
Cookieless intent data derives in-market signals without third-party cookies or user identifiers. Content-inferred intent is the main approach: a model reads a page's topic, funnel position and price signals and assigns purchase-intent segments from a controlled vocabulary (283 segments in 34 groups, PI.* codes, aligned with the IAB Audience Taxonomy 1.1), each with a banded confidence. Because the signal lives in the content, it covers 100% of traffic — including Safari, Firefox and iOS, where third-party cookies are blocked by default — and involves no consent dependency.
What is the difference between B2B intent data and content-inferred intent?
B2B intent platforms watch for surges in a named company's content consumption on a topic, using publisher co-ops, bidstream observation and IP-to-company resolution — the output is account-level ("this company is researching data warehouses"). Content-inferred intent works from the supply side: it scores what pages and domains are about, yielding audience-level signals ("this page's readership is in-market for enterprise software"). The two are complementary — content-inferred signals can enrich account scoring by profiling what a target account's buyers read, without any tracking network.
How accurate is purchase intent data?
It depends on what is being claimed. User-observed intent can be genuinely individual but degrades with stale membership, consent gaps and tracking blind spots. Content-inferred intent makes a deliberately narrower claim: it describes the aggregate in-market skew of a page's audience, with explicit low/medium/high confidence bands rather than decimal probabilities, and it returns no signal rather than a guess when content is ambiguous. Neither approach predicts individual purchases — intent data of any kind concentrates media where consideration is happening; it does not guarantee conversion.
What are in-market audience segments?
In-market segments group consumers (or audiences) actively shopping in a category — "in-market: New Vehicles", "in-market: Hotels and Resorts", "in-market: Mortgage Lenders and Brokers". In content-inferred systems the segment attaches to inventory rather than to users: every page whose content selects a car-shopping readership joins the New Vehicles in-market pool, which can then be activated as a curated deal, a Seller Defined Audience signal, or a planning filter across a 102M-domain dataset.

See the intent profile of any URL

Paste a page into the live demo and watch the PI.* branch light up — in-market segments with confidence bands, alongside demographics, life stage, personas and B2B signals. No signup, no cookies, no tracking.

Stay in the loop

You are on the list!

We will send you updates that matter — no spam.