WebsiteCategorizationAPI
Home
Demo Tools - Categorization
Website Categorization Text Classification URL Database Taxonomy Mapper Carbon Footprint
Demo Tools - Website Intel
Technology Detector Quality Score Competitor Finder
Demo Tools - Brand Safety
Brand Safety Checker Brand Suitability Quality Checker
Demo Tools - Content
Sentiment Analyzer Context Aware Ads
Resources
API Documentation Pricing Login
Try Categorization
Standards Guide · IAB Tech Lab

The IAB Audience Taxonomy: a common language for audience segments

The IAB Tech Lab Audience Taxonomy (current version 1.1) is the industry-standard classification for describing who an audience is — a fixed tree of demographic, interest and purchase-intent segments with stable node IDs, so that a segment means the same thing to a publisher, a data seller, a DSP and an auditor. This guide explains its three branches, how it is used across the programmatic supply chain, and how it differs from the IAB Content Taxonomy.

3Branches: demo / interest / intent
29Interest tier-1 groups
34Purchase-intent groups
1.1Current version
The standard

What the IAB Audience Taxonomy is, and why it exists

Before the Audience Taxonomy, every data vendor named segments its own way. One provider sold “Auto Intenders – Luxury,” another sold “In-Market: Premium Vehicles,” a third sold “High-End Car Shoppers.” A media buyer comparing the three had no way to know whether they described the same people, and a DSP ingesting all three had no key to deduplicate, price or audit them. Audience data was a market with no units of measure.

The IAB Tech Lab's Audience Taxonomy fixed the vocabulary problem. It is a standardized, tiered classification of audience attributes in which every segment concept — “25-29,” “Interest | Travel,” “Purchase Intent | Travel and Tourism | Hotels and Resorts” — has a unique numeric node ID and a fixed position in a parent–child hierarchy. Data sellers map their proprietary segments onto those IDs; buyers and platforms then compare, transact and report on segments in one shared coordinate system. The taxonomy deliberately standardizes the label, not the method: two vendors can still build a “Hotels and Resorts” intent segment from entirely different signals, but the buyer at least knows both claim to describe the same concept — and can ask how each was derived.

Version 1.0 was developed through the IAB Tech Lab's Taxonomy and Mapping working group and released alongside a broader push for data transparency; version 1.1, the current release, refined the node set and naming and is the version referenced by today's programmatic specifications, most visibly Seller Defined Audiences, which transmits Audience Taxonomy IDs in the bid stream. The standard matters more, not less, in a cookieless context: Safari, Firefox and iOS already block third-party cookies — making roughly 40%+ of traffic cookieless today — so audience claims increasingly come from first-party and contextual derivations that buyers cannot verify by retargeting math. A shared taxonomy plus disclosed methodology is what keeps those claims tradable. That is the same reasoning behind our own cookieless audience segmentation approach, which is why our entire output vocabulary is built on this standard.

Structure

The three branches of Audience Taxonomy 1.1

Everything in the taxonomy hangs off three top-level branches. They answer three different questions about an audience — who they are, what they care about, and what they are about to buy — and a real segment is usually a composition across branches: “30-34, Interest | Travel, Purchase Intent | Hotels and Resorts.”

Demographic

Who they are

Stable descriptive attributes of the people in a segment:

  • Age ranges in narrow bands from 18-20 and 21-24 through five-year steps (25-29, 30-34, … 70-74) up to 75+
  • Gender
  • Education & occupation — attainment levels and occupational groupings
  • Household data — household income bands, life stage, home ownership and urbanization (urban / suburban / rural)
  • Marital status
  • Personal finance attributes

This is the branch behind every “A25-54” media plan — see how we infer these signals per page in our guide to website audience demographics.

Interest

29 tier-1 groups · ~500 nodes

Durable affinities — topics people consistently engage with, independent of any imminent purchase. The 29 tier-1 groups fan out into roughly 500 nodes. Sample tier-1 groups:

AutomotiveBusiness and FinanceCareersFood & DrinkHealthy LivingHome & GardenPersonal FinancePetsReal EstateSportsStyle & FashionTechnology & ComputingTravelVideo Gaming

Below each group sit specific nodes — Travel splits into destination and trip-type interests, Sports into individual sports, and so on.

Purchase Intent

34 groups · 800+ nodes

In-market signals — product and service categories a person is actively shopping. It is the largest branch, with 34 groups and over 800 nodes, because commerce needs fine resolution. Sample groups:

Automotive OwnershipClothing and AccessoriesConsumer ElectronicsConsumer Packaged GoodsEducation and CareersFinance and InsuranceReal EstateSoftwareTravel and TourismWeb Services

Intent segments decay: someone in-market for a hotel stops being in-market after booking. Why that matters for pricing and recency is covered in our guide to purchase intent data.

The three branches at a glance

BranchQuestion answeredScale in 1.1Example nodesSignal character
Demographic Who is this audience? Age (18-20 … 75+), gender, education & occupation, household data (income bands, life stage, urbanization), marital status, personal finance 25-29 · 30-34 · household income band · urban Slow-changing; the backbone of reach planning and brand-suitability demos
Interest What do they care about? 29 tier-1 groups, ~500 nodes Travel · Sports · Style & Fashion · Technology & Computing Durable affinity; persists across campaigns; good for upper-funnel targeting
Purchase Intent What are they about to buy? 34 groups, 800+ nodes Travel and Tourism > Hotels and Resorts · Consumer Electronics · Software Perishable in-market signal; highest CPMs; recency and derivation matter most

How nodes are identified

The taxonomy is distributed by IAB Tech Lab as a structured file in which every row is one node. Each node carries a unique numeric ID, a reference to its parent node, and its position expressed as tier columns, so the same concept can be read either as a path (“Purchase Intent | Travel and Tourism | Hotels and Resorts”) or as a single ID. The IDs are the part that machines exchange — a bid request or a segment catalog transmits the number, and both sides resolve it against the same file. This is also why the hierarchy matters operationally: a buyer targeting the “Travel and Tourism” group node can automatically include every child node beneath it, and a reporting system can roll spend up from leaf segments to branch totals without any string matching. When you adopt the taxonomy, the practical work is exactly this mapping exercise: deciding, for each of your segments, which node ID (or set of node IDs) it corresponds to — and recording that crosswalk somewhere buyers can inspect.

In practice

How the Audience Taxonomy is used across the supply chain

The taxonomy is a reference document, not a protocol — but several concrete mechanisms in programmatic advertising are keyed to its node IDs.

Seller Defined Audiences (SDA)

The most direct wiring: SDA lets publishers announce first-party audience segments in OpenRTB bid requests as Audience Taxonomy 1.1 node IDs (the spec's taxonomy registry assigns it segtax 4). The buyer receives standard IDs, not vendor-invented strings. We cover the mechanics, and how we generate the underlying signals, in our Seller Defined Audiences guide.

Data marketplaces and segment catalogs

Data sellers list segments in DSP and curation-platform marketplaces mapped to taxonomy nodes, so a buyer searching “Purchase Intent > Travel and Tourism” sees comparable products from every provider side by side, instead of guessing which proprietary names overlap.

DMP / DSP segment naming and interoperability

Platforms use the taxonomy as an internal schema: segments created in one system can be exported, matched or deduplicated in another because both sides key on the same node IDs. It also gives reporting a stable dimension — delivery by taxonomy branch and group rather than by ad-hoc segment strings.

Transparency and auditability

Paired with the IAB Tech Lab's data-transparency work, taxonomy IDs let a seller disclose, per segment, what concept is claimed and how membership was derived. Standard labels make disclosures comparable — and make it visible when two “identical” segments were built from very different evidence.

Versions & scope

Versioning — and how it differs from the IAB Content Taxonomy

Versioning. The Audience Taxonomy has moved slowly and deliberately: 1.0 established the three-branch structure and node-ID scheme; 1.1 is a refinement release — cleaned-up nodes and naming — and has remained the referenced version in downstream specs for years. That stability is a feature. Segment IDs are embedded in bid streams, deal terms, catalogs and measurement pipelines; a taxonomy that churned quarterly would break the very comparability it exists to provide. Practically, this means systems built against 1.1 IDs today are built on settled ground.

Content vs. audience — two taxonomies, two questions. The IAB Tech Lab also publishes the Content Taxonomy (2.x and 3.x), and the two are constantly confused. The distinction is simple: the Content Taxonomy classifies what a piece of content is about (“this page is about Cruise Travel”) and drives contextual targeting, brand safety and suitability. The Audience Taxonomy classifies people (“these users are in-market for cruises”) and drives audience-based buying. They are separate trees with separate ID spaces — in OpenRTB they are even carried in different places (content signals on the page object, audience signals on the user object, each with its own segtax value) — and an Audience node ID must never be interpreted as a Content node ID or vice versa. The two interlock in practice, though: content classification is often the evidence from which audience is inferred. If you need the content side, we maintain equivalent guides to the IAB Content Taxonomy 2.2 and Content Taxonomy 3.0, both of which our classification API returns.

Rule of thumb: Content Taxonomy = what the page is about. Audience Taxonomy = who the people are. Contextual audience intelligence is the bridge: read the page with the first, infer the second — without observing a single user.

Our implementation

How our audience vocabularies align with Audience Taxonomy 1.1

Our audience segmentation API infers the likely audience of a page or domain from its content — no cookies, no user tracking — and returns every attribute from controlled, versioned vocabularies (v1.0) that are explicitly aligned with Audience Taxonomy 1.1. Alignment is documented, not asserted: each vocabulary value records its crosswalk to the IAB node IDs it corresponds to, so IAB-keyed systems can consume our output directly.

Demographics: coarser by design, crosswalked by ID

Because we infer from content rather than consumer panels, our demographic bands are deliberately coarser than the taxonomy's narrow ranges: 8 age brackets (13-17 through 75+), a 5-point gender skew, 6 income bands, 7 education levels and 14 life stages. Each value documents exactly which IAB Demographic node IDs it spans — e.g. our 25_34 bracket covers the taxonomy's 25-29 and 30-34 nodes — so the mapping is mechanical, never interpretive.

Interests: built on the 29 tier-1 backbone

Our interest vocabulary uses the Interest branch's 29 tier-1 groups as its backbone (INT.* codes such as INT.travel and INT.tech_computing), with 285 sub-interests beneath them. Open-ended model output is snapped onto these canonical codes by embedding similarity; anything that fails to map is dropped, never invented, so the enumeration stays closed.

Purchase intent: full branch coverage as PI.* codes

The purchase-intent vocabulary covers the complete IAB Purchase Intent branch — all 34 groups, expressed as stable PI.* codes (PI.travel.hotels_and_resorts, PI.software.computer_software) across 283 segments, with zero custom additions. Every intent signal carries a banded confidence (low / medium / high). The activation side of these signals is the subject of our purchase intent data guide.

Two derivation notes for completeness: personas are the one deterministic layer — a fixed IAB-category-to-persona mapping (a 1,667-persona taxonomy) that always yields the same personas for the same category — while all other attributes are model-inferred from content with banded confidence. The full code lists, the vocabulary JSON and the per-value IAB crosswalks are published on the audience segmentation taxonomy page, and the same vocabularies drive both the per-URL real-time API and the 102M-domain dataset used for planning and curation.

Worked example

From coded output to IAB-aligned labels

A European city-break guide (“3 days in Lisbon: where to stay and what it costs”) run through the structured endpoint. Left: the coded response — every value from the closed vocabulary. Right: the same response rendered as labels, the way a planner or a deal description would read it.

{
  "vocab_version": "1.0",
  "audience_type": "b2c",
  "demographics": {
    "age_bracket": ["25_34", "35_44"],
    "gender_skew": "balanced",
    "income_level": "middle",
    "life_stage": ["young_professional",
                   "newlywed_couple"],
    "urbanicity": "urban",
    "confidence": "medium"
  },
  "interests": {
    "tier1": ["INT.travel"],
    "tier2": ["INT.travel.europe_travel"],
    "confidence": "high"
  },
  "purchase_intent": {
    "codes": ["PI.travel.hotels_and_resorts",
              "PI.travel.air_travel",
              "PI.travel.sightseeing_tours_and_activities"],
    "confidence": "high"
  }  // each code carries its IAB 1.1 node crosswalk
}

Rendered for humans

Demographics · medium confidence 25-3435-44Balanced genderMiddle incomeYoung ProfessionalNewlywed CoupleUrban
Interests · high confidence TravelEurope Travel
Purchase intent · high confidence Hotels and ResortsAir TravelSightseeing Tours and Activities

Read as a deal description: “25-44 urban city-break travelers, in-market for hotels, flights and tours” — composed from Demographic, Interest and Purchase Intent branches exactly as the taxonomy intends, and derived entirely from what the page is about.

FAQ

IAB Audience Taxonomy: frequently asked questions

It is a standardized classification published by IAB Tech Lab for describing audience segments in advertising. It organizes audience attributes into three branches — Demographic, Interest and Purchase Intent — each a tiered hierarchy of nodes with unique IDs, so that data sellers, publishers, DSPs and buyers can describe, compare and transact audience segments using one shared vocabulary instead of proprietary names.

The Content Taxonomy classifies what content is about (topics of a page or video) and is used for contextual targeting and brand safety; the Audience Taxonomy classifies who people are (demographics, interests, purchase intent) and is used for audience-based buying. They are separate trees with separate node IDs and are carried in different fields of a bid request, though content classification is often the evidence from which audience characteristics are inferred.

Version 1.1 contains three top-level branches. The Interest branch has 29 tier-1 groups expanding to roughly 500 nodes; the Purchase Intent branch is the largest, with 34 groups and over 800 nodes; and the Demographic branch covers age ranges (18-20 up to 75+), gender, education and occupation, household data such as income bands, life stage and urbanization, marital status, and personal finance attributes.

Version 1.1 is the current release. It refined the structure and naming introduced in 1.0 and is the version referenced by downstream specifications — notably Seller Defined Audiences, whose taxonomy registry carries Audience Taxonomy 1.1 as its audience segment vocabulary. The taxonomy is intentionally slow-moving because its node IDs are embedded in bid streams, catalogs and reporting pipelines.

In SDA, a publisher attaches audience segments to OpenRTB bid requests as Audience Taxonomy 1.1 node IDs together with a taxonomy identifier (segtax 4). The buyer's DSP reads standard IDs rather than publisher-specific segment names, which makes first-party, cookieless audience packages comparable across publishers and auditable against the taxonomy's definitions.

Yes. The taxonomy defines segment vocabulary, not collection methodology, so segments derived from first-party data or from content-based inference can be expressed in it just as cookie-derived segments were. Our approach infers the likely audience of a page or domain from its content alone, then expresses the result in vocabularies crosswalked to Audience Taxonomy 1.1 node IDs — privacy-safe by construction, since no user is ever observed.

Related resources

See IAB-aligned audience output on real pages

Run any URL through the live demo and get demographics, INT.* interests and PI.* purchase-intent codes back — every value crosswalked to Audience Taxonomy 1.1, inferred from content alone.

Try the live demo Explore the API
Stay in the loop

You are on the list!

We will send you updates that matter — no spam.