WebsiteCategorizationAPI
Home
Demo Tools - Categorization
Website Categorization Text Classification URL Database Taxonomy Mapper Carbon Footprint
Demo Tools - Website Intel
Technology Detector Quality Score Competitor Finder
Demo Tools - Brand Safety
Brand Safety Checker Brand Suitability Quality Checker
Demo Tools - Content
Sentiment Analyzer Context Aware Ads
Resources
API Documentation Pricing Login
Try Categorization
AdPage Intelligence

Try a live demo

AdPage Intelligence · 4 production APIs

Score any URL for Ad-Tech Risk

MFA detection, COPPA compliance, audience segmentation and carbon footprint — real scores from production endpoints, tested against URLs you already know.

MFA Detection 78 vs 11 Mean score: known MFA sites vs premium publishers. Zero false positives, zero false negatives.
COPPA Compliance 15 / 0 Child-directed classifier fires on all 15 kids sites, zero adult sites mis-flagged.
Audience Segmentation 50 / 50 Full audience profile returned for every URL tested. No gaps, no fallbacks.
Carbon Footprint A – F Letter grade per URL from 5 sub-scores: page weight, ad load, video, supply chain, hosting.

What is in these sample sets

Every sample set was generated with the same public API endpoints a customer calls. Each one ships three artifacts: a CSV with one row per URL and the key scores, a full JSON with the complete API response per URL for technical evaluation, and a one-page written summary of what the data shows. All of them are downloadable from this page, with no form in front of them.

MFA Detection

50 URLs: 15 publicly documented MFA sites, 15 premium publishers, 10 mid-tier publishers, 10 deliberately borderline properties.

Jump to the results →

How MFA scoring works

COPPA Compliance

50 URLs across five buckets: child-directed, general child appeal, adult-directed, kids sites with data collection, and well-governed children's brands.

Jump to the results →

How COPPA scoring works

Audience Segmentation

50 URLs across 10 content verticals: personas, demographics, purchase intent, life stage and B2B firmographics.

Jump to the results →

How segmentation works

Carbon Footprint

Grade any URL A–F for carbon intensity. Resource weight, third-party bloat, green hosting and optimization signals in one sustainability score.

Jump to the results →

How carbon scoring works

1. MFA Detection — 50 URLs

The set is built to be adversarial rather than flattering: 15 domains publicly documented as Made-for-Advertising (Adalytics' 2024 "Made for Arbitrage" study and its network-graph clusters), 15 premium publishers that run heavy programmatic stacks and would break a naive ad-tech-count heuristic, 10 mid-tier regional and enthusiast publishers, and 10 borderline properties (celebrity, listicle and aggregation sites) where reasonable people disagree.

Known MFA

n = 15

78 mean score (min 78 / max 78)

Premium publishers

n = 15

11 mean score (min 0 / max 25)

Mid-tier publishers

n = 10

15 mean score (min 0 / max 25)

Borderline

n = 10

30 mean score (min 0 / max 79)

The headline separation: known MFA averaged 78 against a premium-publisher average of 11. Counting a premium publisher scoring above 45 as a false positive and a known MFA site scoring below 46 as a false negative, this sample produced zero of each. The borderline bucket is where the interesting behaviour lives: its 0–79 range is the model declining to give a single verdict to a mixed bag, which is the correct answer.

Representative rows from mfa_detection_sample.csv
DomainBucketMFA scoreRisk tier
heraldweekly.com known_mfa 78 HIGH
kueez.com known_mfa 78 HIGH
thedaddest.com known_mfa 78 HIGH
trend-chaser.com borderline 79 HIGH
suggest.com borderline 72 HIGH
celebritynetworth.com borderline 29 LOW RISK
boredpanda.com borderline 11 CLEAN
dexerto.com midtier_publisher 25 CLEAN
cleveland.com midtier_publisher 19 CLEAN
cnn.com premium_publisher 18 CLEAN
theguardian.com premium_publisher 23 CLEAN
bbc.com premium_publisher 7 CLEAN

Read the borderline rows carefully. trend-chaser.com and suggest.com were not on any public MFA list — they were scored HIGH on their own structural and content evidence. That is the case the score exists for: the arbitrage site nobody has published about yet. The full CSV carries every underlying signal (ad script count, header-bidding depth, ad-to-content ratio, word count, clickbait, trustworthiness, content originality, the public flag source, and the LLM's evidence sentence).

How MFA scoring works Try the live demo Download CSV Full JSON

2. COPPA Compliance — 50 URLs

Five buckets, chosen so that the sample can fail visibly: 15 unambiguously child-directed sites, 10 with general child appeal (gaming, UGC, wikis), 10 adult-directed business and technology sites, 10 kids' properties with real signup and data-collection flows, and 5 well-governed children's brands. Every page is rendered in headless Chrome with the consent dialog auto-accepted, so consent-gated trackers are actually observable.

Child-directed

n = 15 · mean risk 38 (2 LOW)

Child-directed detection fired on 15 of 15. Mean data-collection sub-score 2.3/20; mean tracking sub-score 10.2/20.

General child appeal

n = 10 · mean risk 27 (6 LOW)

Child-directed on 4 of 10. Data-collection 2.6/20; tracking 5.7/20.

Adult-directed

n = 10 · mean risk 28 (5 LOW)

Child-directed on 0 of 10. Data-collection 3.3/20; tracking 10.3/20.

Child data collection

n = 10 · mean risk 31 (6 LOW)

Child-directed on 4 of 10. Data-collection 2.4/20; tracking 10.0/20.

Good compliance

n = 5 · mean risk 28 (3 LOW)

Child-directed on 5 of 5, but data-collection 0.6/20 and tracking 5.2/20 — the well-governed profile.

The finding that matters

Child-directed sites carry adult-grade tracking loads (10.2/20 vs 10.3/20 on adult sites). Combined with is_child_directed=true, that is precisely the exposure a compliance buyer needs surfaced.

How to read it: the child-directed classifier is the anchor — it fires on every clearly child-directed site in the set and on zero adult sites. The risk score itself is deliberately not a verdict; it is a composite of five sub-scores (child-directed content, personal data collection, tracking and persistent identifiers, absent privacy controls, content appropriateness). Well-governed children's brands — policies that mention COPPA, light tracking — sit at the bottom of the range, exactly where they should.

Representative rows from coppa_compliance_sample.csv
URLBucketRisk scoreLevelChild-directed
www.coolmathgames.com child_directed 49 MEDIUM YES
www.abcya.com child_directed 47 MEDIUM YES
www.mathplayground.com child_directed 44 MEDIUM YES
www.poptropica.com child_directed 40 MEDIUM YES
poki.com general_child_appeal 40 MEDIUM YES
www.roblox.com general_child_appeal 33 MEDIUM YES
outschool.com/signup child_data_collection 42 MEDIUM YES
www.getepic.com/signup child_data_collection 36 MEDIUM YES
tocaboca.com good_compliance 20 LOW YES
techcrunch.com adult_directed 44 MEDIUM no
www.marketwatch.com adult_directed 20 LOW no

How COPPA scoring works Try the live demo Download CSV Full JSON

3. Audience Segmentation — 50 URLs

50 URLs across 10 content verticals — luxury, B2B enterprise, parenting, finance, health, automotive, tech and gaming, food, fashion and beauty, education. All 50 returned full audience profiles. Every URL combines static persona mapping (which costs no LLM call) with LLM-inferred demographics, purchase intent, life stage, B2B firmographics and content context.

B2B detection

5/5 enterprise-content URLs classified is_b2b=true.

Luxury vertical

4/5 URLs assigned a high / affluent income level.

Coverage

50/50 URLs returned a full profile — no gaps to explain away.

Representative rows from audience_segmentation_sample.csv
URLVerticalTop personasB2BTop purchase intent
robbreport.com/motors/cars/ luxury_affluent Exotic Car Enthusiast; Luxury Car Enthusiast; Auto Enthusiast B2C luxury vehicles (0.90)
www.hodinkee.com luxury_affluent Minimalism Enthusiast; Fashion Enthusiast; Organic Food Advocate B2C luxury watches (0.90)
www.snowflake.com/en/blog/ b2b_enterprise Infrastructure Engineer; Data Scientist; Cloud Computing Specialist B2B technology (0.80)
hbr.org b2b_enterprise Financial Analyst; Corporate Executive; Financial Advisor B2B business_management (0.85)
stackoverflow.blog tech_gaming Technology Enthusiast; Cloud Computing Specialist; Data Analyst B2B software tools (0.80)
www.whattoexpect.com/first-year/ parenting_family Pediatrician; Childcare Professional; Parenting Blogger B2C baby products (0.90)
www.fool.com finance_investing Cryptocurrency Investor; Beginner Investor; Financial Analyst B2C investment services (0.90)
www.healthline.com health_wellness Health Enthusiast; Exercise Enthusiast; Detox Devotee B2C healthcare products (0.80)
www.ign.com tech_gaming Gamer; Action Enthusiast; Hardcore Gamer B2C video_games (0.90)
www.harpersbazaar.com fashion_beauty Fashion Forward Female; Fashion Enthusiast; Style Enthusiast B2C fashion (0.90)

The CSV carries considerably more per URL than fits in a table: age band, gender skew, income level, education, life stage, B2B roles, content type, reading level, price signals and IAB categories. The JSON carries the complete response, confidence scores included.

How segmentation works Try the live demo Download CSV Full JSON

4. Carbon Footprint Scoring

Grade any URL A–F for carbon intensity. The score (0–100, lower is greener) combines resource weight analysis, third-party script bloat, green hosting detection, image optimization and render efficiency into a single sustainability metric. Six grade tiers: A (0–15), B (16–30), C (31–50), D (51–70), E (71–85), F (86–100). Use it for sustainability reporting, green advertising programs and brand placement decisions that factor in carbon cost.

Resource weight

Total page weight, JavaScript payload, CSS payload, image weight and font loading.

Third-party bloat

Number and weight of third-party scripts, ad-tech load, tracking overhead and render-blocking resources.

Green hosting

CDN usage, hosting provider green energy status, HTTP/2 adoption and caching efficiency.

How carbon scoring works Try the live demo

Methodology — including the limits

A sample report that only shows what worked is a sales asset, not evidence. These are the caveats we would raise ourselves in front of a media auditor, and they are stated in the README that ships with the data.

Cloaking is real, and it limits what a crawler sees

Several documented MFA operators serve ad-free, harmless-looking pages to datacenter IPs. The monetization_observed field reports what our crawler actually saw, kept separate from the domain flag — a flagged domain showing no ads to us is a cloaking fingerprint, not innocence.

Multi-step signup flows cannot be fully traversed

A single-fetch API cannot walk a multi-step JavaScript registration wizard. Those pages report signup_cta_detected rather than an enumerated list of PII fields. We report the CTA; we do not claim to have completed the form.

MFA supply churns quickly

Every URL was live and returning HTTP 200 with real content at generation time (2026-07-13). MFA domains turn over month to month — re-verify liveness immediately before any live demo, and re-score inclusion lists on a schedule.

Consent is accepted, deliberately

COPPA pages are rendered in headless Chrome with CMP auto-accept. Crawling the EU without accepting consent hides most of the tracking stack, which would make every site look cleaner than it is.

The buckets are curated, and the labels are ours

Bucket assignments (known MFA, premium, borderline) come from public research and editorial judgment, not from the model. That is what makes false positives and false negatives measurable in the first place — and it is also why you should check the domains you know personally.

Sampling

URLs owned by top-100 advertisers were excluded from the sample sets to avoid conflicts of interest in an audit context. Scores are produced by the same public endpoints a customer calls — no offline tuning, no per-domain hand-editing.

Try it yourself — score any URL right now

Each demo scores a live URL in real time using the same production endpoints behind the sample data above. No signup, no API key — just paste a URL and see the result.

MFA Detection COPPA Compliance Audience Segmentation Carbon Footprint
Request a Custom Sample Report Read the API Docs
Stay in the loop

You are on the list!

We will send you updates that matter — no spam.