WebsiteCategorizationAPI
Home
Demo Tools - Categorization
Website Categorization Text Classification URL Database Taxonomy Mapper Carbon Footprint
Demo Tools - Website Intel
Technology Detector Quality Score Competitor Finder
Demo Tools - Brand Safety
Brand Safety Checker Brand Suitability Quality Checker
Demo Tools - Content
Sentiment Analyzer Context Aware Ads
Resources
API Documentation Pricing Login
Try Categorization

Enrichment Data Fields

Unlock deep insights with our expanded categorization API. Get comprehensive metadata including technologies, topics, entities, personas, sentiment, and more for any URL.

15+
Data Fields
4000+
Technologies
Real-time
Analysis
99.9%
Accuracy
Security Technology Content Audience Business

API Metadata Fields

Detailed explanations for each field returned when using the expanded_categories option with our URL Categorization API.

Malware & Security Check

Performs a comprehensive security scan to detect malware, phishing attempts, social engineering, deceptive downloads, and other malicious content. The check returns a boolean flag indicating whether any threat was detected, making it straightforward to integrate into automated pipelines.

Common use cases: Ad verification companies use this field to block ad placements on compromised pages before they serve impressions. Brand safety platforms incorporate the malware flag into their scoring models to prevent advertisers from appearing alongside dangerous content. ISPs and DNS filtering services check URLs in real time to protect end users from visiting phishing sites. Security operations teams feed URL lists from email attachments and network logs through the API to identify threats before they reach employees. Compliance teams in financial services use malware detection as part of vendor risk assessments, scanning the websites of third-party partners to verify they have not been compromised.

Sample Output "malware_check": true/false

Web Technologies Detected

Identifies 4,000+ front-end and back-end technologies including JavaScript frameworks, CMS platforms, analytics suites, advertising networks, payment processors, hosting providers, marketing automation tools, security solutions, customer service platforms, and development libraries. Technologies are detected by analyzing page source code, HTTP headers, and runtime behavior.

Common use cases: B2B sales teams use technology detection to qualify leads by identifying prospects whose websites run a compatible or competing tech stack. A CRM vendor can target companies using Salesforce by filtering for that technology across millions of domains. Marketing agencies build competitive intelligence reports showing which technologies clients and their competitors have adopted. Investment analysts map technology adoption across entire market verticals to identify emerging trends and disruption risks. Migration consultants discover companies still running legacy platforms as potential clients for modernization projects. Cybersecurity teams identify outdated or vulnerable software versions across their organization's web properties.

Sample Output ["React", "Google Analytics", "Stripe", "Cloudflare", "WordPress"]

Topics & Themes

Extracts high-level subjects and themes discussed on the page, going beyond broad category labels to identify specific stories, product features, promotions, and messaging angles. Each topic is returned with a contextual description explaining why it was identified.

Common use cases: Content recommendation engines use topics to match articles to reader interests at a granular level — not just "Technology" but "MacBook Air M4 performance benchmarks." Editorial teams track which topics trend across competitor publications to identify content gaps and opportunities. Programmatic advertising platforms use topic extraction to enable contextual targeting more precise than category-level matching, placing ads next to content about specific product launches or industry events. Market research firms analyze topics across thousands of domains to map emerging narratives and messaging shifts within an industry vertical.

Sample Output [["MacBook Air performance", "M4 chip highlights"],
["Education savings", "College promotion"]]

Similar Domains

Identifies domains similar in industry, content, and market position. The algorithm considers content overlap, audience similarity, and market positioning to surface domains that a human analyst would recognize as operating in the same competitive space.

Common use cases: Competitive intelligence teams use similar domains to discover businesses they may not have identified through manual research. Advertising platforms build lookalike audience segments by finding domains whose visitors share characteristics with a target site. SEO professionals analyze similar domains to identify link-building opportunities and benchmark their content strategy against comparable sites. Investment firms use domain similarity to map market landscapes and identify acquisition targets or partnership opportunities. Brand safety teams expand their blocklists by discovering domains similar to known problematic sites.

Sample Output ["samsung.com", "microsoft.com", "google.com"]

Key Named Entities

Detects and categorizes important named entities including organizations, products, services, locations, people, and technologies mentioned on the page. Each entity is classified by type (Person, Company, Organization, Location, Product, Event, Technology, Brand, Topic) with contextual evidence from the page content.

Common use cases: Knowledge graph construction teams extract entities from millions of pages to build structured databases of relationships between companies, products, and people. Media monitoring services track brand mentions across news sites, blogs, and forums to measure share of voice and detect emerging PR issues. Ad verification platforms use entity extraction to enforce placement rules — ensuring an automotive brand's ads never appear on pages mentioning competitor vehicles. Compliance teams in regulated industries scan web content for mentions of sanctioned entities or restricted products. Academic researchers use entity extraction to analyze large corpora of web content for digital humanities and social science studies.

Sample Output [["Apple Inc.", "Organization"],
["iPhone 16", "Product"],
["FDA", "Organization"]]

Likely Buyer Personas

Infers potential customer profiles based on products, language, pricing, design patterns, and promotional messaging. Our library includes over 2,000 distinct personas spanning professional roles (CTOs, marketers, procurement managers), interest-based segments (tech enthusiasts, fitness buffs, automotive fans), and behavioral profiles (comparison shoppers, impulse buyers, B2B decision-makers). Each persona includes a confidence score from 0 to 1 indicating signal strength. View all 2,000+ personas.

Common use cases: Programmatic advertising platforms use buyer personas to match ad creatives to page audiences — placing enterprise software ads on pages visited by "IT Decision Makers" and consumer electronics ads on pages attracting "Tech Enthusiasts." Sales development teams score inbound leads by analyzing their company website's detected personas to identify whether the site targets enterprise buyers or consumers. Content personalization engines use persona data to dynamically adjust messaging, CTAs, and product recommendations based on the audience profile of the referring page. Media planners use persona distributions across publisher inventory to optimize campaign placement without relying on third-party cookies.

Sample Output [["Tech Enthusiasts", "Interest in latest hardware"],
["Creative Professionals", "MacBook Pro performance"]]

Related Keywords

Extracts important keywords and search terms directly from the page content — actual vocabulary used on the site, not inferred or generated terms. Keywords represent the core semantic footprint of the page and are valuable for understanding what the page ranks for and what topics it covers.

Common use cases: SEO professionals use extracted keywords to analyze competitor pages and identify ranking opportunities they may be missing. PPC campaign managers discover new keyword ideas by extracting terms from top-performing competitor landing pages. Content strategists identify semantic gaps by comparing keywords across a set of pages covering the same topic. E-commerce platforms use keyword extraction to auto-generate product tags and improve internal search relevance. Affiliate marketers analyze keywords across niche sites to identify high-value content topics worth targeting.

macbook air iphone 16 trade-in apple card

Brand Reputation Analysis

Identifies brands mentioned on the page and evaluates how each brand is portrayed in context: Positive, Neutral, or Negative. This field focuses specifically on brand-level portrayal rather than overall page sentiment, making it distinct from the content sentiment field.

Common use cases: Brand monitoring teams track how their brand is portrayed across thousands of web pages, identifying negative coverage before it escalates into a PR crisis. Ad verification platforms use brand reputation data to prevent advertisers from appearing on pages where their brand or products are being criticized. Competitive intelligence analysts monitor how competitor brands are perceived across industry publications and review sites. Investor relations teams scan financial news sites to assess brand sentiment trends that may affect stock performance. Franchise operators monitor brand portrayal across franchisee websites to ensure messaging consistency and quality standards.

Sample Output [["Apple Inc.", "Neutral"],
["MacBook Air", "Positive"],
["iPhone 16", "Positive"]]

Audience Demographics

Estimates the demographic profile of the website's target audience including age ranges, professional roles, education levels, and interest segments. Demographics are derived from content analysis — examining products offered, pricing tiers, language complexity, and design patterns — rather than from tracking data or cookies.

Common use cases: Media planners use demographic estimates to match campaign audiences to publisher inventory without relying on third-party audience data. Market researchers analyze demographic profiles across hundreds of competitor websites to map the addressable market for a new product launch. Compliance teams verify that advertising placements target appropriate age groups, particularly for age-restricted products like alcohol, gambling, or financial services. Content teams use demographic data to calibrate their writing style, vocabulary complexity, and topic selection for their target reader. E-commerce platforms personalize product recommendations by matching visitor demographics with page-level audience profiles.

Sample Output [["Ages 18-24", "Education promotions"],
["Ages 25-45", "Tech professionals"]]

Sentiment Analysis (Entity)

Evaluates sentiment for each named entity mentioned on the page. Returns the entity name, type classification (Person, Company, Location, Product, Event, Technology, Brand, Topic), sentiment label (positive, neutral, negative), numeric score, and a contextual explanation drawn from the page content. Up to 15 entities are analyzed per page.

Common use cases: Brand safety platforms use entity sentiment to catch nuanced placement risks — a page may have neutral overall sentiment but contain strongly negative coverage of a specific company, which category-level filters would miss entirely. PR and communications teams monitor entity-level sentiment across news sites to detect shifts in how their company or executives are being discussed. Investment analysts correlate entity sentiment across financial news with stock price movements to build predictive models. Product teams track sentiment about their products and features across review sites and forums to identify improvement priorities. Media analysis firms measure how political figures, organizations, and policy topics are portrayed across different publications.

Sample Output (in expanded_categories.data) [["MacBook Air", "Product", "Positive", "Sky-high performance with M4 chip"],
["Goldman Sachs", "Company", "Neutral", "Referenced as Apple Card issuer"]]

Content Sentiment (Overall)

Analyzes the overall emotional tone of the entire page content. Returns a sentiment label (positive, neutral, negative, mixed), a numeric score from -100 to +100, a confidence percentage, a distribution showing what percentage of the content is positive, neutral, and negative, detected tones (professional, persuasive, casual, critical, humorous), and a one-sentence summary. Returned in a separate sentiment_analysis field when expanded_categories=1.

Common use cases: Contextual advertising platforms use content sentiment to ensure ads appear in emotionally appropriate contexts — a luxury brand avoids pages with strongly negative sentiment while a crisis management firm may target them. News aggregators use sentiment scoring to balance their feeds, ensuring readers are not overwhelmed by negative content. Social listening platforms analyze sentiment across blog posts and editorial content to gauge public opinion on trending topics. Customer experience teams analyze sentiment across competitor review pages to benchmark satisfaction levels. Content moderation systems use sentiment as one signal among several to flag potentially problematic user-generated content for human review.

Sample Output (sentiment_analysis field) {"content_sentiment": {"overall_sentiment": "positive", "sentiment_score": 65, "confidence": 85, "distribution": {"positive": 60, "neutral": 30, "negative": 10}, "tone": ["professional", "persuasive"], "summary": "Content is promotional with positive product messaging."},
"entity_sentiment": [{"name": "MacBook Air", "type": "Product", "sentiment": "positive", "sentiment_score": 80, "context": "Sky-high performance highlighted"}]}

Language Detection

Detects the primary natural language of the text content on the page. Supports 100+ languages with high-confidence detection even on pages with mixed-language content, short text, or transliterated scripts.

Common use cases: International advertising platforms use language detection to route ad creatives in the correct language, preventing the common problem of serving English ads on French-language pages. Content distribution networks use language data to direct incoming URLs to the appropriate regional processing pipeline. Compliance teams verify that web properties intended for specific markets actually serve content in the expected language. Translation service providers scan client websites to identify untranslated pages that need localization. SEO teams performing international audits use language detection to verify hreflang implementation matches actual page content.

Sample Output "language": "English"

Legal Entity & Address

Identifies the primary legal entity or organization associated with the website. Extracts registered company names and physical addresses from footers, contact pages, terms of service, and privacy policies when available. Supports entity resolution across different naming conventions.

Common use cases: KYC (Know Your Customer) and anti-money-laundering teams use legal entity extraction to verify the identity behind a website as part of onboarding due diligence. Ad verification platforms match legal entities against advertiser blocklists to enforce brand safety at the corporate level rather than just the domain level. Data enrichment services use legal entity data to deduplicate company records across datasets where the same organization operates multiple domains. Supply chain compliance teams verify that supplier websites belong to the expected legal entities. Insurance underwriters scan client websites to confirm business identity and operating address as part of risk assessment.

Sample Output [["Apple Inc.", "Legal entity from footer"]]

Content Tags

Generates categorical tags derived from page content, similar to blog post tags or editorial labels. Tags are more granular than taxonomy categories and more structured than free-text keywords, making them ideal for filtering, faceted search, and content clustering at scale.

Common use cases: Content management systems use auto-generated tags to organize articles without requiring manual tagging by editors, saving hundreds of hours per month on high-volume publishing sites. E-commerce platforms use content tags to improve internal search and product discovery by adding semantically relevant facets that users can filter by. News aggregators cluster articles by shared tags to build topic pages and trending sections automatically. Research platforms use tags to create filterable indexes across large document collections, enabling analysts to drill into specific subtopics without reading every document. Advertising platforms use content tags as an additional signal layer for contextual targeting, enabling advertisers to target or exclude specific micro-topics.

mac ipad entertainment education fitness

Similar Companies

Lists direct competitors and companies operating in similar spaces, with contextual explanations of the competitive relationship. Based on market analysis, product overlap, industry positioning, and content similarity. Distinct from similar domains: this field returns company names with competitive context rather than just domain addresses.

Common use cases: Sales teams use similar companies data to identify cross-sell opportunities — if a prospect's website lists competitors as similar companies, the sales team can tailor their pitch to address competitive switching. Market research firms map competitive landscapes at scale by analyzing similar companies across hundreds of domains in a vertical. Investment analysts use competitive clustering to identify which companies compete most directly in specific segments, informing portfolio construction and risk assessment. Business development teams discover potential partners by finding companies that operate in adjacent rather than directly competing spaces. Advertising agencies build competitive analyses for client pitches by systematically mapping the competitive landscape using API data rather than manual research.

Sample Output ["Samsung: Consumer electronics",
"Microsoft: Computing hardware",
"Google: AI competitor"]

Ready to Enrich Your Data?

Get started with our expanded categorization API and unlock comprehensive metadata for any URL. Free trial available.

View API Documentation
Stay in the loop

You are on the list!

We will send you updates that matter — no spam.