Product deep dive
Content Protector
Purpose-built detection for scrapers: the automation that does not attack you, it simply copies you — prices, catalogues, listings, availability, editorial content and, increasingly, training data.
Why scraping needs its own control
More than half of all web traffic is generated by bots. Some — search engine crawlers, uptime monitors, partner integrations — are good bots you want to keep. Others steal data, generate spam, or hoard inventory. Both categories consume origin capacity, so there is value in managing all bot traffic, not just the malicious slice.
Scrapers are not trying to break in, so WAF signatures never fire. They are often paced below rate limits, run from residential proxy pools, and drive real browsers via headless automation. Yet the business damage is concrete and measurable:
- Skewed analytics. Conversion rate and cart-abandonment figures are polluted by traffic that was never going to buy, so merchandising and pricing decisions are made on corrupted data.
- Wasted ad spend. Advertising budget is consumed by unidentified bots clicking through paid placements.
- Performance and cost. Aggressive crawling raises origin egress, database load and infrastructure cost, and degrades availability for real users.
- Competitive loss. Harvested pricing and inventory data is monetised directly — competitors reprice against you within minutes.
- Scalping and inventory hoarding. Automated buy-out of limited stock or event seats for resale.
- Privacy breaches. Profile-data harvesting exposes your users, with regulatory consequences that land on you, not the scraper.
- Content and AI ingestion. Editorial content republished elsewhere, or absorbed into model training corpora without licence or attribution.
The detection problem is that a good scraper and a good customer both issue well-formed requests for product pages. Content Protector therefore introduces detection methods designed specifically for scrapers, and is able to classify a client even on the first request of a session, before any behavioural history exists.
The four evaluation layers
Content Protector evaluates behaviour, device information and request details to identify clusters of bots. Each anomaly found while processing a request contributes to a compiled score. The evaluation runs across four deliberately independent layers.
1. Protocol-level evaluation
Examines how the client speaks TCP, TLS and HTTP rather than what it says. Cipher-suite ordering, extension set and ALPN in the TLS ClientHello, HTTP/2 settings frames and priority behaviour, and header order and casing all form a fingerprint of the underlying network stack. A client that presents a Chrome User-Agent while negotiating TLS like Python's requests or Go's net/http is exposed immediately — and fixing that mismatch is genuinely hard for the attacker, because it means reimplementing a browser's network stack.
2. Application-level evaluation
JavaScript injected into protected HTML pages runs lightweight challenges in the client environment to establish whether a real, fully-featured browser is executing the page: JavaScript execution correctness, DOM and rendering behaviour, availability and consistency of browser APIs, the client's capability to handle complex or costly operations, and evidence of headless or instrumented automation frameworks. Discrepancies between claimed identity and actual capability are strong signals.
3. User behaviour evaluation
Analyses interaction over time: mouse movement, scroll dynamics, touch events, keystroke cadence, dwell time and navigation sequencing. Humans browse irregularly, arrive from search and internal links, and pause. Harvesters traverse category listings in perfect order at constant intervals, never scroll, and never move a pointer. This layer is the one that survives when an attacker perfectly spoofs the network stack.
4. Browser fingerprint evaluation
Builds a fingerprint from device and browser characteristics — canvas/WebGL rendering behaviour, fonts, screen and viewport metrics, timezone, language, hardware concurrency — and looks for internal inconsistency (a mobile User-Agent with a desktop viewport and no touch support) and for reuse of a single fingerprint across implausibly many sessions or IP addresses.
The four layers are deliberately independent. Defeating one is achievable; defeating all four simultaneously, continuously, at scale, is expensive — which is the actual objective. Anti-scraping is an economics exercise, not an absolutes exercise.
Bot score and response segments
Every anomaly contributes to a bot score between 0 and 100. A score closer to 0 means the traffic is more likely human; closer to 100 means more likely bot. By default the score range is divided into three response segments:
| Segment | Score | Traffic character | Typical action |
|---|---|---|---|
| Cautious | 1–39 | Likely human — “low-risk” traffic | Allow, monitor |
| Strict | 40–69 | Mixture of humans and bots — “medium-risk” | Monitor, or a low-friction challenge |
| Aggressive | 70–100 | Likely bot — “high-risk” | Deny, tarpit, serve alternate content |
You set your own response strategy for the Aggressive segment so that the action is appropriate for the score range you are confident about. Deny and other hard mitigations belong there first.
Segment editability. Not every account can adjust all three response segments. If yours is limited to the Aggressive segment today, your Akamai account team will confirm the path to full control of all segments.
Choosing an action per segment
| Action | Behaviour | When to use it |
|---|---|---|
| Monitor | Score recorded, request passes untouched | Always the first setting for any new segment |
| Challenge (JS / cryptographic) | Client must complete work before proceeding | Medium-risk traffic where a false positive must not block a customer |
| Serve alternate content | Cached, stale, or reduced-fidelity response | Price and inventory pages — the scraper gets data that is useless to them |
| Tarpit / delay | Response deliberately slowed | Destroys the economics of large crawls without a visible block page |
| Deny | 403 or custom deny action | High-confidence Aggressive segment traffic only |
Hard-blocking every detection teaches the operator exactly which of their techniques failed, and they re-tool within days. Mixing delay, alternate content and denial keeps the feedback signal ambiguous and the attacker's cost high.
Cookies and detection state
Content Protector's detection is stateful: cookies carry the relevant detection state between requests so the edge can correlate a session's protocol, behavioural and fingerprint evidence rather than judging each request in isolation. All detection logic and workflows execute in the Akamai cloud; the client-side JavaScript only collects signals.
- Cookies are set on the protected domain and must not be stripped, rewritten or blocked by an intermediate proxy, cookie-consent layer or CDN in front of Akamai.
- If your consent banner blocks non-essential cookies before acceptance, detection quality degrades for pre-consent traffic — classify these cookies correctly as strictly necessary security cookies.
- Sessions that cannot persist cookies (a common scraper property) become a signal in their own right.
Scoping guidelines
Why scoping matters
Content Protector defends websites against scrapers targeting both HTML and API resources, and detection accuracy, performance and user experience all depend heavily on proper integration. Three principles drive the whole exercise:
- Balance HTML and API coverage. Protect both HTML pages and the Ajax/API endpoints behind them. Bots frequently target APIs directly, while legitimate users reach the same APIs through the HTML page (a product page calling an endpoint for price or inventory). Protecting only one side simply redirects the scraper to the other.
- Define content types explicitly. Label each resource as HTML or Ajax in configuration. Relying on auto-detection from HTTP headers causes false positives and false negatives, because headers are frequently inaccurate or deliberately spoofed.
- Protect critical and commonly scraped content only. Avoid full-site protection unless it is genuinely required. Stateful detection adds latency; a focused scope improves accuracy and minimises page-load impact.
Understanding your site type
Before defining scope, understand the structure of the site or application. Use browser developer tools to inspect network requests and see whether interactions load HTML documents or JSON payloads.
| Type | Structure | Scoping implication |
|---|---|---|
| Web 1.0 | Each action loads a new HTML page | Scope is mostly HTML rules on page paths |
| SPA (Web 2.0) | Initial HTML load, then mostly Ajax/API calls refresh content | One HTML rule for the app shell plus careful, selective Ajax rules |
| Hybrid | Most interactions load HTML, some sections use API calls for dynamic content. Common for e-commerce. | Both rule types; the dynamic price/stock endpoints matter most |
Do not scope only around your current pain points. Once protection and mitigation are in place, bot activity shifts to alternate resources carrying the same information. Perform a complete assessment of everywhere product descriptions, prices and inventory can be obtained.
Discovering critical resources
- Site coverage. Recommend broad coverage of HTML pages, but a selective approach to APIs. Protect API calls carrying product and dynamic information — inventory, pricing, availability, search results, reviews. Do not protect calls that merely assemble page structure or return boilerplate such as headers, footers, navigation menus, translation bundles or static config.
- Pre-integration discovery. Identify the high-value content first (product details, prices, inventory, listings, profiles). Then use developer tools to trace which requests deliver it and whether the transport is an HTML document or an Ajax call.
- Watch for HTML snippets over Ajax. Some sites return HTML fragments from Ajax endpoints — typically without
<html>or<body>tags. Even though the payload is HTML, treat it as an Ajax resource: the Content Protector JavaScript cannot be injected into a fragment.
Rule definition
Discovery leaves you with a list of URLs. The next step is finding patterns that identify those URLs using the match conditions available in the configuration.
Identifying HTML pages
- Common path components —
/c/for category pages,/pd/or/p/for product detail pages. - Common filenames or extensions at the end of the URL —
product.html. - The absence of an extension, typical of SEO-friendly URLs.
Identifying Ajax / API requests
- Keywords such as
api,ajax,graphqlor a version segment (/v2/) in the path. - A
.jsonfile extension. - Known service hostnames used only by the front end.
Example scope for an e-commerce site HTML rules /c/* category listing pages /pd/* product detail pages /search* search results /store-locator* location data Ajax rules /api/*/price* pricing service /api/*/inventory* stock levels /graphql (match on operation name where possible) *.json under /catalog/ Explicitly NOT protected /api/nav, /api/i18n boilerplate structure /static/*, /assets/* images, CSS, JS bundles /health, /robots.txt infrastructure endpoints
Setting up Content Protector
- Create a Content Protector configuration and associate it with the security configuration and policy protecting the hostnames in question.
- Add the hostnames in scope, confirming they are already delivered through Akamai and covered by the relevant security policy.
- Define HTML resource rules for the pages that should receive the injected JavaScript.
- Define Ajax/API resource rules for the dynamic endpoints identified during discovery, explicitly typed as Ajax.
- Set every response segment to Monitor and activate to staging, then production.
- Observe for at least one full weekly cycle, covering weekday and weekend traffic, promotions and any batch/partner jobs.
- Allow-list your own automation — synthetic monitoring, load tests, SEO auditors, partner feeds, price-comparison partners you have a commercial relationship with.
- Enable mitigation on the Aggressive segment, then reassess before touching Strict.
- Establish a review cadence and a support feedback path for suspected false positives.
Validate JavaScript injection on a representative sample of pages before enforcing anything. A page whose Content Security Policy blocks the injected script produces no client-side signals, so genuine users on that page can be scored on protocol evidence alone.
Native app traffic protection
Mobile and native applications cannot execute injected page JavaScript, so the equivalent signal collection is provided by the Native App Traffic Protection SDK for iOS and Android.
- The SDK collects device, sensor and interaction telemetry and attaches a signed payload to outbound requests, giving the edge the same class of evidence a browser provides.
- Without it, native app traffic is judged on protocol and request characteristics only — frequently indistinguishable from a scripted client, which risks false positives against your own app.
- Plan SDK adoption alongside your app release cycle: enforcement should not be enabled until the instrumented version has reached a large majority of your installed base, because older versions cannot produce the telemetry.
- Combine with certificate pinning and API-key hygiene — an attacker who extracts the app's API contract will otherwise replay it from a plain HTTP client.
Bot reports and Content Protector data
Reporting is what turns detection into a decision. Expect to work with:
- Traffic composition — humans, verified good bots, and each response segment, over time and per hostname.
- Score distribution — the shape of the 0–100 histogram. A healthy site is bimodal: a large human cluster low down and a distinct cluster high up. A flat middle means your scope or content typing needs work.
- Top offenders — by IP, ASN, fingerprint cluster and user agent, to distinguish a single aggressive crawler from a distributed botnet.
- Per-resource breakdown — which protected paths attract automation, which is the strongest indicator of what your data is worth to someone else.
- Action effectiveness — denied and challenged volumes, and whether they fall while origin traffic to the same pages stays flat. That pattern is evasion, not success.
Bot detection methods and rule IDs
Each detection carries a rule ID identifying the method that fired. Working with rule IDs rather than aggregate counts is essential for tuning:
- Detections group by category — protocol anomaly, browser impersonation, automation framework, behavioural anomaly, fingerprint reuse, and verified-bot categories such as search-engine and social crawlers.
- A specific rule ID producing false positives can be excepted for a specific path or client population without weakening the rest of the detection set.
- Track which rule IDs fire against your top offenders over time. A shift from “protocol anomaly” to “behavioural anomaly” means the operator upgraded from a scripted client to a driven browser — and your response strategy should follow.
Forwarding bot results to origin
You can forward the bot classification to your origin as request headers, so application and business systems can act on it without duplicating detection:
- Analytics hygiene — exclude bot sessions from conversion and abandonment metrics so merchandising decisions run on human data.
- Application logic — skip expensive personalisation, recommendation calls or inventory reservations for high-score sessions.
- Fraud and risk — feed the score into downstream scoring alongside payment and account signals.
- Soft enforcement in-app — the origin can degrade gracefully (hide stock counts, omit exact pricing) rather than block, which is often commercially safer than a 403.
Forwarded headers must be stripped from any request arriving at the origin from outside Akamai, otherwise an attacker can simply set the header themselves and declare their own traffic human.
Forwarding bot status to your analytics tool
The same classification can be exposed to a client-side analytics platform so bot sessions are segmented out at source. Treat this as reporting only — anything evaluated in the browser is visible to and modifiable by the client, so it must never be a security control.
Managing user access
Access to Content Protector is governed by Akamai Control Center roles and permissions:
- Grant configuration-editing rights narrowly — scope and response strategy directly affect production traffic.
- Read-only reporting access is appropriate for analytics, e-commerce and SEO stakeholders who need the data but should not change enforcement.
- Activation to production should be a separately controlled step, ideally reviewed by whoever owns the affected hostnames.
Bot product comparison
| Product | Primary problem | Best suited to |
|---|---|---|
| Content Protector | Scraping and data harvesting of HTML and API content | Retail, travel, classifieds, media, marketplaces — anywhere the content itself is the asset |
| Bot Manager | Broad bot management across the full bot taxonomy, with an extensive known-bot directory | Organisations needing category-level policy over all automation, good and bad |
| Account Protector | Account abuse: credential stuffing, ATO, fake accounts | Any site with logins and stored value |
| App & API Protector | WAF, DDoS and the platform these controls attach to | The baseline; the bot products layer on top |
Content Protector and Bot Manager solve overlapping but distinct problems. Content Protector's detection set is purpose-built for scrapers and classifies on first request; Bot Manager brings breadth of known-bot identification and category policy. Sites with both scraping pressure and a wide partner/automation ecosystem commonly run them together.
AI crawlers and content ingestion
AI crawlers are a distinct category: some identify themselves honestly and respect robots.txt, others do neither. The policy question is commercial, not just technical — some publishers want visibility in AI answers, others want licensing, others want exclusion.
- Separate declared AI crawlers (identifiable, verifiable) from undeclared ingestion traffic that mimics a browser. The first is a policy decision; the second is scraping.
robots.txtis a request, not a control. Enforcement has to happen at the edge for it to mean anything.- Consider differentiated treatment: allow crawlers that drive attribution and referral traffic, deny or meter those that do not.
- Emerging signed-agent standards let a legitimate agent prove its identity cryptographically, which is a far better basis for policy than a User-Agent string.
Tuning workflow
- Start in monitor for at least a fortnight. You need a baseline that covers weekday, weekend and promotional traffic before you enforce anything.
- Inventory your own automation first. Synthetic monitoring, load tests, SEO tooling, partner feeds and internal jobs will otherwise be your first false positives.
- Verify JavaScript injection and cookie persistence on a representative page sample, including pages behind consent banners.
- Check the score histogram. A large undifferentiated middle usually means content typing is wrong, or protected pages are not receiving the script.
- Scope protection to the valuable paths. Site-wide challenges add user friction and latency for no benefit.
- Enable the Aggressive segment first, and only tighten Strict once the false-positive rate on real sessions is measured rather than assumed.
- Expect adaptation. Sophisticated operators re-tool within days. Review weekly: are detections dropping while origin traffic to product pages stays flat? That is evasion, not success.
- Watch SEO. Verify that Googlebot, Bingbot and legitimate social unfurlers are validated and allowed — check server logs and Search Console, not just the security console.
- Close the loop. Customer support tickets about blocked access are your ground truth on false positives; build that path before you enforce.
Developer tools
| Interface | Use |
|---|---|
| Application Security API | Manage security configurations, policies, activations and the surrounding protections programmatically |
| Bot Manager API | Manage bot detections, categories, response actions and custom bot definitions |
| Terraform provider | Declare Content Protector scope, rules and response strategy as versioned code, reviewed like any other change |
| Native App Traffic Protection SDK | iOS/Android telemetry collection for native app traffic |
Manage scope and response strategy as code. Rules drift quickly as sites are re-platformed, and a Terraform-managed configuration makes it obvious when a new checkout flow or API version has escaped protection.