workingproxysites.

API directory

Structured Web Extraction APIs: Compare Parsers & Output Contracts

Compare scraping APIs with documented structured-extraction capabilities, outputs, authentication requirements and billing qualifications.

Shortlist extraction APIs and distinguish parsed target data from a JSON response envelope.

  • Which target types or extraction rules are supported?
  • Is JSON parsed target data, or only a wrapper containing HTML?
  • Do custom fields, browser rendering and asynchronous delivery change the request contract?
7 documented listings

This collection covers our reviewed proxy and web-data catalog. It is not an exhaustive market survey or a performance ranking.

Category reviewed . Review due .

Compare documented requirements

7 of 7 listings
Clear filters

Swipe or scroll horizontally to compare requirements. Scroll back to the first column to see listing names.

API interfaces, outputs and billing · alphabetical
API / serviceInterface & outputDocumented behaviorBilling & restrictionsEvidence
Nimble Extract APINimble · official publisher

Fetch HTML, Markdown and screenshots with configurable rendering, CSS parsing, browser actions and asynchronous batches.

Details & setup →Provider discussionSourced JSON record
HTTP API

HTML, JSON, markdown, PNG

Scraping API

For this task: A CSS parsing schema defines extracted fields separately from the formats array. Match selectors to the target page and choose rendering independently; requesting a JSON response alone does not define a field schema.[1][4]

Formats
The formats array selects HTML, Markdown, Base64 PNG screenshots and response headers.[1]
Rendering
render defaults to false; true enables JavaScript and auto lets Nimble select the configuration.[1]
Parsing and actions
CSS parsing schemas, browser actions and network-capture rules are supported.[1]
Async processing
Async and batch submissions return before processing finishes; poll task or batch progress for completion.[2]
Billing
Extraction Tools are listed at USD 1 per 1,000 URLs, with charges only for successful requests.[3]

Successful URL extractions; public API pricing lists USD 1 per 1,000 URLs for Extraction Tools.

Requirements & limitations
  • The Extraction Tools price is distinct from paid Extraction Templates, Search and Web Search Agent products.
  • Exact concurrency limits and all driver-specific constraints were not verified in this pilot.
Documentation checkedReview due
4 source references
Oxylabs Web Scraper APIOxylabs · official publisher

Fetch public pages or use dedicated target parsers, with rendered HTML and asynchronous batch delivery.

Details & setup →Provider discussionSourced JSON record
HTTP API

HTML, JSON, PNG

Scraping API

For this task: Dedicated supported targets accept parse=true for parsed fields. A universal page fetch does not establish automatic parsing for every site; check the supported source and target before choosing a schema.[1]

Authentication
Use a Web Scraper API user's username and password with HTTP Basic authentication, rather than dashboard credentials.[1]
Rendering and parsing
render=html enables JavaScript rendering. Dedicated supported targets accept parse=true for structured JSON.[1]
Delivery choices
Realtime keeps the connection open; Push-Pull retrieves jobs asynchronously; Proxy Endpoint offers a proxy interface.[2]
Batch size
Push-Pull accepts up to 5,000 query or URL values per batch; Realtime and Proxy Endpoint do not support batch.[2]
Screenshots
render=png returns a Base64-encoded screenshot.[2]

Product-specific plan and target rates; exact costs and failed-request rules were not verified in this pilot.

Requirements & limitations
  • Dedicated parsers apply to supported targets, not every arbitrary URL.
  • Proxy Endpoint accepts fewer additional parameters than the HTTP integrations.
  • Exact plan pricing and failed-request billing remain unverified.
Documentation checkedReview due
2 source references
ScrapeOps Proxy API AggregatorScrapeOps · official publisher

Route scraping requests across upstream services, with rendering, proxy-pool and extraction options.

Details & setup →Provider discussionSourced JSON record
HTTP API

HTML, JSON, markdown, screenshot

Scraping API

For this task: The LLM extraction path supports JSON output and page-type schema hints. llm_extract adds 25 credits to the underlying fetch cost; this is separate from receiving an ordinary JSON response wrapper.[4][1]

Billing outcomes
The credit policy counts both HTTP 200 and 404 responses as successful, billable requests.[1]
Rendering costs
Standard rendering costs 10 credits; residential plus rendering costs 25; render_js_cheap costs 5 using a limited provider set.[1]
Extraction add-on
llm_extract adds 25 credits on top of base proxy costs.[1]
Cheaper rendering tradeoff
render_js_cheap is intended for simpler sites and does not provide the full upstream provider pool.[2]
Response options
The feature list documents screenshot capture and JSON or Markdown LLM extraction responses.[3]

Monthly API credits determined by target and options; HTTP 200 and 404 responses consume credits.

Requirements & limitations
  • Certain domains and premium configurations use different credit schedules.
  • Unused credits do not transfer into the next subscription period.
  • A target 404 can still consume quota.
Documentation checkedReview due
4 source references
ScraperAPIScraperAPI · official publisher

A scraping HTTP API with optional rendering, premium proxy pools and asynchronous requests.

Details & setup →Provider discussionSourced JSON record
HTTP API

HTML, JSON

Scraping API

For this task: autoparse=true and dedicated structured-data endpoints work with supported sites and target types. An ordinary URL fetch does not promise a general parser for arbitrary websites.[1][4]

Rendering
render=true returns HTML after JavaScript execution; wait_for_selector requires rendering.[1]
Credit examples
The documented parameter table lists 10 credits for rendering, 25 for premium plus rendering, and 75 for ultra-premium plus rendering.[2]
Billable outcomes
HTTP 200 and 404 responses are charged, as are client-cancelled requests when the client allows less than 70 seconds.[2]
Cost inspection
The account/urlcost endpoint estimates request cost and the sa-credit-cost response header reports actual credits.[2]
Concurrency
Concurrency is plan-dependent; the Async API supports batch requests.[3]
Structured extraction
autoparse=true and dedicated structured-data endpoints parse supported sites and target types; ordinary URL fetching is not a universal field-extraction contract.[1][4]

API credits determined by target domain and request options.

Requirements & limitations
  • The listed rendering credits are configuration examples; protected domains and special targets can change total cost.
  • A target 404 can still be billed.
  • Unused subscription credits do not roll over.
Documentation checkedReview due
4 source references
Scrapfly Scrape APIScrapfly · official publisher

A scraping API with managed rendering, JavaScript scenarios and configurable anti-blocking options.

Details & setup →Provider discussionSourced JSON record
HTTP API

HTML, JSON, screenshot

Scraping API

For this task: Integrated extraction accepts extraction_model, extraction_prompt or extraction_template. These select model, prompt and template approaches with different output contracts; choose the mechanism and required fields for the target.[3]

Rendering and actions
render_js enables browser rendering; js_scenario defines page interactions before results return.[1]
Result handling
Examples return scraped content and browser data through the API result rather than exposing a browser connection.[1]
Credit model
A scrape can consume multiple credits depending on browser, proxy and protection configuration.[2]
Concurrency examples
Published monthly plans list 5 concurrent requests for Discovery, 20 for Pro, 50 for Startup and 100 for Enterprise.[2]
Quota expiry
Unused subscription quota does not roll over to the next subscription period.[2]
Integrated extraction
extraction_model, extraction_prompt and extraction_template select documented model, prompt and template extraction approaches within a scrape.[3]

API credits per scrape; proxy network, rendering and protection settings affect consumption.

Requirements & limitations
  • This API listing does not claim Puppeteer or Selenium browser control; the pricing FAQ explicitly distinguishes that capability.
  • Exact target costs and failed-request billing were not verified in this pilot.
Documentation checkedReview due
3 source references
ScrapingBee HTML APIScrapingBee · official publisher

Fetch rendered pages, apply CSS extraction rules or run browser scenarios through an HTTP API.

Details & setup →Provider discussionSourced JSON record
HTTP API

HTML, JSON, markdown, text, screenshot

Scraping API

For this task: CSS extract_rules defines the fields to extract. AI extraction rules are a separate option with additional credit requirements; choose the extraction mechanism before comparing costs.[1]

Rendering default
render_js defaults to true. HTML, page text, Markdown and screenshots are documented output options.[1]
Request credits
Classic requests cost 1 credit without rendering or 5 with it; premium with rendering costs 25 and stealth with rendering costs 75.[1]
Extraction and interaction
CSS extract_rules and js_scenario support structured extraction and page interactions.[1]
Custom proxy
The own_proxy parameter supports a user-supplied proxy provider.[1]
Automatic cost control
GET-only mode=auto chooses a successful rendering/proxy tier; max_cost caps eligible tiers and all-failed auto attempts cost zero.[1]

Monthly API credits; rendering and proxy tier determine base credits, with AI extraction add-ons.

Requirements & limitations
  • AI extraction adds credits beyond base fetch costs.
  • Auto mode does not supply page-specific waits or browser scenarios.
  • The service returns extracted content through HTTP; this record does not claim a remotely controllable CDP browser.
Documentation checkedReview due
1 source references
Zyte APIZyte · official publisher

Choose HTTP response bodies, rendered HTML, screenshots or supported structured extraction in one API.

Details & setup →Provider discussionSourced JSON record
HTTP API

HTML, JSON, screenshot

Scraping API

For this task: Choose a typed extraction output such as product or article. Returned HTML and a JSON response envelope are separate from these extracted objects; select the page type and extraction source for the workload.[3]

HTTP output
httpResponseBody=true returns the response body Base64-encoded; decode it before parsing.[1]
Browser output
Browser requests return rendered HTML, screenshots or both, and support browser actions and network capture.[2]
Structured extraction
The API supports dedicated extraction types including articles, products, job postings and SERPs.[3]
Billing
Unsuccessful and rate-limited responses are free; browser actions, screenshots and extraction can add costs.[4]
Request model
The /v1/extract endpoint blocks until its single-URL result is ready.[3]

Successful responses, priced by target and HTTP/browser tier, with feature add-ons.

Requirements & limitations
  • This listing covers the HTTP extraction API; Zyte's live CDP browser has separate session and billing rules.
  • Initial browser requests do not support arbitrary HTTP methods, bodies or custom headers beyond Referer.
  • Target tier assignments can change; use the provider's cost estimator for a workload.
Documentation checkedReview due
4 source references

How to use this comparison

Compare the documented request interface, output and account product before checking price. A JSON response can contain raw HTML; it does not by itself establish structured extraction. A browser action in one request does not by itself establish a reusable browser session.

Each row keeps the publisher’s billing basis and restrictions alongside its sources. Unknown prices or limits remain unknown. These are documentation reviews; no live connection, purchase or comparative performance test is claimed.

Overdue records remain visible with a label. This collection leaves search indexing when its category review is overdue or fewer than 3 current, visible records qualify. Individual detail pages retain their own review policy.

The collection JSON link covers the whole category before filters.

Suggest a correction · Affiliate disclosure · Collection JSON records