🎉 All Proxy Prices Reduced — Save Up To36%newResidential Lite Proxies$0.50/GB

Crawl entire websites into LLM-ready data with a single API request

Turn any website into clean, structured data — Markdown, HTML, JSON, Links, and PDF — through a proxy-powered crawl engine. Built for developers and AI teams building RAG pipelines, market research, and data products at scale.

  • 99.8% success rate
  • 10M+ pages processed
  • Markdown/JSON/HTML/PDF output
Free
Crawl // Site to data

From one URL to a complete, LLM-ready dataset.

Crawl starts with a target URL, discovers reachable pages, applies rendering and proxy rules, then returns normalized outputs for every page in the job.

  • Scalable
  • Controlled
  • Structured

Site Discovery

Find every crawlable URL from sitemaps, page links, and custom subpage rules. No page left behind.

Recursive Crawl

Control max pages, depth, and exclude paths before a crawl begins.

JavaScript Rendering

Render dynamic, JS-heavy pages in a real browser before extracting content.

Crawl API

Generate ready-to-use code from the Playground and run crawl jobs from your own app.

Data Output

Review crawl records, inspect results, and export clean data in one click.

Team Isolation

Keep crawl records, subscription credits, and usage isolated by account or team.

Proxy-powered Access

Use residential, datacenter, or custom proxies to reach pages more reliably.

Flexible Formats

Return Markdown, HTML, JSON, Links, and PDF from the same crawl workflow.

How Crawl works

Designed for developers and data teams who need repeatable, full-site data collection — without maintaining crawler infrastructure.

01

Submit a target URL

Start with a website, documentation hub, or product category.

02

Set crawl boundaries

Define depth, max pages, include/exclude rules, and request options.

03

Render and route

Enable JavaScript rendering and route traffic through Nstproxy or a custom proxy.

04

Export page data

Receive clean, structured outputs per page, ready for your pipeline.

Use Crawl visually or through API.

Start from the Playground to test parameters, then copy an API example that matches the same crawl configuration.

Playground first

Try URLs, JS rendering, proxy options, and output formats without writing code.

Code-ready examples

Copy Node.js, cURL, or SDK snippets generated from current settings.

Usage-aware

Crawl uses subscription credits first, then recharge credits when needed.

View Examples
import os
from nstdata_ai_crawl import CrawlRequestDto, Format, NstDataClient

with NstDataClient(os.environ["NSTDATA_API_TOKEN"]) as client:
    crawl = client.submit_crawl_task(CrawlRequestDto(
        url="https://docs.example.com/",
        maxDepth=3,
        maxPages=10,
        formats=[Format.MARKDOWN],
        onlyMainContent=True,
        timeout=60000,
    ))

    print("crawl task id:", crawl.id)

Discover an Easy and Cost-Effective Proxy Infrastructure

Pay per use

Free

$0

$1.80 / GB Residential
$1.20 / 1k URLs Crawl
No included credits

Small Projects

Starter

$79/ mo

$1.60 / GB Residential
$1.00 / 1k URLs Crawl
$79 included credits
Recommended

Growth

$249/ mo

$1.30 / GB Residential
$0.80 / 1k URLs Crawl
$0.15 / GB Proxy Manager
$249 included credits

High Volume

Scale

$699/ mo

$1.10 / GB Residential
$0.60 / 1k URLs Crawl
$0.10 / GB Proxy Manager
$699 included credits

Built for full-site data workflows

Use Crawl when one URL needs to become many pages of clean, structured, reusable data.

Documentation to knowledge base

Crawl docs sites and convert every page into clean Markdown for AI and support workflows.

Product catalog collection

Collect category pages, product details, pricing, availability, and links from commerce sites.

Research and monitoring

Monitor public websites and turn raw pages into repeatable research inputs.

SEO and content audits

Crawl domains to collect metadata, internal links, page content, and crawl coverage.

AI data pipelines

Feed crawled Markdown and JSON into RAG, agents, indexing, and enrichment workflows.

Page archives

Capture HTML and PDF outputs for review, records, compliance, and repeatable audits.

Frequently asked questions

Everything you need to know about Crawl, pricing, and technical integration.

01

What is Crawl?

Crawl is Nstproxy's product for turning entire websites into clean, LLM-ready data with a single API request. It discovers pages, renders JavaScript when needed, and returns structured Markdown, JSON, HTML, Links, or PDF.

02

How much does website crawling cost per page?

03

What's the best Firecrawl alternative?

04

Can Crawl render JavaScript-heavy websites?

05

Is Crawl good for building a RAG knowledge base?

06

Can I control which pages are crawled?

07

Do I pay for failed or blocked requests?

08

How is Crawl usage billed?

Nstproxy

Crawl

Turn any website into clean, LLM-ready data with Nstproxy's proxy-powered crawl engine. Start free — no monthly minimum, pay only for what you crawl.

Nstproxy logo©2026 NST LABS TECH LTD. All RIGHTS RESERVED.