🇯🇵 Tokyo is live! 🚀 Launch your VPS and enjoy 2 months off — use code KONNICHIWA50 🎉 Get Started Today →

Bot-vs-Human Bandwidth: What AI Scrapers Really Cost Your VPS (and How to Block or Throttle Them)

VPS server connected to automated web traffic and infrastructure services

Automated crawlers can be useful, unwanted, or simply too aggressive for the capacity of a site. The expensive part is not the bot label in a user-agent string; it is the work each request triggers: TLS, application code, database queries, cache misses, rendered pages, and response bytes. Before blocking anything, measure which clients and paths are consuming resources.

This guide builds a reversible response for an Nginx-hosted site: establish a baseline, classify traffic from logs, publish crawler preferences, test rate limits in dry-run mode, and escalate only when the evidence supports it. It avoids fabricated benchmarks and fixed thresholds because a safe limit depends on your application, cache behavior, legitimate bursts, and shared-client traffic.

⚡ Spin up a Premium VPS in 2 minutes
17 locations worldwide
NVMe  ·  Unmetered 1 Gbps  ·  Full root access  ·  From $10/mo

Establish a Scraper-Traffic Baseline

Start with a representative window before changing robots.txt, Nginx, a firewall, or a CDN. Preserve the raw logs, note the timezone, and collect application metrics for the same period. Request count alone is not enough: one cached asset and one dynamic search page can have very different CPU, database, and transfer costs.

  • Count requests and response bytes by client address, user agent, status, and path.
  • Compare Nginx request volume with CPU, memory pressure, disk latency, upstream response time, and cache-hit data.
  • Separate successful responses from redirects, client errors, and server errors.
  • Keep at least one known normal period for comparison, including legitimate traffic bursts.

The command below summarizes requests and bytes by source address for a standard combined Nginx access log. Confirm your own log_format first; if the fields differ, adjust the parser rather than trusting incorrect columns.

sudo awk '
$10 ~ /^[0-9]+$/ {
    requests[$1]++
    bytes[$1] += $10
}
END {
    for (ip in requests)
        printf "%12d %12d %s\n", bytes[ip], requests[ip], ip
}' /var/log/nginx/access.log | sort -nr | head -30

Read the first column as response bytes recorded by Nginx, not total network usage. Headers, retransmissions, upstream traffic, and traffic terminated before Nginx may be absent. For a broader monitoring stack, see the VPS.us guide to open-source monitoring tools.

Identify Automation Without Treating User Agents as Proof

A user-agent string is a claim made by the client. It helps group requests, but any client can copy a well-known crawler name or a browser signature. Combine user agent, verified provider ranges where available, request rate, path sequence, response size, cache behavior, and repeated access patterns before classifying traffic.

Google explicitly recommends verifying Google crawler requests with its published IP ranges or forward-confirmed reverse DNS because Googlebot user agents can be spoofed. OpenAI and Anthropic also document separate crawler identities for different purposes. As of this review, OpenAI distinguishes crawler purposes, while Anthropic documents ClaudeBot, Claude-SearchBot, and Claude-User. Re-check the providers’ current documentation before editing a permanent policy.

Extract the most common user-agent strings from a combined log, then inspect the corresponding addresses and paths. Do not automatically block a client merely because the string contains curl, python, or another general-purpose tool; legitimate integrations and health checks may use them.

sudo awk -F'"' 'NF >= 6 {count[$6]++}
END {
    for (agent in count)
        printf "%10d %s\n", count[agent], agent
}' /var/log/nginx/access.log | sort -nr | head -40
sudo grep -F 'suspected-agent-fragment' /var/log/nginx/access.log | tail -50

Decide Which Crawlers and Paths You Want

Write down the policy before implementing it. You may want search discovery while opting out of model training, or you may allow public articles while denying expensive search, export, archive, or API paths. A purpose-based policy is safer than a blanket list copied from an old blog post.

  • Allow: traffic that supports search visibility, monitoring, integrations, or another documented business goal.
  • Disallow by crawler convention: compliant crawlers you do not want accessing specific paths.
  • Rate-limit: traffic that is legitimate in moderation but expensive in bursts.
  • Block: abusive behavior with strong evidence, a defined duration, and a rollback path.

Protect private content with authentication and authorization. A robots.txt rule is a crawler preference, not an access-control boundary. Likewise, a noindex directive addresses indexing; it does not make a resource private.

Publish Explicit robots.txt Preferences

OpenAI, Anthropic, and Common Crawl each document robots.txt controls for their crawlers. The example below blocks three training or corpus crawlers while leaving other agents unaffected. Confirm that this matches your policy; blocking search-specific crawlers can reduce visibility in their search experiences.

cat > /tmp/robots-crawler-policy.txt <<'EOF'
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
EOF
curl -fsS https://example.com/robots.txt

Back up the existing file and merge the candidate rules from /tmp/robots-crawler-policy.txt into it; preserve current crawler groups, path rules and sitemap entries. Replace the example domain with your own. Serve the file at the root of every hostname covered by the policy and test the public URL. Compliant crawlers can still take time to re-fetch it, and noncompliant clients may ignore it. Keep enforcement controls separate so a crawler can read the policy even when other paths are restricted.

Test Nginx Rate Limits Before Enforcing Them

Nginx’s limit_req module uses a leaky-bucket method and can key limits by client address. Start with limit_req_dry_run on: Nginx records excessive requests without limiting them, which lets you observe false positives against real traffic. Choose a rate from measured application capacity and legitimate burst behavior rather than copying the values below.

This example defines a per-address zone and applies it only to an expensive endpoint. If Nginx sits behind a proxy or load balancer, configure and verify the real-client-address module first; otherwise every visitor may appear to share the proxy address. This is a configuration fragment: merge the limit directives into the existing location that actually handles the expensive request, retaining its proxy/FastCGI or other content handler. Test internal redirects and shared-IP clients as part of the dry run. The error-log threshold must include notice events for the example observation commands to show them.

# In the http context:
limit_req_zone $binary_remote_addr zone=per_client:10m rate=5r/s;
limit_req_status 429;
limit_req_log_level notice;
error_log /var/log/nginx/error.log notice;
# In the relevant server block:
location /expensive-path/ {
    limit_req zone=per_client burst=20 nodelay;
    limit_req_dry_run on;
    # Keep your existing upstream or content-handler directives here.
}

Validate the complete configuration with sudo nginx -t, reload it, and review the error log during a normal traffic cycle. When the observed excess events match your intended policy, change only limit_req_dry_run to off, validate again, reload, and watch both 429 responses and application health. The VPS server optimization guide explains how to pair request controls with capacity and service monitoring.

Escalate Blocking Carefully

Long-lived firewall blocks are a poor first response to rotating or shared addresses. They can deny legitimate users, are easy for distributed scrapers to evade, and may become stale. Prefer scoped application or proxy controls that preserve useful logs and return an explicit response. Use a firewall only for addresses or networks you have independently verified as abusive, with an expiry and a documented owner.

Fail2Ban can turn repeated, well-defined log events into temporary bans, but its filter must match your actual Nginx log format. Test the filter against saved logs with fail2ban-regex before enabling a jail, add trusted administrative addresses to ignoreip, and confirm the jail’s backend matches the firewall stack your server already owns. Do not mix unmanaged iptables, nftables, UFW, provider-firewall, and Fail2Ban rules without understanding their order.

Before changing remote-access or firewall policy, keep the provider console available and verify your SSH path in a second session. The VPS.us server-access and hardening checklist covers that recovery-first workflow.

Verify Results and Keep a Rollback

Measure the same indicators you captured in the baseline. A successful change reduces unwanted expensive work without increasing errors for normal users or blocking required discovery. Review at least response codes, request and byte volume, upstream latency, CPU, database load, cache-hit behavior, and support reports.

  • Fetch /robots.txt from outside the server and confirm the intended host serves the current file.
  • Use a controlled client against a path you own to verify dry-run and enforced rate-limit behavior.
  • Inspect Nginx logs for the real client address, selected path, response code, and limit events.
  • Keep the previous configuration, validate it before restoration, and record who changed the policy and why.
sudo nginx -t
sudo systemctl reload nginx
curl -fsS https://example.com/robots.txt
sudo grep -E 'limiting requests|dry run' /var/log/nginx/error.log | tail -50
sudo awk '$9 == 429 {count++} END {print count + 0}' /var/log/nginx/access.log

If legitimate traffic is affected, return the location to dry-run mode or restore the reviewed prior configuration, run sudo nginx -t, reload, and verify again. Never let an incident response become an undocumented permanent rule.

Build a Sustainable Bot-Traffic Policy

Crawler control is an operating process, not a static denylist. Keep raw evidence, separate crawler purposes, protect private data with real access controls, and enforce limits at the narrowest expensive paths first. Review provider identities and product goals periodically because crawler names, published ranges, and business priorities change.

The strongest outcome is not “zero bots.” It is predictable application performance with explicit, testable rules: useful crawlers can reach what you intend, unwanted compliant crawlers see clear preferences, abusive behavior encounters measured limits, and every enforcement step has a safe rollback.

Frequently Asked Questions

Does robots.txt block every scraper?

No. It communicates preferences to compliant crawlers but does not enforce authorization. Protect private content with authentication, and use measured server, application, or network controls for abusive traffic.

Can I identify a crawler only from its user-agent string?

No. User agents are easy to spoof. Combine them with provider verification methods where available, source addresses, request patterns, paths, rates, response sizes, and application impact.

What Nginx request limit should I use?

There is no universal safe value. Derive it from measured capacity and legitimate traffic bursts, apply it to a narrow expensive path, and observe limit_req_dry_run events before enforcement.

Should I block known AI-crawler IP addresses?

Usually not as a first step. Some providers publish verification data while others do not, addresses can change, and shared or spoofed traffic complicates attribution. Prefer documented robots.txt controls for compliant crawlers and behavior-based limits for abuse.
Facebook
Twitter
LinkedIn

Table of Contents

Get started today

With VPS.US VPS Hosting you get all the features, tools

Image