Over the past several months we have watched a clear shift in the traffic reaching the sites and portals we work with. A growing share of it is not coming from people. It’s coming from bots.

At first glance, an uptick in activity can look like growing interest. But a closer look reveals much of it is automated: crawlers, scrapers, and AI agents moving across the web at a scale that now outpaces human visitors. Link Digital believes it is important to be open with its clients about what is happening, why it is happening, and what it means for the portals your operations rely on.

What is happening?

In 2025, automated traffic on the web grew far faster than human activity, as more people turned to AI assistants for everyday questions. By 2026, bots had, for the first time, become the majority of web requests. The web was built on the assumption that a person sits on the other side of the screen. That assumption is increasingly not the case.

Why it is happening

Most of this traffic relates to AI activity, and it falls into three patterns:

  • Training crawlers that gather content to build and update AI models
  • Real-time fetchers that pull a page when someone asks an AI assistant a question
  • Agentic browsers that navigate sites on a user’s behalf
Infographic: Automated traffic is outpacing humans, with AI daily requests on the left and bots’ share of web traffic on the right.

A lot of the traffic comes from a small number of operators. According to HUMAN Security’s 2026 State of AI Traffic report, OpenAI’s bots account for roughly 69% of observed AI-driven traffic by volume, Meta around 16%, and Anthropic around 11%. These figures vary by measurement method and period, so treat them as directional, but the pattern of concentration is consistent across all the sources.

One aspect of this that is worth stating plainly is that the heaviest AI traffic lands on retail, media, and travel sites, which together absorb more than 95% of it. Government, research, and open-data portals may see a different pattern, which is why your own server logs, not industry averages, are the most reliable measure of your exposure.

Why this matters for portal owners

There are two practical consequences.

First, measurement. Standard analytics tools often cannot tell an AI agent from a human user, so raw visit counts are no longer a clean signal of genuine human interest. Engagement depth and authenticated actions are more dependable signals.

Second, infrastructure. Bots consume real resources: bandwidth, compute, and server capacity. AI bots drive increased scraping, operational load, and high-frequency access, and at scale that can translate into infrastructure cost, particularly for organisations running a public data portal.

What you can do about it

The good news is that this is a manageable problem. Much of this traffic is legitimate, and, for open data, research, and government portals, being discoverable is usually the point. The goal is to let the useful visitors in, slow down the heavy ones, and keep out the genuinely abusive ones, without losing visibility or overspending on capacity. In practice, this means:

  • Measure what matters. Lean on engagement depth, authenticated actions, and qualified enquiries rather than raw visit counts, and review your server logs to see which bots are actually hitting you.
  • Decide who you let in. You can set the rules for crawlers with a small file on your site called robots.txt, which tells them what they can and cannot access. The catch is that it works on trust: well-behaved crawlers, like search engines, follow those rules, but abusive ones ignore them. So robots.txt takes care of the polite bots, while the heavy crawlers can be slowed down with rate-limiting (a cap on how often they hit your server), and the genuinely abusive ones can be blocked outright. Done this way, your portal stays discoverable without the load getting out of hand.
  • Protect the infrastructure. A few standard tools can take most of the strain off your server. Caching and a CDN (a content delivery network) keep copies of your pages so repeat requests do not hit your main server every time. Server-side rate-limiting and a web application firewall (WAF) give you enforceable control that bots cannot simply ignore, unlike robots.txt. And bot-management tooling spots and handles automated traffic at scale.
Bot traffic readiness checklist

Free download: Grab our Bot Traffic Readiness Checklist, a one-page guide portal owners can work through step by step.

Planning for capacity and cost

This is the part of the issue Link Digital wants to be candid about. 

As automated traffic grows, some portals will need more infrastructure headroom to stay fast and reliable, and that can mean a future conversation about capacity and budget. We want to flag it earlier rather than later. Right now, Link Digital is not introducing new charges related to what we are talking about in this article. Instead, we are working through how best to manage the additional costs, resources, and expertise involved so that any infrastructure decisions are planned together with clients, based on evidence from your own traffic, rather than being sprung on you unilaterally.

How we can help

If you would rather not manage any of these tasks in-house, that is exactly what our managed hosting service is for. We handle the caching, rate limiting, monitoring, and bot management on your behalf and keep you informed on how your portal traffic is trending. You stay focused on your data and your users; we keep the infrastructure steady underneath. 

If you would like a quick review of what bot traffic currently looks like on your portal, we are happy to look at your logs and talk you through it. Contact us here if you would like to discuss your situation and what might be involved.

Sources: