If you keep your own visitor log (and a business site should: it is the only record of your traffic that nobody else can change), you will soon meet the question that hosted analytics quietly answer for you: which of these visits were people? The old answer was to look at the user-agent string, the label a browser sends naming itself. Crawlers used to say "bot" in it. Many still do. The ones that matter no longer do.
What the log showed
When we reworked the log on a small-business site this week, we took a sample of one day. It counted 94 page views as people, from 22 visitors. Two thirds of those views were the owner testing pages. Among the rest were addresses in Google Cloud and Amazon Web Services running current versions of Chrome on Linux, an address belonging to a security vendor's scanner running a Chrome from two years ago, a web-hosting company and a content delivery network. Every one of them announced itself as an ordinary browser. Every one of them had been counted as a customer.
The pixel made it worse. A visitor log usually has a fallback for browsers that do not run scripts: a one-pixel image whose request is recorded. Those loads had been counted as people too. Almost none of them were.
Judge the visit by what it did
The fix was not a better list of bot names. It was to stop deciding at the moment of writing and decide at the moment of reading, with everything the visit left behind. A visit is now classified when a report is drawn, by rules that can improve without touching the stored records. In order:
- A request that never ran the page's script (the pixel) is its own category, "no script ran". Not a person, not a bot; reported separately.
- A visit that says it is automated (a bot user agent, or a browser that reports it is being driven by automation software) is a bot.
- A visit from a datacenter, hosting or security-vendor network is a bot, unless the visit shows a sign of a person.
- A browser many major versions behind the newest of its family seen in the same report is a bot, again unless it shows a sign of a person. Scanners run whatever browser they were built with; people update.
- Everything else is a person.
The phrase doing the work is "a sign of a person". The log records how long a page was active and whether the visitor scrolled or clicked. A visit with at least three seconds of active time and any scroll or click counts as a person whatever network it came from, because people use VPNs that exit in datacenters, and a crawler does not read the page. The rule errs toward people once there is evidence of one.
Two exceptions worth knowing
Apple's iCloud Private Relay routes Safari on iPhones, iPads and Macs through Akamai, Cloudflare and Fastly. A naive datacenter rule would mark every one of those visitors as a bot. Safari from one of those networks is treated as a person.
The second exception is you. The owner and staff generate a surprising share of a small site's traffic: checking a page after a change, opening the admin console, testing a form. Those visits are real, but they are not customers. We tag the browser when it signs in to the admin surface, and the log also remembers devices seen there, so staff visits become a fourth category and every "people" figure excludes them. Nothing about this needs a person's name; a first name from the signed-in account is enough to label the row.
What changed in the numbers
On a sample of the most recent four thousand records from the same site, public page views by people fell from 399 to 91 once staff and datacenter browsers were separated out; 308 views were staff and 449 were bots. The report did not lose anything. It stopped counting a scanner's nightly visit as a prospect and the owner's own checks as demand, which is what the numbers had quietly been.
Why do this yourself?
Hosted analytics products make similar judgements, invisibly, and they are usually right. The reasons to keep your own log alongside them are that it is yours (no consent banner from a third party, no sampling, no retention limit you did not choose), that it can be joined to your own records (which visit became which inquiry), and that when a number looks wrong you can open the rows and see why. The classification described here is about a hundred lines of plain JavaScript with a test suite. Its value is not the code; it is that the rules are written down where the business can read them, and the report says which rule it applied to each visit.
If you keep a log, check three things this week
- Do no-script loads (a tracking pixel, if you have one) count as people? They should be their own line.
- Filter one day's "people" by network owner. If Google Cloud, AWS, Microsoft Azure or a security vendor appears, your bot rule is a user-agent list and it is out of date.
- Subtract yourself. Compare a week of traffic with and without the devices you and your staff use. The difference is the size of the story your reports have been telling you about your own habits.
This is part of the first-party analytics work we do inside website and platform projects; the PC NET TECHS project shows the reporting it feeds. If your reports have never separated people, bots and staff, ask us what it would take on your site.
