How Sites Tell Bots From People (And Why It’s Getting Harder to Fake)

134 Views

The idea of bots causing havoc on the world wide web predates this latest era of LLM and agentic AI. For many years now, bots have been causing problems for competitive product drops and concert ticket releases, with human users frequently left in the dust because bots have been deployed to buy-up all available [insert product here] and resell them for profit down the line. 

Sites have been grappling with queue management and bot identification for a long time, but the recent advances in AI (and its availability for just about anyone with a computer) mean that the problem has escalated. Combine that with the rise of legitimate AI agents, which are increasingly being used to shop on behalf of their users, and figuring out how to keep a website for humans is harder than ever.

Or is it? 

Browser Fingerprinting

Browsers reveal a collection of technical traits, like how they render graphics, which fonts they have installed, how the audio engine processes a signal, the precise order in which it negotiates a secure connection, to every site that is opened. Even without cookies, logins, or tracking IDs, this fingerprint means that the browser (and the device using it) reveals itself with quite a high degree of clarity. 

It’s a distinct concept from cookies, and existed independently of anti-bot measures for a long time. 

person using silver laptop computer on desk

How Browser Fingerprinting Reveals Bots

While fingerprinting was originally introduced to track human users, it turns out that it’s also highly effective at identifying non-human ‘users’, too, thanks to the signals it sends to the site. These include:

A few of the most common signals worth naming specifically:

  • Canvas and WebGL rendering
    The browser is asked to draw a hidden image or 3D shape, and subtle pixel-level differences caused by GPU, driver, and OS combinations create a near-unique output.
  • Audio fingerprinting
    A similar idea applied to how a device’s audio stack processes a synthesized waveform.
  • Font and hardware enumeration
    The specific list of installed fonts, screen dimensions, and device metadata, which in combination narrow a device to a small population very quickly.
  • TLS-level fingerprinting (JA3/JA4)
    A layer below the browser entirely, based on how a client negotiates its encrypted connection, which means fingerprinting isn’t purely a JavaScript-visible signal anymore.

Some properties contained within a fingerprint have no real reason to be present unless an automation is present, while others are more general-purpose fingerprinting traits that only become meaningful once cross-referenced against everything else the browser is reporting.

This means it’s possible to risk score each user that passes through the metaphorical door. 

Why Headless Environments are More Likely to Fail This Test

A headless browser operates without a graphical user interface, which means it’s operated purely on code commands and doesn’t have that GPU rendering pipeline generating the canvas and WebGL output a genuine device would produce. 

What’s more, headless execution frequently produces memory footprints and process signatures (and even timing patterns) that differ meaningfully from a browser running twith a visible UI attached.