How Evaluat Uses Real Browser Load Testing to Solve Modern Web Performance Challenges
Founder of Evaluat, a real-browser performance testing platform. I spent years building and load-testing Magento storefronts, and got tired of load tests that passed while real customers struggled at peak.
By Ahmad Farzan
· Spent years building and load-testing high-traffic Adobe Commerce and Magento storefronts.
· London, United Kingdom
Published August 25, 2026 · 6 min read
Published August 25, 2026 · 6 min read
This case study is based on responses submitted directly by the founder or member of the team from Evaluat. They have verified ownership of their domain evaluat.com on SaaS Browser.
How Evaluat got started
I spent years building and load-testing high-traffic Adobe Commerce and Magento storefronts, and the moment that tipped me over came during a peak sale on one of them. We had load-tested beforehand and everything was green. Server response times looked fine, error rates were low, the report said we were ready. Then the sale started and customers were telling us checkout was crawling. Pages took forever to become usable, and people were abandoning baskets while our dashboards showed a healthy server.
The tools weren't lying, they just weren't measuring the right thing. Protocol-level load tests fire HTTP requests and time the responses. They never render a page, so they never see the JavaScript, the third-party tags, or the layout shifts, which is where the customer experience actually lives. The server was fine. The experience was broken. And when I went looking for evidence of what an individual shopper had gone through, there was nothing to look at. Just averages and percentiles that buried the failures.
I wanted to run every virtual user in a real, isolated browser and keep the evidence, Core Web Vitals, a video of each session, network and console logs per user. Nothing did that in a way I could put in front of a retailer, so I started building it myself. That became Evaluat.
Growing Evaluat: what worked and what didn't
Worked, writing the comparison post nobody honest had written. I published a roundup of the best real-browser load testing tools that included our direct competitors, with straight pros and cons and clear notes on which famous names don't actually run real browsers. It hit number one on Google in the US within three days of going live, and it became the first page an AI assistant cited us on. Buyers at the comparison stage are exactly the people we want to meet, and being honest about competitors is still rare enough that search engines and answer engines both seem to reward it.
Flopped, expecting anyone to find us by name. We're called Evaluat, without the final e. Google autocorrects the search to 'evaluate' and serves dictionary definitions, so even people who knew the name and typed it correctly landed on a dictionary page. Effort spent on brand-first visibility did nothing measurable. The lesson was that at our size, distribution comes from owning the specific questions buyers type, not from brand searches or the big head term, which is guarded by Wikipedia and a wall of very high-authority sites. We rank for the narrow questions and let the brand catch up later.
What Evaluat customers really think
The two complaints I hear most are regions and cost versus protocol tools.
Regions first. We generate load from the UK and EU, and a test's data stays in the region it ran in, which our European customers like. But anyone with a big audience elsewhere asks when we'll run tests closer to their users, and the honest answer today is a roadmap conversation. I handle it by being straight, the region list will grow, here's the residency model any new region will follow, and I'd rather say that plainly than pretend coverage we don't have.
Second, price per virtual user. Next to a protocol tool that simulates thousands of connections, running hundreds of real browsers looks expensive, and prospects say so. A real browser is a much heavier thing to run than an HTTP loop, but the way through isn't arguing about compute, it's making the value visible. Every session ships with its own video, network log, and console output, so when a test catches a failure, you can watch the exact user who hit it and hand your engineer the evidence. Once someone has seen that in a report, the price conversation usually goes quiet.
What most people get wrong about Load & Performance Testing
Most people think load testing means firing HTTP requests at a server and reading response times, and that if the server answers quickly under load, users are fine. That misses where modern sites actually fail. A server can return every response in 300 milliseconds while the page itself is unusable because rendering, JavaScript, third-party tags, and layout shifts all happen after the response arrives, in the browser, and a protocol test never opens a browser.
The second layer of the misconception is that people assume this is already solved. When I mapped 25 tools that show up in load testing shortlists, only around six actually run each virtual user in a real browser. Several famous names in the category simulate protocol traffic and offer a browser mode as a garnish— a handful of real browsers running alongside thousands of simulated connections. Buyers read "browser-based" on a pricing page and assume they're getting real-user conditions. Mostly, they aren't.
That is the gap Evaluat sits in. Every virtual user is a real, isolated browser, and the test measures what those browsers experienced under load, Core Web Vitals included, not just what the server sent back.
What's next for Evaluat
Two things. First, we're launching Pulse, a free website speed test that loads your page in a real browser and gives you Core Web Vitals, a composite grade, and a video of the load. It's built and in final testing now. It's the top of our funnel, and it's useful on its own even if you never buy anything.
Second, more test regions. We run from the UK and EU today, with test data staying in the region it ran in, and the region list will grow toward wherever our customers' audiences are. Underneath both, we keep building for a world where AI agents do the evaluating. The platform already exposes an MCP server, so an agent can run a speed test against a site directly.
Ahmad's background
I wasn't starting from scratch on the problem, but I was on the product. I'd spent years building and load-testing high-traffic Adobe Commerce and Magento storefronts, so I'd lived the pain from the practitioner side. I was the person running the tests, reading the reports, and then explaining to a retailer why the site still fell over on sale day. I knew exactly what the reports weren't telling me, and I knew the gap wasn't going to be closed by yet another protocol tool with prettier charts.
What I hadn't done was build and sell a SaaS product. Running a platform, pricing it, doing sales-led go-to-market, positioning against established names - all of that I've had to learn on the job, and some of it the hard way.
Biggest lesson building Evaluat
Positioning. I launched talking like the incumbents, describing Evaluat as a load testing platform, because that's the category buyers know. In a market where the head term is held by Wikipedia and a wall of very high-authority sites, sounding like a smaller version of the big names got us nowhere. Worse, the people who did find us couldn't tell why we existed.
The fix took a proper repositioning this summer, leading everything with the one claim that is actually different and true, that every virtual user is a real browser, and building pages around the specific problems that claim solves, like Magento checkouts that die under load. Rankings, AI citations, and the quality of demo conversations all improved once the story got narrower.
What I learned is that the differentiator you're almost embarrassed to lean on because it feels too simple is usually the whole company. Say it first, say it everywhere, and stop borrowing the incumbents' language.
I'd get advisors in the room from month one instead of trying to hold strategy, positioning, sales, and finance in my own head. I only put a proper advisory board together this July, and the speed of decisions since then has made it obvious what the delay cost.
I'd also start publishing from day one. The content that now drives most of our discovery could have been compounding a year earlier.
Evaluat at a glance
Website
Category
MRR
$10-50k
Founded
2022
Employees
2–10
Country
United Kingdom
Target market (B2B/B2C)
Business
Pricing
From $0/mo
Growth model (Product/Sales)
Sales led
Uses AI
Yes
Tech stack
Social