The reference-bot build test

One small bot, specified in advance, to be built on every platform and run through the same 40 scripted conversations. This page is the protocol. The test has not been run yet.

By the chatbots.reviews teamPublished 18 September 2026Updated 18 September 2026Prices verified 18 September 2026
Status: not run. No bot has been built and no conversation has been scored. There are no results on this site, and no review here describes a product as tested. This page exists so the protocol is public before the first run.

Why one bot, built everywhere

Feature tables cannot tell you how long a builder takes to learn, or what a bot does with a question it was not prepared for. The only fair way to find out is to build the same thing on each platform and ask it the same questions. The design has to be fixed first, so that no platform's strengths shape the test.

The reference bot

The bot serves a fictional online shop that sells houseplants. It is small on purpose: a competent person should be able to build it in an afternoon on a good tool.

  1. Answer from a knowledge base. Twelve short help articles (shipping, returns, plant care, payment, gift cards) supplied as web pages and as one PDF.
  2. Look up an order. Ask for an order number, call a public mock API, and report the status, including the cases where the order does not exist or the API fails.
  3. Hand off to a person. Offer a human on request, when the customer is upset, and when the bot cannot answer twice in a row. Pass the conversation so far.
  4. Say it is a bot in the first message, as EU and some US state rules expect.

The 40 scripted conversations

GroupConversationsWhat it checks
Answerable from the knowledge base12Correct answers in plain wording
Same questions, badly typed or paraphrased6Tolerance for how people really write
German and Spanish8Answers in the customer's language from English content
Order lookup6Valid order, unknown order, missing number, API failure
Out of scope or unanswerable4Declining instead of inventing an answer
Handoff4Explicit request, frustration, repeated failure, out-of-hours
The scripts, the knowledge base and the mock API will be published with the first results so anyone can repeat the run.

What will be measured

  • Build time, in minutes, from a new account to a published bot, recorded on screen by a builder who has not used the tool before.
  • Correct, contained conversations out of 40, scored by two people against a written answer key.
  • Confident wrong answers, counted separately, because one of those costs more than one handoff.
  • Handoff behaviour: did a person get the conversation, with its history?
  • What the run cost on the vendor's own meter, so the pricing pages can be checked against a real bill.

Planned weights for rubric v2

CriterionWeightSource
Answer quality30%Reference-bot test
Build time20%Reference-bot test
Price at volume20%Product database
Channels15%Product database
Handoff to humans15%Reference-bot test
Published before the first run. If the weights change, the changelog will say why.

Limits we already know about

  • Products sold only through a sales team cannot be tested without the vendor's cooperation. They will keep a documentation-only score and be labelled that way.
  • Bots for Instagram, Messenger and WhatsApp marketing are built for a different job. They will be run on their main channel with the same scripts where the scripts apply, and the pages will say where they do not.
  • One small bot cannot show how a platform behaves with hundreds of intents, large teams or heavy traffic. The test measures getting started, not scale.
  • AI answers vary from run to run. Each conversation will be run three times and the spread reported.

How today's scores are produced without this test is explained on How we review.

Researched and drafted with AI assistance from the sources listed on this page. We have not built or tested bots on these platforms yet. Method: How we review