AI tool demos are unusually good at hiding weaknesses. The demo runs on clean data, on a prepared question, in a controlled setting, and the product looks remarkable. Then you deploy it against your actual messy business and it does not.
This is not usually deliberate deception. It is that demos are built to show capability, and what you need to know is reliability, which is a different thing and harder to show.
Here is a checklist that surfaces the difference.
Test it on something you already know
The single most useful evaluation technique, and the most skipped.
Do not ask the tool a question you want answered. Ask it a question you already know the answer to, ideally something specific to your business or your jurisdiction where a wrong answer is obviously wrong to you.
You learn two things. Whether it is accurate. And more importantly, what it does when it is wrong. Good systems hedge or say they are unsure. Poor ones state incorrect things with complete confidence, which is far more dangerous, because in production you will not know the answer.
Run five or six of these. It takes twenty minutes and tells you more than any demo.
Ask what it does when it does not know
Follow up directly. Ask the vendor how the system behaves outside its competence.
The answer you want describes a mechanism: it declines, it flags uncertainty, it escalates to a human, it cites sources you can check. The answer that should worry you is a reassurance that it is very accurate. That is not an answer to the question.
Separate what is live from what is coming
Almost every AI product page mixes shipped features with roadmap items, usually with a small label that is easy to miss in a grid.
Ask for a written list of what is available today. Then check the changelog and the roadmap, which most vendors publish and which are more honest than the marketing pages because they are written for existing customers.
Pay particular attention when something on the pricing page is quantified but not built. A plan that allocates a monthly quota of a feature still marked as planned is a signal about how the whole page was written.
Look hard at the integrations
For business AI, integrations are where the value is. An assistant that cannot see your data is a general chatbot with a different logo.
Ask three questions about each integration you care about:
Is it live today, or planned? Does it read only, or can it write? What is the sync frequency?
The read versus write distinction matters for risk. An integration that reads your bank transactions is very different from one that can initiate payments. Most of the value sits in reading. Most of the risk sits in writing.
Read the data processing agreement
Under GDPR, any vendor processing personal data for you must have one. If they cannot produce it, that is your answer.
Go to the sub-processor annex at the back. It has to name every third party touching your data, what they do and where they are. It is the most honest page a vendor publishes, and it tells you the real architecture in a way the marketing site never will.
Check that the model providers are named, that non-EU entities have a transfer safeguard listed, and that there is a notification clause for changes.
Then ask one question that catches a lot of vendors out: is the commitment not to train on your data written in the contract, or only on the website? A great many homepages carry that promise and a great many contracts do not.
Check the claims that can be checked
If a product advertises a review score, click the link. It should go to the vendor's profile on the review platform, not the platform's homepage. A score with no verifiable profile behind it is not evidence, and it tells you something about how the rest of the page was written.
If a product lists customer logos, see whether they are real linkable companies. Text names with no link, no case study and generic-sounding titles are placeholder content.
None of this proves a product is bad. Plenty of good products have overenthusiastic marketing pages. But it calibrates how much of the rest you should take at face value.
The commercial questions
What happens when you exceed a limit? Message caps, minute bundles, per-seat overages. Get the overage rate in writing.
What is the exit path? Can you export your data, in what format, and how quickly is it deleted afterwards?
What is the actual support commitment? If an SLA is advertised, check it against the terms of service. It is common for a marketing page to promise 99.9% uptime while the contract disclaims any availability guarantee. When they conflict, the contract wins.
Who is the company? For an EU vendor, company registration details should be published. Their absence is not necessarily sinister, but it is worth asking about before you route business data through the product.
What good looks like
A vendor worth buying from tends to publish more than they have to. A real changelog with dated entries. A roadmap that admits what is not built. A DPA with a full sub-processor list. An acceptable use policy that says plainly that output can be wrong and must be checked.
Mirage Cloud, a French AI platform for small businesses, publishes all four, including a sub-processor annex naming its model and voice providers with their locations and transfer safeguards. That level of documentation is not universal, and it is a reasonable proxy for whether a vendor expects to be audited.
The one-line version
Test it on something you already know, read the sub-processor list, and get the difference between live and roadmap in writing. Those three steps catch most of what a demo hides.
Comments
Log in or sign up to join the conversation.