Vendor evaluation for AI tools tends to focus on capability. The data questions get a glance at the privacy policy and a check in a box.

These are the questions worth asking, and what a good answer looks like.

Data use

1. Is our data used to train your models, or any third party's? The answer should be an unambiguous no for business tools, stated in the contract rather than in a settings toggle. "Not by default" means yes under conditions you have not read.

2. Which subprocessors see our data? A named list, including the model providers. A vendor who cannot produce this does not know their own data flow.

3. Does our data cross jurisdictions? Where is it processed and stored. Matters for regulation and increasingly for your own customers' contracts.

Retention and deletion

4. How long is our data retained, and where? Distinguish the data you upload, the prompts and outputs, and the logs. Logs are the one people forget and the one that most often retains sensitive content indefinitely.

5. What happens when we leave? Deletion within what period, verified how, including backups. "Deleted from active systems" while persisting in backups for a year is a materially different commitment.

6. Can we export everything? Data, configurations, history — in a usable format. An export that produces an unusable dump is lock-in with extra steps.

Access and security

7. Who at your company can see our data? Under what circumstances, with what approval, and is it logged. Support access is legitimate and should be controlled and auditable.

8. Encryption at rest and in transit? Table stakes. A vendor who fumbles this fails the evaluation here.

9. What is your breach notification commitment? Within what period, through what channel. Should be contractual.

The AI-specific questions

10. Which models do you use, and what changes when you switch? If they use one, ask about their contingency. If they route, ask how they validate quality across models. Either answer can be good; no answer cannot.

11. What do you record about AI-assisted decisions? For anything consequential: inputs, model, output, human approver, timestamp. This is what makes your process defensible later, and it is a capability you cannot add retroactively.

12. What is the human oversight design? Where is the gate, who approves, what happens to their edits. If the answer is that the system is fully autonomous, that is a decision you should be making deliberately rather than inheriting.

The three where answers are usually evasive

Training on your data. Watch for scope games: "we don't train our models on your data" while a subprocessor might. Ask about the whole chain.

Log retention. Frequently the honest answer is that logs contain prompts and outputs, retention is longer than the data policy implies, and nobody has examined it closely.

Deletion verification. Many vendors can delete and cannot prove it. That may be acceptable to you; it should be a known acceptance rather than an assumption.

How to use this list

Send it before the demo. The response tells you a great deal before you have spent any time.

Vendors with real practices answer quickly and specifically, often with documentation already prepared. Vendors without them return marketing language, take three weeks, or reply that the questions are unusual.

They are not unusual. They are the questions your own customers will eventually ask you.