Ten Questions to Ask Before You Buy an AI-Powered Product
Every software vendor is shipping AI features right now, and most of the marketing pages say roughly the same thing. The questions below cut through that by focusing on the only thing that actually changes your risk profile: where your data goes, who touches it, how long it stays there, and what happens if the vendor is wrong.
This is not a legal review, and it is not a substitute for your attorney or compliance officer. It is the working checklist we hand to clients in construction, medical and dental, law, and professional services when they are evaluating an AI-powered tool and want to know what to ask before signing. Send it to the vendor in writing. Ask for written answers. A vendor who can answer these in a day is a different kind of vendor than one who needs three weeks and a call with legal.
Before the checklist: know what you are actually buying
“AI-powered” describes at least four different architectures, and they carry different risks.
Some products run a model inside your existing tenant, under your existing agreements. Microsoft 365 Copilot is the common example here: it operates on data already in your Microsoft 365 environment and inherits the permissions and commitments you already have in place. Some products call a third-party model API (OpenAI, Anthropic, Google) on your behalf, meaning your data crosses into a second vendor’s infrastructure. Some host an open-weight model on their own servers. And some are thin wrappers on a public chatbot with a login screen.
The first question in any vendor call should be “which of those are you?” Everything downstream depends on the answer.
The ten questions
1. Where does our data physically go when we use the AI feature?
Ask for the data flow in plain language: our user types something, it goes to which systems, in which countries, and who operates each hop. You want the answer to name specific infrastructure, not “our secure cloud.” If the vendor routes to a third-party model provider, you need that named too. For firms with contractual data-residency obligations — common in manufacturing work tied to defense supply chains, and increasingly in professional services contracts — a vendor who cannot tell you the country is disqualifying on its own.
2. Is our data used to train your models, or anyone else’s?
The question has two halves, and vendors often answer only the first. “We do not train on customer data” can still be true while the underlying model provider retains prompts for abuse monitoring. Ask specifically: is our input used for training, fine-tuning, model improvement, evaluation, or human review by you or by any sub-processor? Get the answer in the contract, not the FAQ page. Also ask whether the default is opt-in or opt-out, and whether the setting is per-tenant or per-user, because a per-user opt-out means one new hire can undo your policy.
3. Who are your sub-processors, and how do we learn when the list changes?
Most AI features involve at least one sub-processor, often several: the model provider, a vector database host, a transcription service, a monitoring vendor. Ask for the current published list and the notification mechanism for additions. A good vendor maintains a sub-processor page with an email subscription. A vendor who says “we don’t disclose that” is telling you that your data governance stops at their front door, which is not a governance program.
4. How long do you retain prompts, outputs, and uploaded files — and can we set it?
Retention is where AI tools quietly become a discovery and breach-exposure problem. A chatbot that keeps every prompt for 18 months has built a searchable archive of your staff’s questions, including the ones containing client names, patient details, or settlement figures. Ask for the default retention window for prompts, outputs, attachments, and logs separately. Ask whether you can shorten it, whether deletion is real deletion or soft-delete, and how long backups persist after that. For law firms weighing privilege exposure, this is the single most consequential answer on the list.
5. How is our tenant isolated from other customers?
Multi-tenant is fine. Multi-tenant with weak isolation is not. Ask whether isolation is logical (row-level filtering in a shared database) or physical (separate instances), how embeddings and vector indexes are partitioned, and whether any caching layer is shared across customers. The failure mode you are guarding against is a retrieval bug that surfaces another company’s document in your search results. It has happened publicly to consumer AI products, and the architecture question is how you assess the likelihood.
6. Does the AI feature respect our existing permissions?
This one catches people. If a tool indexes your file shares or SharePoint to answer questions, it may index content the asking user is not entitled to see. Ask whether the AI enforces source-system permissions at query time — meaning the answer is filtered per user, live — or whether it indexes everything once and trusts the interface. Then ask how permission changes propagate and how fast. We have seen this drive real cleanup work: before turning on tenant-wide AI search, most organizations discover folders that were over-shared years ago and never corrected. That cleanup is worth doing regardless, and it is a normal part of the readiness work we cover under AI security and governance.
7. What compliance attestations do you hold, and what is in scope?
Ask for the SOC 2 Type II report, not the badge on the website. Read the scope section: a report covering the vendor’s core platform may explicitly exclude the new AI module. For medical and dental practices, ask whether the vendor will sign a Business Associate Agreement and whether the BAA covers AI processing and all sub-processors in the chain. For manufacturers in defense supply chains, ask directly about CMMC alignment and whether any AI processing touches controlled unclassified information. “We’re SOC 2 compliant” without a report and a scope section is a marketing sentence.
8. How do you handle prompt injection, jailbreaks, and adversarial inputs?
If the tool reads untrusted content — inbound email, uploaded PDFs, web pages, vendor invoices — an attacker can hide instructions in that content to manipulate the AI’s behavior. Ask what controls exist, whether the vendor runs adversarial testing, and how findings are handled. You are listening for evidence that the vendor has thought about it at all. Many have not.
9. What are your breach notification terms, and who is liable?
Read the notification window, the trigger definition, and the liability cap. A cap set at twelve months of fees is common and means the vendor’s financial exposure is small relative to yours. For SMBs in the 25 to 300 employee range, recovery from a serious data incident typically runs several days to a couple of weeks of degraded operations, and industry benchmarks put downtime cost for firms this size in the range of roughly $5,000 to $25,000 per day. Compare that projected exposure to the cap before you sign.
10. What happens to our data if we leave?
Ask for the export format, the deletion timeline after termination, the certificate-of-deletion process, and whether derived data — embeddings, fine-tuned weights, aggregated analytics — is deleted too. Derived data is frequently excluded. Get it in writing before the contract, because you will have no leverage after.
How to use this without stalling every purchase
Not every tool needs a full ten-question review. Tier it. A tool that touches patient records, client matters, or engineering drawings gets the full checklist and a legal read. A meeting transcription tool that only touches internal standups gets questions 2, 4, and 7. A marketing copy generator that never sees customer data gets a quick look at question 2 and a note in your acceptable-use policy.
That tiering only works if you have an AI policy that defines the tiers and names an owner. Without one, the practical result is shadow AI: staff signing up for free accounts with business credit cards and pasting whatever helps them finish a task. A written AI policy, a short approved-tools list, and one person who owns the review is enough structure for most organizations under 300 people. It also makes vendor conversations faster, because you know what you are asking for.
Where Century fits
We run this process with clients across medical and dental, legal, construction, and manufacturing in Atlanta and across Georgia. In practice the work is: inventory what AI tools are already in use, classify data sensitivity, run the checklist against the vendors that matter, and write a policy your staff will actually follow. Most of it takes weeks, not quarters.
If you are evaluating an AI-powered product right now and want a second set of eyes on the vendor’s answers, schedule a short discovery call. Bring the vendor’s security documentation and we will walk through it with you, flag the gaps, and tell you plainly whether the tool is a reasonable fit for your data. No obligation, and no pitch for software you do not need.

