A new AI tool looks great in a 90-second demo. The demo doesn't show you the data policy, the failure modes, the price after the trial, or how hard it is to leave. This is a due-diligence checklist for a freelancer or small business deciding whether to trust a tool with real work and real client data — a 10-point check, red and green flags, a scoring worksheet, and a worked example on a made-up tool.
The 10 things to check
| # | Check | Where to look / what "good" looks like |
|---|---|---|
| 1 | Company identity | A real company name, address, and team are findable. An "About" with no names and no legal entity is a caution. |
| 2 | Data handling | Privacy policy states what's collected, where it's stored, and which sub-processors/model providers see it. |
| 3 | Training on your inputs | Do they train models on what you upload? Is there an opt-out, and is it on by default? |
| 4 | Content ownership & licence | Terms should say you own your inputs and outputs, and grant the vendor only the licence needed to run the service. |
| 5 | Retention & deletion | Can you delete data and your account? Does deletion actually remove it, and within a stated window? |
| 6 | Integrations & permissions | When it connects to email/drive/calendar, does it ask for read-only where possible, or full access to everything? |
| 7 | Security posture | Look for specifics (encryption in transit/at rest, SSO, an audited compliance report you can request) — not just the word "secure". |
| 8 | Pricing & trial terms | Price after the trial, whether a card is required up front, auto-renew terms, and how cancellation works. |
| 9 | Export & lock-in | Can you get your data and work out in a standard format if you leave? Or is it trapped in their UI? |
| 10 | Model/provider dependency | Which underlying model powers it? If that provider changes pricing or access, does the tool break? |
Then test the behaviour, not just the paperwork
- Low-stakes first. Run it on a draft only you'll see, or a task with no deadline, before anything client-facing.
- Provoke a failure. Feed it something ambiguous or outside its scope. Does it flag uncertainty, or state a wrong answer in the same confident tone as a right one? Quiet, confident failure is the dangerous kind.
- Check the workflow fit. Does it remove a bottleneck you actually have, or is it an impressive feature attached to no real problem? The second kind gets abandoned in a month.
Red flags
- No company name, no legal entity, no way to contact a human.
- Privacy policy is generic boilerplate that never names a data location or sub-processor.
- Training on your data is on by default with a buried or missing opt-out.
- Terms claim a broad licence to your content beyond running the service.
- No export. Your work only lives inside their app.
- Card required for a "free" trial with auto-charge and an awkward cancellation path.
- Security page is adjectives ("bank-grade", "military-grade") with zero specifics.
- Requests full account access when read-only would do.
- Roadmap and pricing change frequently with no changelog.
Green flags
- Named company, clear docs, responsive support channel.
- Privacy policy names storage regions, sub-processors, and the model provider.
- Training opt-out that's either off by default or a single clear toggle.
- Plain statement that you keep ownership of inputs and outputs.
- One-click export in a standard format; documented account deletion.
- Granular, least-privilege permissions on integrations.
- Transparent pricing page, trial without a card, self-serve cancellation.
- A public changelog and status page.
Scoring worksheet
Tool: __________________________ Date checked: __________ Score each 0 (fail) / 1 (partial) / 2 (good): [ ] 1. Company identity clear [ ] 2. Data handling documented [ ] 3. Training opt-out (off by default = 2) [ ] 4. You own inputs/outputs [ ] 5. Retention & deletion stated [ ] 6. Least-privilege integrations [ ] 7. Concrete security details [ ] 8. Pricing & trial terms clear, no card trap [ ] 9. Export exists, low lock-in [ ] 10. Model dependency understood Behaviour: [ ] Tested on low-stakes work [ ] Failure mode is visible, not silent [ ] Solves a real bottleneck I have Total: ___ / 26 Data sensitivity of intended use (low / medium / high): ______
When to skip the tool
- Any score of 0 on checks 2, 3, 4, or 5 and you'd feed it client data.
- It fails silently and you can't build a reliable review step around it.
- No export and it would hold work you can't afford to lose.
- It duplicates a tool you already have (see avoiding AI tool overload).
- The problem it solves would be fixed by a free setting change elsewhere.
Worked example — a hypothetical tool
Hypothetical, invented for illustration. "InboxZero AI" promises to draft all your client email replies. On the checklist: company is a named entity with docs (2). Privacy policy names storage region but not the model provider (1). Training on inputs is on by default, opt-out is in an account sub-menu (0). Terms confirm you own outputs (2). Deletion is documented, 30-day window (2). It requests full Gmail access, no read-only option (0). Security page lists encryption and offers a compliance report on request (2). Free trial needs a card, auto-renews monthly, cancellation is self-serve (1). Export is copy-paste only (0). Built on a single third-party model with no stated fallback (1). Behaviour: tested on a personal draft, it invented a meeting time that was never discussed and stated it plainly — a silent-ish failure (0).
Total ≈ 11/26, and it would touch a client's entire inbox. Verdict: don't connect it to real accounts. It might be usable as a standalone draft assistant you paste into — never wired into live email — and only after turning off input training.
After it passes
Approval isn't permanent. Models update and behaviour drifts, often with no announcement. Keep a one-line log per tool (name, date checked, score, notes) and re-run the check if output quality shifts or the pricing/terms change. This pairs with our rundown of common AI automation mistakes that cost freelancers clients.