Which AI model should your firm actually use?

Which AI model should my firm use: Claude, ChatGPT or something else?
For most professional firms the choice matters far less than the account you use it through. Pick a business or API tier rather than a consumer plan, hold the account and keys in the firm's name, and standardise one model per workflow so the output stays consistent. Then test two providers on your own real work for a fortnight and let that decide, because capability differences shift with every release while data handling and cost control do not.
Why is "which model is best" the wrong question?
Because the answer expires. Every provider ships new versions on their own schedule, and whichever one leads on a given task this quarter may not next quarter. Choosing a firm-wide model based on a benchmark you read is choosing based on something you cannot verify and that will not hold.
The better question is which provider you want a commercial relationship with, on what terms, with what data going where. That decision is durable. Inside it, the specific model becomes a component you can change without renegotiating anything.
There is a real cost to constant switching. Prompts that produce a clean, consistent client letter with one model may produce something subtly different with another, and your team will notice the inconsistency before they notice the improvement. Standardise one model per workflow. Review it on a schedule, not on impulse.
What actually differs between the providers?
The durable differences are about how you buy, where the processing happens, and what the provider is already connected to. Everything else moves.
| Criterion | Claude (Anthropic) | ChatGPT (OpenAI) | Gemini (Google) | Open-weight, self-hosted |
|---|---|---|---|---|
| How you buy it | Consumer and team plans, direct API, or through AWS Bedrock and Google Cloud Vertex AI | Consumer, team and enterprise plans, direct API, or through Microsoft Azure | Consumer and Workspace plans, or through Google Cloud Vertex AI | Your own infrastructure or a hosting provider |
| Where it fits naturally | Long document work, drafting, workflows called by code | Broadest staff familiarity, wide tooling support | Firms already deep in Google Workspace | High-volume classification and extraction |
| Control over processing location | Good, if bought through a cloud platform with region selection | Good, if bought through a cloud platform with region selection | Good, if bought through a cloud platform with region selection | Complete |
| Cost shape | Fixed per seat, or metered per token | Fixed per seat, or metered per token | Fixed per seat, or metered per token | Infrastructure cost, plus your time |
| Main watch-out | Confirm which models are available in your chosen region | Consumer plans have different data terms to business plans | Workspace integration inherits existing file permissions | Quality gap on hard reasoning, and you own the security |
Microsoft Copilot deserves a separate note because it is not really a model choice. It is a distribution choice: it sits inside your Microsoft 365 tenant and inherits the permissions your staff already have. That is genuinely useful and genuinely risky, because it will happily surface anything a user can already open, including the folder someone shared too widely in 2023. If you are considering it, audit your permissions first.
How should your firm handle data going into these tools?
Decide the tiers before you decide the model, and write the rule down where staff can find it. The most common failure in a professional practice is not the model choice, it is a staff member pasting a client document into a personal consumer account because it was the tool already open in their browser.
A workable standard for a small firm:
- No client-identifying information goes into any personal or consumer-tier account. This is the rule that actually needs enforcing.
- Firm-approved accounts are business or API tier, paid on the firm's card, listed in your asset register.
- Read the current data terms for the exact plan you are on, not the marketing page, and record the date you read them.
- Know which countries the processing happens in, and confirm your privacy policy and client agreements are consistent with that.
- For anything touching client money, advice or correspondence, a human approves before it leaves the building.
- If you carry professional confidentiality obligations, have that reviewed by someone engaged to advise your practice. General guidance cannot do that job.
Whose account and whose keys?
The firm's, in the firm's name, on the firm's card, with a partner able to log in without asking anyone. That is the whole rule, and it is worth more than any model comparison.
This goes wrong quietly. A consultant builds something clever using their own API key because it was faster, or a staff member sets up the workspace under their personal email. Nothing breaks until the relationship ends or the person leaves, and then the workflow stops on a Friday and nobody can rotate the key.
Ask any provider who builds automation for you to show you the account, the billing and the key in your own name before handover. If the answer involves their platform or their key, you are renting something you were told you owned.
How do you keep the cost predictable?
Choose the pricing shape that matches the use, and put a cap on the metered one. Per-seat pricing is fixed and predictable, which suits people doing ad hoc work. Metered API pricing is variable and suits workflows, where you pay only for runs that happen.
Most firms end up with both. A few seats for the humans who are thinking with the tool, and one API key for the automations. The mistake is buying seats for everyone because it feels tidier, then discovering half of them go unused, or wiring an automation to metered billing with no ceiling and no alert.
Set a monthly spend limit on the API account. Set a billing alert well below it. Then check the actual usage after the first month against what you expected, because that gap is the most useful number you will get all quarter.
How do you test models on your own work?
Run a small, honest comparison instead of reading someone else's. This takes an afternoon to set up and answers the question properly for your firm.
- Pick ten real tasks you would actually hand to the tool. Real client enquiries with names removed, real quote summaries, real file notes.
- Write down what a good output looks like, before you run anything. Length, tone, structure, what it must never do.
- Run all ten through two providers using the same prompt.
- Have the person who normally does the work grade the results blind, without knowing which is which.
- Count the ones that needed no edit, some edit, or a rewrite. That ratio is your answer.
- Repeat on the same ten tasks in six months to see whether the choice still holds.
Consistency should weigh more heavily than peak quality. A model that produces a solid result nine times out of ten is more useful in a practice than one that produces a brilliant result seven times and something odd the other three.
What to do next
Sort the account structure first: business or API tier, firm's name, firm's card, spend cap and alert set. Write the one-line rule about consumer accounts and tell the team. Then run the ten-task test with two providers and let your own work decide.
Standardise one model per workflow, note the date and the reason in your documentation, and diarise a review in six months. When Shift builds a workflow that calls a model, it runs under the firm's own key for exactly these reasons: you can see the spend, change the model, and switch it off without asking anyone.