Which AI model should your firm actually use?

2026-07-26 · 6 MIN READ · TOOLS

Which AI model should your firm actually use?

Which AI model should my firm use: Claude, ChatGPT or something else?

For most professional firms the choice matters far less than the account you use it through. Pick a business or API tier rather than a consumer plan, hold the account and keys in the firm's name, and standardise one model per workflow so the output stays consistent. Then test two providers on your own real work for a fortnight and let that decide, because capability differences shift with every release while data handling and cost control do not.

Why is "which model is best" the wrong question?

Because the answer expires. Every provider ships new versions on their own schedule, and whichever one leads on a given task this quarter may not next quarter. Choosing a firm-wide model based on a benchmark you read is choosing based on something you cannot verify and that will not hold.

The better question is which provider you want a commercial relationship with, on what terms, with what data going where. That decision is durable. Inside it, the specific model becomes a component you can change without renegotiating anything.

There is a real cost to constant switching. Prompts that produce a clean, consistent client letter with one model may produce something subtly different with another, and your team will notice the inconsistency before they notice the improvement. Standardise one model per workflow. Review it on a schedule, not on impulse.

What actually differs between the providers?

The durable differences are about how you buy, where the processing happens, and what the provider is already connected to. Everything else moves.

CriterionClaude (Anthropic)ChatGPT (OpenAI)Gemini (Google)Open-weight, self-hosted
How you buy itConsumer and team plans, direct API, or through AWS Bedrock and Google Cloud Vertex AIConsumer, team and enterprise plans, direct API, or through Microsoft AzureConsumer and Workspace plans, or through Google Cloud Vertex AIYour own infrastructure or a hosting provider
Where it fits naturallyLong document work, drafting, workflows called by codeBroadest staff familiarity, wide tooling supportFirms already deep in Google WorkspaceHigh-volume classification and extraction
Control over processing locationGood, if bought through a cloud platform with region selectionGood, if bought through a cloud platform with region selectionGood, if bought through a cloud platform with region selectionComplete
Cost shapeFixed per seat, or metered per tokenFixed per seat, or metered per tokenFixed per seat, or metered per tokenInfrastructure cost, plus your time
Main watch-outConfirm which models are available in your chosen regionConsumer plans have different data terms to business plansWorkspace integration inherits existing file permissionsQuality gap on hard reasoning, and you own the security

Microsoft Copilot deserves a separate note because it is not really a model choice. It is a distribution choice: it sits inside your Microsoft 365 tenant and inherits the permissions your staff already have. That is genuinely useful and genuinely risky, because it will happily surface anything a user can already open, including the folder someone shared too widely in 2023. If you are considering it, audit your permissions first.

How should your firm handle data going into these tools?

Decide the tiers before you decide the model, and write the rule down where staff can find it. The most common failure in a professional practice is not the model choice, it is a staff member pasting a client document into a personal consumer account because it was the tool already open in their browser.

A workable standard for a small firm:

  1. No client-identifying information goes into any personal or consumer-tier account. This is the rule that actually needs enforcing.
  2. Firm-approved accounts are business or API tier, paid on the firm's card, listed in your asset register.
  3. Read the current data terms for the exact plan you are on, not the marketing page, and record the date you read them.
  4. Know which countries the processing happens in, and confirm your privacy policy and client agreements are consistent with that.
  5. For anything touching client money, advice or correspondence, a human approves before it leaves the building.
  6. If you carry professional confidentiality obligations, have that reviewed by someone engaged to advise your practice. General guidance cannot do that job.

Whose account and whose keys?

The firm's, in the firm's name, on the firm's card, with a partner able to log in without asking anyone. That is the whole rule, and it is worth more than any model comparison.

This goes wrong quietly. A consultant builds something clever using their own API key because it was faster, or a staff member sets up the workspace under their personal email. Nothing breaks until the relationship ends or the person leaves, and then the workflow stops on a Friday and nobody can rotate the key.

Ask any provider who builds automation for you to show you the account, the billing and the key in your own name before handover. If the answer involves their platform or their key, you are renting something you were told you owned.

How do you keep the cost predictable?

Choose the pricing shape that matches the use, and put a cap on the metered one. Per-seat pricing is fixed and predictable, which suits people doing ad hoc work. Metered API pricing is variable and suits workflows, where you pay only for runs that happen.

Most firms end up with both. A few seats for the humans who are thinking with the tool, and one API key for the automations. The mistake is buying seats for everyone because it feels tidier, then discovering half of them go unused, or wiring an automation to metered billing with no ceiling and no alert.

Set a monthly spend limit on the API account. Set a billing alert well below it. Then check the actual usage after the first month against what you expected, because that gap is the most useful number you will get all quarter.

How do you test models on your own work?

Run a small, honest comparison instead of reading someone else's. This takes an afternoon to set up and answers the question properly for your firm.

  1. Pick ten real tasks you would actually hand to the tool. Real client enquiries with names removed, real quote summaries, real file notes.
  2. Write down what a good output looks like, before you run anything. Length, tone, structure, what it must never do.
  3. Run all ten through two providers using the same prompt.
  4. Have the person who normally does the work grade the results blind, without knowing which is which.
  5. Count the ones that needed no edit, some edit, or a rewrite. That ratio is your answer.
  6. Repeat on the same ten tasks in six months to see whether the choice still holds.

Consistency should weigh more heavily than peak quality. A model that produces a solid result nine times out of ten is more useful in a practice than one that produces a brilliant result seven times and something odd the other three.

What to do next

Sort the account structure first: business or API tier, firm's name, firm's card, spend cap and alert set. Write the one-line rule about consumer accounts and tell the team. Then run the ten-task test with two providers and let your own work decide.

Standardise one model per workflow, note the date and the reason in your documentation, and diarise a review in six months. When Shift builds a workflow that calls a model, it runs under the firm's own key for exactly these reasons: you can see the spend, change the model, and switch it off without asking anyone.

Common questions

Is it safe to put client information into ChatGPT or Claude?

It depends entirely on which tier you are using. Consumer plans and business or API tiers have different data handling terms, and providers generally state that business and API data is not used to train their models by default. Check the current terms for the exact plan you are on, confirm where the processing happens, and make sure your privacy policy and client agreements match. If you are bound by professional confidentiality obligations, get advice specific to your practice rather than relying on a general article.

Should we buy seats or use the API?

Seats suit people doing ad hoc thinking work, because the cost is fixed and predictable per user per month. The API suits workflows, because you pay only for what runs and you can cap it. Most firms end up with both: a few seats for the humans and one API key for the automations, billed to the firm.

Do we need to run our own model to keep data private?

Rarely. Self-hosting an open-weight model removes the third-party processor but adds hosting, security patching and a quality gap on harder tasks. For most professional firms, a business-tier account with clear data terms and a documented processing location answers the question at far lower cost and risk.

Will we have to switch models constantly to keep up?

No, and switching constantly is how firms end up with inconsistent output. Standardise one model per workflow, write your prompts so the model is a swappable component, and review the choice on a schedule rather than every time a new release is announced.

Next step

Work out what yours is costing.

The calculator on the home page takes about ten seconds, and the fit call is thirty minutes with no deck. If the honest answer is "not yet", you'll hear that.