On-premise AI versus cloud AI — an honest comparison
We sell on-premise systems, so treat our view accordingly. Here is the case against us as well as for us.
Published 26 July 2026
We build on-premise AI systems. That is a conflict of interest, and the useful thing we can do about it is put the argument against ourselves in writing.
Where cloud AI is genuinely better
Frontier capability. The largest commercial models remain ahead on novel, complex reasoning. If your work involves genuinely hard problems where the difference between a very good answer and an excellent one matters commercially, the leading cloud services will do it better. This gap has narrowed considerably and continues to narrow, but it has not closed.
Zero capital outlay. A subscription costs tens of dollars a month per person and can be cancelled. Hardware is a capital purchase. For a small team, that difference alone can settle the question.
No hardware to think about. No physical unit, no failure to plan for, no upgrade cycle. Someone else’s problem entirely.
Immediate access to new capability. When a provider ships something better, you have it that day. On-premise upgrades are constrained by what your hardware can hold.
Multimodal breadth. Cloud services handle images, audio, video and live web access as a matter of course. Local systems can do some of this, less comprehensively.
Scale on demand. Ten new staff on Monday is a billing change, not a capacity problem.
If none of your work involves confidential material, cloud AI is almost certainly the right answer and we would tell you so.
Where on-premise genuinely wins
Nothing is disclosed. This is the entire argument, and it is architectural rather than contractual. There is no third party, no overseas recipient, no terms of service to assess, no retention policy to audit. Not “protected by a good agreement” — simply absent.
Cost does not scale with use. Once purchased, it costs the same whether the team sends ten requests a day or ten thousand. This matters more than it first appears: per-message pricing quietly teaches staff to use the free consumer tool for anything routine, which is precisely the behaviour you were trying to stop.
It works offline. No internet, no problem. For a site office on a satellite link, or a practice during an outage, this is the difference between a tool and an ornament.
Predictable cost. No per-seat creep, no repricing at renewal, no surprise when usage grows.
Contractual permission. Some client agreements and confidentiality deeds simply do not permit data to be sent to third parties. On-premise is often the only way to use AI at all on that work.
Nothing is retained anywhere else. When you delete it, it is gone. There is no copy on someone else’s infrastructure that you are trusting a policy to remove.
The honest limitations
Model size is capped by hardware. A local system holds what it can hold. Large open models run well. The very largest run slowly, or not at all. Anyone claiming a desk-side machine runs the biggest models available at full speed is misleading you.
Speed varies by model architecture, considerably. This is the technical detail most vendors gloss over, and it matters. Modern “mixture of experts” models only activate a fraction of their parameters per token, so a very large MoE model can run quickly on modest hardware. Traditional dense models activate every parameter, so a large dense model on the same machine can be several times slower — sometimes slow enough to be genuinely annoying for interactive use.
The practical consequence: two models described as “70 billion parameters” and “120 billion parameters” can have wildly different speeds, and the larger one may well be the faster. If a vendor quotes a single throughput figure without naming the model and its architecture, the number is meaningless. Ask which specific models, at which speeds, and ask to see it.
Someone has to maintain it. Updates, backups, hardware failure. Usually your provider, but it is a relationship you now depend on.
It is a capital decision. Money committed up front, against a technology moving quickly. That is a real risk and it deserves to be weighed rather than waved away.
It does not make you compliant. It removes one category of risk. Your security, consent and record-keeping obligations are entirely unchanged.
The hybrid nobody mentions
These are not mutually exclusive, and for many firms the sensible answer uses both.
Confidential work — client documents, patient information, financial records, anything under an NDA — runs on the local system. General work with no confidential content runs on a cloud service, where frontier capability is available and it does not matter.
This is usually the most rational position for a mid-sized firm, and it is worth saying that it sells less hardware than the alternative. A smaller local system handling the confidential 60% of work, alongside a cloud subscription for the rest, is often better value than a large system sized to do everything.
How to work out which side you are on
Cloud is probably right if: few people use AI, usage is occasional, little of the work involves confidential material, budget is tight, or the work genuinely needs frontier capability.
On-premise starts making sense if: a whole team would use it daily, most of the work involves confidential material, a single disclosure would be a notifiable breach or a privilege problem, client contracts prohibit third-party disclosure, you operate somewhere with poor connectivity, or per-seat subscription costs across the team are already substantial.
A hybrid is probably right if you recognised your firm in both lists — which most firms of any size do.
The question that actually decides it
Not “is local AI as good as cloud AI?” — for the work most professionals do daily, the honest answer is that it is close enough that the difference rarely determines the outcome.
The question is:
What would it cost us if a client document ended up somewhere we could not account for?
If the answer is “not much”, buy a subscription and write a sensible policy. That is a legitimate answer and we will not argue with it.
If the answer involves a regulator, an insurer, a court, or a client relationship you cannot replace, then the calculation is not really about capability at all. It is about whether the exposure is one you are willing to keep carrying.
This article is general information about common obligations under Australian privacy and professional conduct rules. It is not legal, medical or financial advice and does not account for your circumstances. Obtain your own advice before acting on it.