Local models and remote services: where inference actually runs
Where the weights sit decides who operates what, and what leaves your machine.

A model can be used in two ways, and the difference is not “cloud versus on-premise”. It is who holds
the weights and who operates the engine that runs them. Locally, you install an inference engine, put a
weights file on a disk you control, and drive it yourself; the engine happens to expose an HTTP API, but
that API is on your own machine — Ollama serves http://localhost:11434 and llama.cpp’s server listens on
127.0.0.1:8080 by default. Remotely, someone else does all of that and you hold an endpoint and a
credential. Everything else that follows — cost, failure modes, who decides when the model changes, and
what leaves the machine — follows from that single fact.
The distinction, stated precisely
| Local | Remote service | |
|---|---|---|
| Weights | a file on your disk | on the provider’s hardware |
| Engine | you install and operate it | the provider’s |
| Interface | an API on your machine (localhost) |
an HTTPS endpoint and a credential |
| Data path | prompt and response stay inside your network | both are sent to the provider |
| Capacity | what your hardware can hold | a quota you share, or a reserved deployment |
| Cost | hardware, electricity, your time | metered per token, or a subscription |
| Typical failure | the model does not fit in memory | rate limits, overload, provider incidents |
The table is a model of the trade-off, not a rule about which side is better: the same model can sit on either side, and a local engine can be reached over the network exactly like a remote one.
What the choice actually decides
Where the data goes. With a remote service the prompt and the response are transmitted to a third party, and what happens next is a policy question you must read rather than assume. OpenAI states that, as of 1 March 2023, data sent to its API is not used to train or improve its models unless the customer explicitly opts in, that abuse-monitoring logs are retained for up to 30 days by default, and that eligible customers can be approved for Zero Data Retention. Anthropic states that by default it does not use inputs or outputs from its commercial products, including the API, to train its models. Local inference does not remove this question — it changes who answers it: the engine’s own logs and your client’s logs are still data at rest on machines you operate, which is the subject of the sheet on secrets and sensitive data.
Who operates the engine. Running locally means owning the unglamorous parts: choosing and updating the engine, keeping the weights file, sizing the hardware, watching memory, and reading why a request failed. A remote service hides those, and in exchange hides the causes of its own failures behind status codes.
Who decides the version. A local weights file stays exactly as you downloaded it until you replace it. With a remote service the provider decides when a model is updated or retired, and your request keeps working until it does not. If reproducibility matters, that is a real asymmetry — and it cuts both ways: you are also the one who has to notice that a local model has become outdated.
The cost shape. Remote inference is metered per token, so cost scales with use and is bounded by nothing unless you bound it; that is the economic face of unbounded consumption, the risk OWASP lists as LLM10, where excessive or uncontrolled inference leads to denial of service, economic loss or service degradation. Local inference moves the cost into hardware and has a hard capacity ceiling instead: the model either fits or it does not.
The risk that only the local path has
Downloading a weights file is a supply-chain act. A pre-trained model is a binary artefact whose internal behaviour cannot be inspected by reading a manifest, and OWASP’s LLM03 entry notes that there are currently no strong provenance assurances for published models: model cards describe a model but guarantee nothing about where it came from, so an attacker who compromises a repository account can substitute a file. A remote service moves that risk to the provider, and adds a different one: your prompts are now outside your boundary.
What neither side gives you
Local is not automatically private, compliant or cheap: the hardware has a price, the operator is you, and the data still exists on your disks. Remote is not automatically insecure or careless: the provider may retain less than you do. The decision inputs are the classification of the data, the volume, the latency target, the acceptable cost shape, the need to control the version, and whether you have the ability to operate an engine at all.
Level and prerequisites. L2 — operational: enough to choose where inference runs and to name what the choice decides. The L1 sheets on the model/runtime/application stack and on privacy and data disclosure give the vocabulary; this sheet does not re-explain them, and it does not compare specific products.
Where to go next
- Automation & AI — the macro-area this sheet belongs to.
- AI foundations — the node this sheet sits in.
References
- Ollama — README and API reference — a local engine that runs models on your own machine and serves
them on
localhost:11434; streaming and non-streaming responses. - ggml-org — llama.cpp server — the local server’s default bind address and port, and its API key option.
- OpenAI — Data controls in the OpenAI platform — API data not used for training by default, abuse monitoring retention, Zero Data Retention.
- Anthropic — Is my data used for model training? — commercial products, including the API, not used for training by default.
- OWASP — Top 10 for LLM Applications (2025), LLM03 Supply Chain and LLM10 Unbounded Consumption — weak model provenance; excessive inference as a risk class.
- vLLM — documentation — an inference and serving engine you run yourself.