Skip to content

Deployment decisions

Can Ollama run offline? Setup and an air-gap readiness checklist

Separate local inference from cloud models and connected tools. Prepare model files, disable cloud features and test the complete workflow without a network.

By AI Codex · Reviewed · 6 min read · Editorial policy (Chinese)

Ollama can run local inference with model files already downloaded, but installing Ollama does not make an entire application offline. You must also account for cloud models, web search, remote embeddings and connected tools, then test the workflow with the network disconnected.

This is a preparation and acceptance guide. It is not a claim that we performed a network audit on a particular device.

Define which layers must work offline

Layer Prepare locally Common hidden dependency
Model inference Complete model files and a working Ollama service A cloud model selection or missing files
Chat interface An installed interface that opens locally Remote page assets, sign-in or history storage
Document questions Documents, an index and local embedding model Local generation paired with cloud embeddings
Agent tools Local tools and their input data Search, web fetching and remote MCP services

Write an acceptance task such as “answer three questions whose answers are present in these local text files.” A task that requires today’s online information cannot retain that capability after you remove network access.

Prepare files and test cases while connected

Install Ollama on the target machine and download a local model appropriate for its resources. Check ollama list and record the exact name. A name shown in a separate chat interface does not prove the complete model is present on disk.

Run a minimal local task while the network still works, and record how the service was started. This makes it easier to distinguish an existing setup problem from a dependency exposed by disconnection.

Save short and representative long inputs, plus expected answers. A successful greeting will not exercise document parsing, indexing or long-context behavior. Test material should resemble the work you actually need to perform.

Disable Ollama’s cloud features

The official Ollama FAQ documents disable_ollama_cloud in ~/.ollama/server.json, or the service environment variable OLLAMA_NO_CLOUD=1. Disabling cloud features removes Ollama cloud models and web search.

Configuration field example:

{
  "disable_ollama_cloud": true
}

Merge this field into an existing configuration rather than overwriting other settings. Restart Ollama afterward and verify the result. An environment variable must reach the process that runs the service; setting it in an unrelated terminal does not reconfigure an already-running desktop application or service.

This controls Ollama’s cloud features. It does not disable networking in a third-party interface, plugin or other program. Requirements covering the complete application need checks at those layers as well.

Run an offline acceptance test

  1. End the previous test session, disconnect external networking, and reopen Ollama and your actual interface.
  2. Run ollama run MODEL_NAME, replacing the placeholder with the local name recorded from ollama list.
  3. Provide a locally saved short document and ask for specific fields. Compare the answer with the source yourself.
  4. Add a longer input, then exercise the parsing, embedding, indexing and retrieval steps you will use.
  5. Close and reopen the application and repeat. This helps expose reliance on an old session or cached state.
  6. Record the model, software version, startup method, result and exact failure message.

Success establishes offline functionality for the cases tested. It does not prove that no program ever attempts an outbound connection. If that is part of your requirement, examine firewall or network logs separately. Functioning without a network and never attempting a connection are different claims.

Troubleshoot by the first failing layer

Symptom Check first Next action
Model not found Download completeness and exact model name Prepare the model while connected, then repeat offline
Command line works but the UI fails UI service address, account and remote resources Compare the UI path with the minimal local call
Chat works but document questions fail Embedding model, parser and index location Test parsing, embedding and retrieval separately
Slow responses or process exit Memory, context length and concurrent requests Reduce the input or model size, then retry realistic material
Features disappear after disabling cloud Previous cloud model or web-search dependency Identify a local replacement or document that the feature needs connectivity

Decide whether the deployment is ready

Mark a workflow offline-ready only after its real task survives both a restart and network disconnection. List untested features, including speech, image parsing and plugins. Adding one later requires a new acceptance test for that dependency.

Local operation does not automatically settle file permissions, logging or model-license requirements. Evaluate those for the application and data involved. Configuration facts come from the Ollama FAQ; the layered acceptance procedure is our own editorial framework.

Keep exploring