---
title: "Run Ollama on Windows Server Without Exposing Your LLMs"
description: "Run Ollama as a hardened Windows Server service with NSSM, keep the API on loopback, and front it with an IIS reverse proxy, TLS, and authentication."
canonical: "https://adamtheautomator.com/run-ollama-on-windows-server-without-exposing-your/"
---

# Run Ollama on Windows Server Without Exposing Your LLMs

> Run Ollama as a hardened Windows Server service with NSSM, keep the API on loopback, and front it with an IIS reverse proxy, TLS, and authentication.

Source: https://adamtheautomator.com/run-ollama-on-windows-server-without-exposing-your/

---

ATA Learning

Tap to hide

[

ATA Learning

](/)

*   [Home](/)
*   [Tutorials](/tutorials/)
*   [Instructors](/author/)
*   [Advertising](/advertising/)
*   [Recommended Resources](/resources/)
*   [About Adam](/about-adam/)

Search for:  

*   [](https://twitter.com/adbertram)
*   [](https://github.com/Adam-the-Automator)
*   [](https://www.linkedin.com/company/adam-the-automator-llc)
*   [](/feed/)

![Run Ollama on Windows Server Without Exposing Your LLMs](https://adamtheautomator.com/wp-content/uploads/publisher/3ec5d9c85b2b8151b299cdf1129b5930/07fca0c5d7b72581bdbccb06e9058b4f8b00f553e0bd60f4e2724cfa489a164d.png)

# Run Ollama on Windows Server Without Exposing Your LLMs

[![](https://secure.gravatar.com/avatar/d0b9d42e21e5622713f8b693aa5c0f9244d5f7dd200ed29b8398f52dee5de337?s=192&d=mm&r=g)Adam Bertram](https://adamtheautomator.com/author/adam-bertram/)6 October 202611 min. read

Categories: [AI](/category/ai/)

Tags:[AI](/tag/ai/)[Security](/tag/security/)[IIS](/tag/iis/)[Microsoft Windows Server](/tag/microsoft-windows-server/)

Table of Contents

*   [Why keep an LLM on premises](#why-keep-an-llm-on-premises)
*   [What the default Windows install gets wrong](#what-the-default-windows-install-gets-wrong)
*   [Prerequisites and sizing](#prerequisites-and-sizing)
*   [Install the standalone binary](#install-the-standalone-binary)
*   [Harden the Service](#harden-the-service)
*   [Run Ollama as a Windows Service with NSSM](#run-ollama-as-a-windows-service-with-nssm)
*   [Set the Environment Variables That Matter](#set-the-environment-variables-that-matter)
*   [Keep the Bind Address Honest](#keep-the-bind-address-honest)
*   [Expose It Safely](#expose-it-safely)
*   [Put IIS in Front as a Reverse Proxy](#put-iis-in-front-as-a-reverse-proxy)
*   [Terminate TLS and Add Authentication](#terminate-tls-and-add-authentication)
*   [Tune memory and model storage](#tune-memory-and-model-storage)
*   [Patch, monitor, and audit](#patch-monitor-and-audit)
*   [Verify the hardened endpoint](#verify-the-hardened-endpoint)
*   [Which deployment shape fits](#which-deployment-shape-fits)
*   [FAQ](#faq)

Standing up an Ollama Windows Server endpoint sounds like the safe option. The prompts never leave the building, no vendor reads your data, and the bill is a power outlet instead of a per-token invoice. Then someone sets `OLLAMA_HOST=0.0.0.0` so a teammate can reach it, and within a day every machine on the subnet can run inference on your GPU, pull models to your disk, or delete the ones you already had.

The gap between “runs on my hardware” and “secure” is where most local LLM deployments, models that answer prompts on hardware you control, fail. Ollama ships as a convenience tool built for a developer’s laptop. But on a server it needs a real service wrapper, a locked bind address, a reverse proxy, and TLS in front of it. This walkthrough builds all four, using Windows-native tooling and one small helper called the Non-Sucking Service Manager ([NSSM](https://nssm.cc/)).

## Why keep an LLM on premises

The reason to run an LLM locally is data governance. When a prompt goes to a hosted API, your source code, incident logs, and customer records travel to somebody else’s servers. A private LLM on premises keeps all of it on disks and in memory you control. Ollama’s FAQ states the position verbatim: “Ollama runs locally. We don’t see your prompts or data when you run locally.” ([Ollama FAQ](https://docs.ollama.com/faq))

Keeping prompts and data on hardware you control matters for security operations, code review, and anything touching regulated data. A SOC analyst can paste a suspicious log into a local model without opening a ticket with legal. A developer can hand a model an internal repository that will never leave the network. Deployment cost is one server and a GPU instead of a metered API, which is a real line item once daily token volume climbs.

Local describes where the data lives, not how well the software is defended. Ollama’s API is unauthenticated by design, and its Windows build defaults to a per-user desktop footprint that no server should accept.

## What the default Windows install gets wrong

The official installer targets a single desktop user. It drops the binaries into `%LOCALAPPDATA%\Programs\Ollama`, stores models under that user’s profile, and starts a tray application when the user logs in. Ollama’s [Windows installation documentation](https://docs.ollama.com/windows) describes this per-user layout. On Windows Server that model fails in three ways.

*   The process dies when the user who launched it logs off, so an RDP session ending takes the endpoint down with it.
    
*   Models land on the system volume by default, which is exactly where you don’t want tens of gigabytes of weights accumulating.
    
*   The desktop build polls an auto-update mechanism, and that update path has been a remote code execution vector on Windows.
    

Running the standalone binary as a service under a system account sidesteps all three desktop-install failures. That single decision removes the update path, moves the process off any interactive session, and gives you a place to inject the environment variables the daemon actually reads.

## Prerequisites and sizing

This walkthrough was tested on Windows Server 2025 with Ollama 0.22.0. Before you touch the server, confirm the following.

*   Windows Server 2022 or 2025 with local administrator rights.
    
*   An NVIDIA GPU at compute capability 5.0 or newer with driver 550 or newer, or a supported AMD Radeon GPU. See [Ollama’s hardware support documentation](https://docs.ollama.com/gpu) for the older-card and Radeon driver requirements.
    
*   A dedicated, fast volume for model storage. Models are read from disk into VRAM at cold start, so an NVMe SSD keeps first-token latency in seconds while a spinning disk pushes it into minutes.
    
*   The NSSM service wrapper, downloaded to a folder on the system PATH.
    
*   IIS with the URL Rewrite and Application Request Routing (ARR) modules if you plan to expose the endpoint beyond the server.
    

Size the model volume before you pull anything. A 7B parameter model takes roughly 4 to 5 GB of disk (the quantized build you actually run), while a 70B model passes 40 GB, and the [Ollama model library](https://ollama.com/library) lists the size of each build. Add 20 percent headroom for the cache and future pulls.

## Install the standalone binary

Download `ollama-windows-amd64.zip` from the [Ollama release page](https://github.com/ollama/ollama/releases) rather than running `OllamaSetup.exe`. The archive holds the CLI and the GPU libraries without the tray application or the updater.

```powershell
  # Run from an elevated PowerShell prompt
Expand-Archive -Path "$env:USERPROFILE\Downloads\ollama-windows-amd64.zip" `
  -DestinationPath "C:\Program Files\Ollama" -Force

  # Confirm the binary landed
& "C:\Program Files\Ollama\ollama.exe" --version
```

Put the folder on the machine PATH so the CLI resolves from any shell.

```powershell
[Environment]::SetEnvironmentVariable(
  "Path",
  [Environment]::GetEnvironmentVariable("Path", "Machine") + ";C:\Program Files\Ollama",
  "Machine"
)
```

## Harden the Service

Running the binary as a supervised service is the first half of the job. The second half is pinning the configuration the daemon reads and keeping its API on the loopback adapter.

### Run Ollama as a Windows Service with NSSM

`ollama.exe` does not implement the Windows Service Control Manager interface. Point the SCM straight at it and the service fails with error 1053, because the executable never answers the start control within the timeout. NSSM is the fix, and its [service documentation](https://nssm.cc/usage) covers the full parameter set. It registers as the service binary, spawns `ollama serve`, and supervises the child so a crash restarts it automatically.

The diagram below shows how NSSM registers with the SCM and supervises the daemon as a child process.

![NSSM supervision](https://adamtheautomator.com/wp-content/uploads/publisher/668563705cbbe88b48e63f120d59fff1427b2b7c7f1beac7cdc4689b799de241.png)

The commands below build the service in one pass.

```powershell
  # Download NSSM and place nssm.exe somewhere on PATH first
nssm install Ollama "C:\Program Files\Ollama\ollama.exe"

  # The serve argument turns the client binary into the background API server
nssm set Ollama AppParameters "serve"
nssm set Ollama AppDirectory "C:\Program Files\Ollama"

  # Run as a system account so the endpoint survives every logout
nssm set Ollama ObjectName "LocalSystem"

  # Raise process priority; inference is CPU and GPU bound
nssm set Ollama AppPriority HIGH_PRIORITY_CLASS

  # Inject the configuration the daemon reads at startup
nssm set Ollama AppEnvironmentExtra "OLLAMA_HOST=127.0.0.1:11434" "OLLAMA_MODELS=D:\OllamaModels"

  # Log stdout and stderr for troubleshooting
nssm set Ollama AppStdout "D:\OllamaLogs\ollama.out.log"
nssm set Ollama AppStderr "D:\OllamaLogs\ollama.err.log"

nssm start Ollama
```

Create `D:\OllamaModels` and `D:\OllamaLogs` first, and grant the LocalSystem account full control of both. After the service starts, confirm it is healthy.

```powershell
Get-Service Ollama
Invoke-RestMethod http://127.0.0.1:11434/api/tags
```

If you can’t take a third-party dependency, Task Scheduler can start `ollama serve` at boot under the SYSTEM account, triggered at startup rather than at logon. It works for a simple keep-alive pattern, but you’ll lose SCM restart semantics and the built-in logging that NSSM gives you for free.

### Set the Environment Variables That Matter

Ollama reads its configuration once, when `ollama serve` starts. Export a variable in an interactive shell and the already-running service never sees it. Set these values through the NSSM `AppEnvironmentExtra` parameter or as machine-scope variables, then restart the service.

| Variable | Value to set | Why it matters |
| --- | --- | --- |
| `OLLAMA_HOST` | `127.0.0.1:11434` | Keeps the API on the loopback adapter so only local processes can reach it. |
| `OLLAMA_MODELS` | `D:\OllamaModels` | Moves weights off the system volume and onto fast storage. |
| `OLLAMA_KEEP_ALIVE` | `-1` | Keeps the model resident in memory for immediate answers instead of reloading after five minutes idle. |
| `OLLAMA_ORIGINS` | Your internal web app URL | Allows the browser client you trust, and nothing else, past Cross-Origin Resource Sharing (CORS). |
| `OLLAMA_MAX_LOADED_MODELS` | `1` | Caps concurrent models so a second pull cannot exhaust VRAM. |
| `OLLAMA_NUM_PARALLEL` | `2` | Limits simultaneous generations per model to protect GPU scheduling. |

One trap catches almost everyone. When `ollama serve` runs under LocalSystem, the account’s home directory is not your user profile, so the default model path resolves to `C:\Windows\System32\config\systemprofile\.ollama\models`. Models you pulled as your own user are invisible to the service. Setting `OLLAMA_MODELS` to a fixed local path is what makes the service and your interactive `ollama` CLI agree on where the weights live.

* * *

_**Warning: Under LocalSystem, Ollama’s default model path resolves to _**`C:\Windows\System32\config\systemprofile\.ollama\models`**_. Models you pulled as your own user stay invisible to the service until _**`OLLAMA_MODELS`**_ points at a fixed local path.**_

* * *

Pull a model to confirm the path works, then check it against `ollama list`, documented in the [Ollama CLI reference](https://docs.ollama.com/cli), before you delete anything.

```powershell
ollama pull llama3.1:8b
ollama list
```

### Keep the Bind Address Honest

Left alone, Ollama listens only on `127.0.0.1:11434`, as documented in [Ollama’s API documentation](https://docs.ollama.com/api/introduction). Nothing off the box can reach it, and that is the state you want to preserve. So leave it there. An exposed port 11434 is administrative control of the model runtime; anyone who reaches it can pull models, delete them, and run inference on your hardware.

* * *

_**Warning: A port bound to _**`0.0.0.0`**_ is administrative control of the model runtime. Anyone who reaches it can pull, delete, and run models on your GPU.**_

* * *

If a specific set of internal machines genuinely needs direct API access, you have two conditions to satisfy, and they fail in a misleading order. The binding must move to a routable address, and the Windows Defender Firewall must allow the traffic. While the service sits on loopback, no firewall rule helps, because the packets never arrive.

When you do open it, never pair `0.0.0.0` with a blanket allow rule. Scope the rule to the Private profile and to the client subnet. The two scopes below differ in who can reach the port.

![Firewall scope comparison](https://adamtheautomator.com/wp-content/uploads/publisher/2fac44bf377e64519ed606816620b7d46cd7c1a3b9aaca6c04ef11ea21f427bd.png)

```powershell
New-NetFirewallRule -DisplayName "Ollama API" `
  -Direction Inbound -Action Allow -Protocol TCP -LocalPort 11434 `
  -Profile Private -RemoteAddress 192.168.1.0/24
```

Leave the Public profile out. That profile applies on untrusted networks, which is the last place an unauthenticated API should be reachable. Then prove the change landed by reading the actual listener.

```powershell
Get-NetTCPConnection -LocalPort 11434 -State Listen
```

A `LocalAddress` of `0.0.0.0` means the bind took effect. If it still reads `127.0.0.1`, the bind never moved.

## Expose It Safely

Once the service itself is contained, exposing it to other machines becomes a deliberate step instead of a side effect of the bind address.

### Put IIS in Front as a Reverse Proxy

Firewall rules alone do not authenticate the caller, especially once clients carry dynamic addresses. The stronger pattern keeps `OLLAMA_HOST` on the loopback adapter (`127.0.0.1`) and places a reverse proxy in front. Clients talk HTTPS to IIS on port 443, IIS authenticates them, and only then does it forward the request to the loopback port.

The flow below shows where the caller is authenticated before any request reaches the loopback API.

![IIS loopback proxy](https://adamtheautomator.com/wp-content/uploads/publisher/55ca59a1ed9094bf6e39df97e953b9b9ccb20387cc3fea0f7a4314fcd9abb46e.png)

Install the URL Rewrite and Application Request Routing (ARR) modules, documented in [Microsoft’s ARR module guide](https://learn.microsoft.com/en-us/iis/extensions/planning-for-arr/using-the-application-request-routing-module). Enable the proxy from IIS Manager under **Application Request Routing Cache**, **Server Proxy Settings**, or set it from PowerShell with `Set-WebConfigurationProperty -PSPath 'MACHINE/WEBROOT/APPHOST' -Filter 'system.webServer/proxy' -Name 'enabled' -Value 'True'`.

Create a site bound to a hostname like `ai.corp.example.com`, then drop a `web.config` in its root.

```xml
<configuration>
  <system.webServer>
    <rewrite>
      <rules>
        <rule name="OllamaProxy" stopProcessing="true">
          <match url="(.*)" />
          <action type="Rewrite" url="http://127.0.0.1:11434/{R:1}" />
        </rule>
      </rules>
    </rewrite>
  </system.webServer>
</configuration>
```

One setting matters more than the rest. LLM responses stream token by token, and if the proxy buffers the whole response before forwarding, the client stares at a blank screen until generation finishes. Set the response buffer threshold to `0` in the ARR server proxy settings so chunks flow through immediately, as described in [Ollama’s streaming documentation](https://docs.ollama.com/api/streaming). If you run a chat interface like [Open WebUI](https://openwebui.com/) behind the proxy, enable WebSockets on the same site.

* * *

_**Key Insight: Ollama streams each response token as it is generated. If ARR buffers the full response, the client shows nothing until generation finishes, so the response buffer threshold must be _**`0`**_.**_

* * *

Prefer a Linux front end? An Nginx server block does the same job with a `proxy_pass http://127.0.0.1:11434` directive and the same buffering disabled. On Windows Server, IIS is the integration that ships with the operating system, so the additional software is one fewer thing to patch.

### Terminate TLS and Add Authentication

Bind the IIS site to port 443 with a certificate from your internal certificate authority or a public CA. Force the redirect from HTTP to HTTPS, and set `OLLAMA_ORIGINS` to the exact internal URL of any browser client so CORS resolves without widening access.

Authentication is the whole point of the proxy. IIS can require Windows Integrated Authentication and restrict access to an Active Directory group, so a user’s existing domain identity gates the endpoint. For service-to-service calls, an API key checked by a URL Rewrite rule, or an OAuth 2.0 and OpenID Connect flow against Microsoft Entra ID, keeps the token out of the URL. Whatever you choose, the request reaches `127.0.0.1:11434` only after it carries a valid credential. That single check blocks the unauthenticated memory-disclosure attacks against the model loader, which otherwise need only a reachable port. The “Bleeding Llama” bug, tracked as CVE-2026-7482, is the clearest example, exploiting the model creation endpoint to read server memory and push it back out ([Cyera research](https://www.cyera.com/research/bleeding-llama-critical-unauthenticated-memory-leak-in-ollama)).

## Tune memory and model storage

Ollama’s defaults assume a workstation that idles between prompts. On a shared server you want the opposite. Setting `OLLAMA_KEEP_ALIVE` to a negative value holds the model in memory until the service restarts, so the first request of the afternoon answers as fast as the first request of the morning. `ollama ps` lists which models are resident and when each unloads, as documented in [Ollama’s running models reference](https://docs.ollama.com/api/ps). In exchange, the model owns its VRAM and RAM the entire time, which is why you cap `OLLAMA_MAX_LOADED_MODELS` at one on a single-GPU host.

Storage tuning is simpler. Keep the model directory on a dedicated NVMe volume, point `OLLAMA_MODELS` at it, and never move a directory while a model is loaded. Stop the service, move the files, update the variable, and start the service again. Confirm the move with `ollama list` before you delete the old folder.

## Patch, monitor, and audit

The standalone service never touches the auto-updater, which removes the update-path vulnerabilities that affected the Windows desktop builds, tracked as CVE-2026-42248 and CVE-2026-42249 ([Striga research](https://www.striga.ai/research/ollama-windows-auto-update-rce)). You still own the patch cycle. Watch the Ollama release notes, test a new build on a staging server, and swap the binary during a maintenance window rather than letting a background process do it for you.

Restrict outbound traffic from `ollama.exe` to your own registry mirror so the service cannot pull a tainted model from a public source. Forward IIS logs and Ollama’s stdout logs to your SIEM. For visibility into who calls the API, enable auditing on the model directory and alert on new entries, and add an IIS rate limiting rule so one runaway client cannot starve the GPU. The routine behind these controls is ordinary Windows server administration, which is the point. A private LLM endpoint should show up in your monitoring stack as just another well-behaved service.

## Verify the hardened endpoint

Run the same checklist from the server and from a client machine before you call the deployment done. If a check fails, [Ollama’s troubleshooting guide](https://docs.ollama.com/troubleshooting) covers the common failure signatures.

1.  On the server, `Invoke-RestMethod http://127.0.0.1:11434/api/tags` returns the model list. The service is alive.
    
2.  From the client, `Test-NetConnection <server> -Port 11434` fails. Direct access is closed.
    
3.  From the client, `Invoke-RestMethod https://ai.corp.example.com/api/tags` succeeds only with valid credentials.
    
4.  On the server, `Get-NetTCPConnection -LocalPort 11434 -State Listen` shows the loopback or the scoped address you intended.
    
5.  Send a prompt through the proxy and confirm tokens arrive progressively instead of all at once at the end.
    

If the direct port test succeeds from a client, stop and fix the binding or the firewall rule before moving on.

## Which deployment shape fits

| Shape | Interactive session needed | Auth built in | Best for |
| --- | --- | --- | --- |
| Desktop installer | Yes | No | A developer workstation, never a server. |
| Standalone binary plus NSSM | No | No | A single server where only local processes call the API. |
| NSSM plus IIS reverse proxy | No | Yes | Shared endpoints that need TLS and directory-based access. |
| Container or WSL2 | No | No | Teams already standardized on containers and accept the extra layer. |

Most Windows Server teams choose the standalone binary plus NSSM plus an IIS reverse proxy. That shape is the only one that delivers TLS and identity integration without adding a whole new runtime to maintain, and Ollama’s [documentation](https://docs.ollama.com/) covers the standalone binary it builds on.

## FAQ

**Does Ollama have authentication?** No. The API accepts any request that reaches the port, as [Ollama’s authentication reference](https://docs.ollama.com/api/authentication) states, which is why the reverse proxy isn’t optional for a shared endpoint.

**Where should models live?** On a dedicated NVMe volume, pointed at by `OLLAMA_MODELS`. Never on the system volume, where a large pull can fill the OS drive.

**Can I keep the model loaded all the time?** Yes. Set `OLLAMA_KEEP_ALIVE=-1` and the model stays in VRAM until you restart the service, trading memory for instant answers.

**Do I still need a firewall if I use a reverse proxy?** Yes. The firewall blocks the direct port from the network, and the proxy adds the credentialed path on top of that.

A laptop experiment becomes a production endpoint when you add the wrapper, the bind address, and the proxy in front of it. Get those three right and your prompts stay private, with the API port reachable only by the proxy.

Share this article

[Share on X](https://twitter.com/intent/tweet?url=https%3A%2F%2Fadamtheautomator.com%2Frun-ollama-on-windows-server-without-exposing-your%2F&text=Run%20Ollama%20on%20Windows%20Server%20Without%20Exposing%20Your%20LLMs)[Share on Facebook](https://www.facebook.com/sharer/sharer.php?u=https%3A%2F%2Fadamtheautomator.com%2Frun-ollama-on-windows-server-without-exposing-your%2F)[Share on LinkedIn](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fadamtheautomator.com%2Frun-ollama-on-windows-server-without-exposing-your%2F)

## Related Posts

![](https://adamtheautomator.com/wp-content/uploads/2026/07/featured_image-4.webp)

### [Stop Copilot Auto-Approvals with a Human-in-the-Loop Flow](/stop-copilot-auto-approvals-human-loop/)

Build a Copilot workflow that reads business requests, auto-approves the routine ones, routes judgment calls to Teams, and escalates timeouts.

![](https://adamtheautomator.com/wp-content/uploads/2026/06/featured_image_v2.png)

### [When AI Automates the Joy Out of Work](/ai-automation-work/)

A personal look at how AI agents and automation can boost productivity while draining the satisfaction that makes hands-on technical work feel meaningful.

![](https://adamtheautomator.com/wp-content/uploads/2025/05/Untitled.jpg)

### [When a Script is Worth a Thousand AI Prompts](/when-script-worth-thousand-ai-prompts/)

When traditional scripts outshine AI: a developer’s journey back to automation basics.

## Categories

*   [IT Ops](/category/it-ops/)
*   [Cloud](/category/cloud/)
*   [DevOps](/category/devops/)
*   [Home Ops](/category/home-ops/)
*   [Information Security](/category/infosec/)
*   [Software Development](/category/software-development/)

## Site

*   [Home](/)
*   [Tutorials](/tutorials/)
*   [Instructors](/author/)
*   [Advertising](/advertising/)
*   [Recommended Resources](/resources/)
*   [About Adam](/about-adam/)

Copyright 2026© ATA Learning | [Privacy Policy](/privacy/)
