Sovereign Inference

Private AI that runs entirely on your machine. Your models, your conversations, your hardware.

Scroll
Run the models everyone else puts behind an API on your own machine, where nothing leaves the device.

Sovereign Inference is a desktop app for running AI language models locally on your own computer. Point it at your models folder, pick one, and start talking.

The app, your chats and your models all live on your machine and never leave it. No queue, no rate limit, no subscription. fast, private, and yours.

Available now · CPU edition shipping, GPU & Vulkan close behind

Fully local inference

Runs open .gguf models on your own hardware. Your data never touches a server, and never leaves the machine.

Live streaming replies

Answers appear word by word as the model generates them - never in one delayed lump - with smart caching so follow-ups come back faster.

Works with any model

Qwen, Llama, Mistral, Gemma, Phi and more. Prompts are formatted correctly automatically, even for brand-new models it hasn't seen before.

Full control over output

Temperature, max tokens, top-P, top-K, repeat penalty, context size and a custom system prompt, all saved between sessions.

Memory that's yours

Save what you'd like it to remember across every chat. Export your memory or any conversation as a portable .sovmem file.

A proper native app

Built natively for Windows with live RAM, CPU and GPU readouts, an accent colour that recolours everything, and an away mode that speeds up replies.

The workspace

Your models, your chats, your local server, in one window.

Sovereign Inference
Sovereign Inference application interface

Browse and load models, manage chats, and watch the live local inference servers speed, latency and context all reported in real time.

Private by design

No telemetry. No analytics.
No account.

Nothing phones home. Everything the app remembers - settings, chats, memory - stays in one folder on your machine. Problem reports are never sent automatically, and never include your keys or private data.

Reach further, only when you choose

Two doors out - both locked until you open them.

Everything below is off by default. Turn it on, and the data goes straight from your machine to the provider you picked - never through us.

Off by default

Web search

Let the assistant look things up when a question needs current information - with the best privacy based providers available. It doesn't just skim headlines; it opens and reads the top pages to find the real answer.

Your provider: DuckDuckGo (no key), Serper, Tavily or Brave.

Off by default

Cloud models

Want a frontier model for a particular task? Bring your own key for OpenAI, OpenRouter or Cerebras the fastest online provider to have ever existed - which reaches GLM 5.2, DeepSeeks latest models, Claude, Gemini, GPT, Llama and more. Local and cloud share the same chat, search and memory, so you can switch freely.

Your key is stored only on your machine, and sent only to that provider. Cloud model terms and conditions, privacy and data are handled by the chosen provider. Full app memory features and capabilities are utilised at a cost of the users tokens which is managed by the user and provider.

Built for real use

The details you only notice when they're missing.

Your models, your way

  • Bring your own. Point it at any folder of .gguf files - subfolders included - and it finds and manages them all.
  • Tidy browser. Grouped by family, sorted and searchable, with size and quantisation shown at a glance.
  • One-click load / unload. Free up memory the moment you're done.
  • Honest guidance. Flags very low-quality quants, and tells you plainly when a model won't fit inside of your GPU, CPU or RAM instead of just failing.

Chat management

  • Pin & reorder. Keep the conversations that matter at the top.
  • Split a chat. Branch off a tangent without disturbing the original.
  • Regenerate or edit & resend your last message in a click.
  • Export anywhere. Save any chat as Markdown for your notes, or as portable memory.
  • Never lose your place with jump-to-newest and auto-scroll.

Make it yours

  • Accent colour that recolours the whole app, plus adjustable chat text size.
  • Live usage for RAM, CPU and GPU - on Intel, AMD and Nvidia.
  • Away / night mode quietly uses more CPU for faster replies when you're away.
  • Start with Windows to the system tray, ready when you are.
Install

Pick the build for your hardware.

Same app, same privacy - accelerated for your personal machine.

Nvidia CUDA 13
Special edition · for RTX 50-series / Blackwell · context limited only by your model and your memory, not by us
Nvidia GPU Edition
GPU-accelerated for Nvidia cards · needs an up-to-date driver
Vulkan Edition
GPU acceleration for AMD, Intel and other non-Nvidia cards
CPU Edition
Runs on any Windows PC · no graphics card required