Sovereign Inference

Private AI that runs entirely on your machine. Your models, your conversations, your hardware.

Run the models everyone else puts behind an API on your own machine, where nothing leaves the device.

Sovereign Inference is a desktop app for running AI language models locally on your own computer. Point it at your models folder, pick one, and start talking.

The app, your chats and your models all live on your machine and never leave it. No queue, no rate limit, no subscription. fast, private, and yours.

Available now · CPU and Vulkan editions shipping, NVIDIA CUDA builds close behind

Fully local inference

Runs open .gguf models on your own hardware. Your data never touches a server, and never leaves the machine.

Live streaming replies

Answers appear word by word as the model generates them - never in one delayed lump - with smart caching so follow-ups come back faster.

Works with any model

Qwen, Llama, Mistral, Gemma, Phi and more. Prompts are formatted correctly automatically, even for brand-new models it hasn't seen before.

Full control over output

Temperature, max tokens, top-P, top-K, repeat penalty, context size and a custom system prompt, all saved between sessions.

Memory that's yours

Save what you'd like it to remember across every chat. Export your memory or any conversation as a portable .sovmem file.

A proper native app

Built natively for Windows with live RAM, CPU and GPU readouts, an accent colour that recolours everything, and an away mode that speeds up replies.

The workspace

Your models, your chats, your local server, in one window.

Sovereign Inference
Sovereign Inference application interface

Browse and load models, manage chats, and watch the live local inference servers speed, latency and context all reported in real time.

Private by design

No telemetry. No analytics.
No account.

Nothing phones home. Everything the app remembers - settings, chats, memory - stays in one folder on your machine. Problem reports are never sent automatically, and never include your keys or private data.

Reach further, only when you choose

Two doors out - both locked until you open them.

Everything below is off by default. Turn it on, and the data goes straight from your machine to the provider you picked - never through us.

Off by default

Web search

Let the assistant look things up when a question needs current information - with the best privacy based providers available. It doesn't just skim headlines; it opens and reads the top pages to find the real answer.

Your provider: DuckDuckGo (no key), Serper, Tavily or Brave.

Off by default

Cloud models

Want a frontier model for a particular task? Bring your own key for OpenAI, OpenRouter or Cerebras the fastest online provider to have ever existed - which reaches GLM 5.2, DeepSeeks latest models, Claude, Gemini, GPT, Llama and more. Local and cloud share the same chat, search and memory, so you can switch freely.

Your key is stored only on your machine, and sent only to that provider. Cloud model terms and conditions, privacy and data are handled by the chosen provider. Full app memory features and capabilities are utilised at a cost of the users tokens which is managed by the user and provider.

Built for real use

The details you only notice when they're missing.

Your models, your way

  • Bring your own. Point it at any folder of .gguf files - subfolders included - and it finds and manages them all.
  • Tidy browser. Grouped by family, sorted and searchable, with size and quantisation shown at a glance.
  • One-click load / unload. Free up memory the moment you're done.
  • Honest guidance. Flags very low-quality quants, and tells you plainly when a model won't fit inside of your GPU, CPU or RAM instead of just failing.

Chat management

  • Pin & reorder. Keep the conversations that matter at the top.
  • Split a chat. Branch off a tangent without disturbing the original.
  • Regenerate or edit & resend your last message in a click.
  • Export anywhere. Save any chat as Markdown for your notes, or as portable memory.
  • Never lose your place with jump-to-newest and auto-scroll.

Make it yours

  • Accent colour that recolours the whole app, plus adjustable chat text size.
  • Live usage for RAM, CPU and GPU - on Intel, AMD and Nvidia.
  • Away / night mode quietly uses more CPU for faster replies when you're away.
  • Start with Windows to the system tray, ready when you are.
Install

Pick the build for your hardware.

Same app, same privacy - accelerated for your personal machine.

CPU Edition
Runs on any 64 bit Windows PC · no graphics card required.
Also the right choice if your graphics are built into the processor, such as Intel Iris or UHD, or AMD Radeon Graphics on a Ryzen chip. Those share memory with the processor, so sending work to them is usually slower than this build.
Vulkan Edition
GPU acceleration for separate graphics cards: AMD Radeon RX, Intel Arc, and NVIDIA too.
For a card with its own memory. If your graphics are built into the processor, the CPU edition will be faster.
NVIDIA · CUDA 12 Coming soon
For NVIDIA cards on a CUDA 12 driver. The safe NVIDIA choice, and the only one that works on GTX 10 series and older cards. Supports the RTX 50 series as well.
NVIDIA · CUDA 13 Coming soon
For NVIDIA cards on a CUDA 13 driver. The newest toolkit, built for Ampere, Ada and Blackwell. Needs an RTX 20 series or GTX 16 series card or newer, because CUDA 13 dropped everything older.

Which NVIDIA build? It depends on your driver, not your card. Run nvidia-smi and read the CUDA version in the top right, or check the NVIDIA Control Panel under System Information, then match it. If in doubt, CUDA 12 covers more cards, the newest ones included. The Vulkan edition also runs on NVIDIA and is a safe fallback.

Every edition runs the same models with the same privacy. Nothing leaves your computer on any of them. The only difference is which part of your hardware does the work. Not sure whether your graphics card is earning its keep? Install a GPU edition and use Settings › GPU offload › Test which is faster. It times your own model both ways and tells you.