Private AI that runs entirely on your machine. Your models, your conversations, your hardware.
Sovereign Inference is a desktop app for running AI language models locally on your own computer. Point it at your models folder, pick one, and start talking.
The app, your chats and your models all live on your machine and never leave it. No queue, no rate limit, no subscription. fast, private, and yours.
Available now · CPU and Vulkan editions shipping, NVIDIA CUDA builds close behind
Runs open .gguf models on your own hardware. Your data never touches a server, and never leaves the machine.
Answers appear word by word as the model generates them - never in one delayed lump - with smart caching so follow-ups come back faster.
Qwen, Llama, Mistral, Gemma, Phi and more. Prompts are formatted correctly automatically, even for brand-new models it hasn't seen before.
Temperature, max tokens, top-P, top-K, repeat penalty, context size and a custom system prompt, all saved between sessions.
Save what you'd like it to remember across every chat. Export your memory or any conversation as a portable .sovmem file.
Built natively for Windows with live RAM, CPU and GPU readouts, an accent colour that recolours everything, and an away mode that speeds up replies.
Browse and load models, manage chats, and watch the live local inference servers speed, latency and context all reported in real time.
Nothing phones home. Everything the app remembers - settings, chats, memory - stays in one folder on your machine. Problem reports are never sent automatically, and never include your keys or private data.
Everything below is off by default. Turn it on, and the data goes straight from your machine to the provider you picked - never through us.
Let the assistant look things up when a question needs current information - with the best privacy based providers available. It doesn't just skim headlines; it opens and reads the top pages to find the real answer.
Want a frontier model for a particular task? Bring your own key for OpenAI, OpenRouter or Cerebras the fastest online provider to have ever existed - which reaches GLM 5.2, DeepSeeks latest models, Claude, Gemini, GPT, Llama and more. Local and cloud share the same chat, search and memory, so you can switch freely.
Same app, same privacy - accelerated for your personal machine.
Which NVIDIA build? It depends on your driver, not your card. Run nvidia-smi and read the CUDA version in the top right, or check the NVIDIA Control Panel under System Information, then match it. If in doubt, CUDA 12 covers more cards, the newest ones included. The Vulkan edition also runs on NVIDIA and is a safe fallback.
Every edition runs the same models with the same privacy. Nothing leaves your computer on any of them. The only difference is which part of your hardware does the work. Not sure whether your graphics card is earning its keep? Install a GPU edition and use Settings › GPU offload › Test which is faster. It times your own model both ways and tells you.