Sovereign AI Platform

The Sovereign AI Platform is a self-hosted, multi-node LLM inference cluster paired with a multi-user web chat interface. Same productivity uplift as the public chat assistants — without a single prompt, attachment, or piece of in-house knowledge ever leaving your own infrastructure. In active production.

The case for self-hosted AI is not about replicating every cloud feature. It is about the simple fact that every prompt your team types into a vendor-hosted assistant is a piece of business context handed to a third party — strategy discussions, customer-specific work, code under development, draft contracts. For most organizations that is acceptable; for some it is not.

PropertyValue
HostingSelf-hosted on a multi-node inference cluster, on your hardware
ModelsBest-in-class open-weights models (instruction-following, coding, multilingual). Switchable from the UI; the operator decides which models are available.
Chat interfaceMulti-user web UI with login, conversation history, prompt templates, model switching, file upload, retrieval-augmented chat against your own document corpus
APIOpenAI-compatible HTTP API — existing tooling and integrations work out of the box
ScalingHorizontal: add more inference nodes for capacity; nodes share model storage
AvailabilityLoad-balanced; any node can take full traffic during maintenance
Cost modelFixed hardware + electricity — no per-token, per-user, or per-month vendor fees
DeploymentIn production
  flowchart LR
    A[Users] --> B[Self-hosted web UI]
    B --> C(("Load balancer"))
    C --> D[Inference node A]
    C --> E[Inference node B]
    C --> F[Inference node C]
    B <--> G[(Vector DB · RAG)]
    D <--> H[(Shared model storage)]
    E <--> H
    F <--> H
    B -.->|SSO| I[Identity provider]

What this is about

The current generation of cloud AI assistants — ChatGPT, Claude, Gemini, Copilot — has changed how knowledge work gets done. They are also, by design, a continuous exfiltration channel for whatever the team types into them: prompts, pasted documents, code snippets, customer data, contract drafts.

The vendor terms vary. Some promise not to train on the data; some offer enterprise tiers with stronger commitments; some routinely use the data anyway under a different name. Even where the privacy commitments are watertight, the content still lives on the vendor’s infrastructure — and is therefore reachable by subpoena, government request, breach, vendor pivot, or simple operational mistake.

The Sovereign AI Platform removes that channel. The inference happens on your machines; the chat UI runs on your network; the document index for retrieval-augmented chat lives in your data store. The productivity uplift is the same; the data exposure is zero.

Operational features

  • 💬 Multi-user chat interface. Login, per-user conversation history, prompt templates, model picker, attachment upload, voice input, code-block syntax highlighting. Same modern experience the team is used to from public chat assistants.
  • 🤖 Best-in-class open-weights models. A curated set of current open-weights models for instruction following, coding, multilingual work, and reasoning. The operator decides which models are exposed in the UI. New models can be added as they release.
  • 📚 Retrieval-augmented chat against your documents. Upload internal documents, contracts, knowledge bases, manuals — the platform indexes them locally and provides retrieval-augmented answers. Documents stay on your storage; the embedding model is local; the vector database is local.
  • 🔌 OpenAI-compatible API. Endpoint exposes the same HTTP API as the public OpenAI service. Existing integrations (IDE plugins, custom tools, in-house apps) work by changing one URL.
  • 🌐 Multi-node inference cluster. Two or more inference nodes behind a load balancer. Each node can serve full load; rolling restarts and model updates happen with zero user-visible downtime.
  • 🔐 Authentication and role separation. SSO via LDAP, SAML, or OIDC against your existing identity provider. Role-based access — who can see which models, who can upload documents to the shared corpus, who can administer.
  • 📝 Conversation audit trail. Per-user conversation history retained according to your retention policy. Optional admin-side audit log for compliance use.
  • 🚦 Rate limits and quotas. Per-user, per-team, or per-model quotas — keeps a single noisy user or integration from monopolizing the cluster.
  • 🔄 Model lifecycle management. Operators can pull, test, and promote new model versions without downtime. Roll back instantly if a new version regresses on internal benchmarks.
  • 📊 Observability. Per-model latency, throughput, token usage, GPU utilization metrics. Prometheus exporters; Grafana dashboards ready to plug into your existing monitoring.

Cluster shape

The cluster runs in active/active: multiple inference nodes are reachable through a single load-balanced endpoint; each node holds a copy of the active model in memory; new requests are routed to the least-busy node.

  • Capacity scaling. Add nodes to increase concurrent user / token throughput. Throughput scales close to linearly with node count until model-loading and storage bandwidth limits kick in.
  • HA. Any node can be drained for OS patches, GPU driver updates, or hardware maintenance without interrupting users. The load balancer routes around it.
  • Model storage. Models live on shared storage (NFS, S3, or a clustered filesystem) so all nodes see the same catalog. Adding a new model on the shared store makes it available cluster-wide.
  • Vector store for RAG. A separate dedicated database node holds document embeddings for retrieval-augmented chat. Replicated for HA; query-only from the inference nodes.
  • Independent of cloud APIs. No external dependency for inference. The cluster keeps working through any internet outage, vendor disruption, or regulatory event affecting cloud AI providers.

Typical use cases

  • 🏢 Replace public chat assistants for internal use. Same productivity uplift for the team — drafting, summarising, coding help, translation, research — with no business context leaving the perimeter.
  • ⚖️ Privileged work. Legal, M&A, audit, executive drafting — work that simply cannot transit a vendor’s infrastructure for confidentiality, privilege, or regulatory reasons.
  • 🩺 Regulated sectors. Healthcare, financial advisory, insurance — sectors where the audit conversation about “where does the data go” becomes trivial when the answer is “nowhere”.
  • 🔐 IP-sensitive code. Engineering teams that need AI assistance but cannot send their proprietary code to a third-party assistant. Self-hosted inference gives the same suggest-the-next-line capability against your own codebase without exfiltration.
  • 📚 Internal knowledge assistant. Index the wiki, the ticketing system, the runbooks, and a decade of board minutes — and provide a chat interface where any team member can ask questions and get cited answers, without any of that content leaving the company.
  • 🤝 Client-facing AI features. Build a customer assistant or product-internal AI feature without paying per token, without sending your customers’ data to a third party, and without your roadmap being subject to a vendor’s pricing or policy changes.
  • 🌍 Geopolitically constrained environments. Where using a US-based or China-based AI provider is not legally or politically acceptable, sovereign self-hosted is the only option.

Why this is a CEO-level topic

  • 💰 Cost trajectory. Cloud AI assistants charge per user per month (€20–€60) plus per-token API fees. At 200 staff using AI tools daily, the annual cost is €60k–€200k and growing as adoption deepens. Self-hosted is a fixed hardware investment (€20k–€80k depending on cluster size) plus electricity — break-even is typically under 18 months.
  • 🛡️ Prompt sovereignty. Every prompt your team types into a cloud assistant is business intelligence handed to that vendor. The vendor’s terms are not the issue — the existence of the channel is the issue. Removing the channel removes the class of risks it creates.
  • 📋 Regulatory tailwind. GDPR, AI Act, sector-specific AI governance — all converge on the question of what data is sent to which AI provider. The answer “none” is the only answer that scales across jurisdictions.
  • 🔓 Vendor lock-in dissolved. Cloud AI vendors routinely deprecate models, change pricing, restrict use cases, or pivot strategically. Your AI capability should not depend on any one of those decisions going your way. Self-hosted lets you choose, test, and switch models on your timeline.
  • 🤝 Customer trust. If you build AI features into your product, “your data never leaves our infrastructure” is a material commercial advantage in regulated markets.
  • 🔗 Integration with the rest of the stack. Same identity provider, same monitoring, same data residency guarantees as the rest of the sovereign-platform estate — not a separate vendor relationship with separate compliance obligations.
  • ⏱️ Continuity. Outages and policy changes at major cloud AI providers in the past 24 months have repeatedly knocked out customer workflows. A self-hosted cluster removes that systemic risk.

Technology foundation

Built on the mature open-source AI inference ecosystem with operational discipline around clustering, observability, and model lifecycle.

LayerImplementation
Inference enginellama.cpp — fast, audited, broadly-deployed open-source inference for open-weights LLMs. CPU, CUDA, ROCm, Apple Silicon supported.
Alternative inferencevLLM for batched high-throughput serving when GPU-dense nodes are available
Web chat interfaceOpenWebUI — multi-user, conversation history, file upload, prompt templates, model switching
API gatewayOpenAI-compatible HTTP API exposed by the inference layer; existing clients work unchanged
ModelsOpen-weights models from the current best-in-class catalogue — Llama, Mistral, Qwen, DeepSeek, Gemma. Operator picks what is exposed.
Embedding modelLocal sentence-embedding model for the document index — no embedding calls to external services
Vector databaseQdrant, Weaviate, or pgvector — your choice based on existing storage stack
Load balancernginx or HAProxy fronting the inference nodes; least-conn routing
Model storageShared NFS or S3-compatible object store, accessible by all inference nodes
Identity integrationLDAP / SAML / OIDC bridge to your existing identity provider
ObservabilityPrometheus exporters for inference latency, token throughput, GPU utilisation, queue depth, per-model usage. Grafana dashboards.
Container runtimePodman or Docker — each inference node runs the same image; nodes are interchangeable.

Reach out and we will scope a deployment for your typical concurrent-user volume, model size requirements, and integration needs with the rest of the platform estate.