1. On-Premise (Customer Managed)
You run kiLM inside your own perimeter, on your Linux hosts or Kubernetes, with three egress postures: full air-gap (no inbound or outbound), intake-only (no outbound), and limited outbound (allow-listed egress).
Deployment patterns
kiLM ships as the same self-contained image set whatever the topology. Pick the pattern that matches your network posture — fully isolated, cloud-managed, or a blend — and the data-flow and licensing guarantees follow from it.
Four ways to run kiLM. They differ in WHO operates it and WHERE it runs — from a fully sealed on-premise install to a POLYCRACY-managed GCP cloud, or a hybrid of the two.
You run kiLM inside your own perimeter, on your Linux hosts or Kubernetes, with three egress postures: full air-gap (no inbound or outbound), intake-only (no outbound), and limited outbound (allow-listed egress).
You run kiLM yourself on any cloud IaaS — AWS, GCP, or Azure — in your own account or project. The same single-tenant product; you own the infrastructure and its operations.
We operate a dedicated single-tenant instance for you on GCP — as a public-cloud footprint or an isolated private-cloud footprint. You get the product without standing up or running the infrastructure.
On-premise (customer managed) for the control plane and sensitive data, combined with POLYCRACY-managed GCP cloud for specific non-sensitive workloads. You draw the line; kiLM honours it.
One product, four tiers. Trial and PoC run on CPU via Docker Compose; Evaluation and Production use an NVIDIA GPU for high-throughput inference, on Compose-GPU or Kubernetes.
| Tier | Inference runtime | Substrate | GPU | Typical use |
|---|---|---|---|---|
| Trial | Ollama (small Model) | Docker Compose | Not required | Single host, fastest to stand up |
| PoC | Ollama (real-quality Model) | Docker Compose | Optional | Low-cost evaluation on real content |
| Evaluation | vLLM (batched) | Compose-GPU or K8s GPU pool | Recommended | Representative throughput for sign-off |
| Production | vLLM (batched + prefix cache) | Compose-GPU or K8s GPU pool | Recommended | Continuous load; HA via multi-node Kubernetes |
A customer-managed on-premise install runs in one of three egress postures. All three keep the product inside your network; they differ in whether kiLM may reach out at all.
Fully sealed both ways: kiLM accepts no inbound connections and makes no outbound connections of any kind. Content, models, and updates enter only through a reviewed import path (signed offline bundles); nothing leaves. The strictest posture — for classified or regulated enclaves. No internet connection is required at run time.
Sealed to outbound traffic — kiLM makes no outbound connections — but a reviewed inbound intake path is permitted (for example a one-way data feed or import gateway). Nothing leaves your perimeter. No outbound internet is required at run time.
Stays sealed to inbound traffic but permits controlled, allow-listed outbound egress — for example to a sanctioned update mirror or an approved enterprise data source — through your egress controls. You define the allow-list; everything else is denied.
The same kiLM, operated by us on GCP as a dedicated single-tenant footprint. GCP is the default; AWS or Azure can be supported on request. Prefer to run it yourself? See the Customer Managed Cloud pattern above.
A dedicated, single-tenant GCP project with managed compute, storage, and networking on Google's public cloud, sized to your tier — the fastest managed footprint to stand up.
A more isolated GCP footprint — private networking (VPC Service Controls / private endpoints, no public ingress) for stricter data-residency and network-isolation requirements, still operated by us.
Tell us your network constraints and data sensitivity in the quote form and we will recommend a pattern. Each pattern also carries a specific licensing posture — see the Licensing Policy.
Tell us about your use case. We review every request and a member of our team will be in touch within one business day.