Deployment patterns

Run kiLM the Way Your Security Model Demands

kiLM ships as the same self-contained image set whatever the topology. Pick the pattern that matches your network posture — fully isolated, cloud-managed, or a blend — and the data-flow and licensing guarantees follow from it.

Three Core Patterns

Four ways to run kiLM. They differ in WHO operates it and WHERE it runs — from a fully sealed on-premise install to a POLYCRACY-managed GCP cloud, or a hybrid of the two.

kiLM deployment patterns — On-Premise (Customer Managed) with full air-gap / intake-only / limited-outbound modes, Customer Managed Cloud, POLYCRACY Managed Cloud (GCP public/private), and Hybrid.
Illustrative representation.

1. On-Premise (Customer Managed)

You run kiLM inside your own perimeter, on your Linux hosts or Kubernetes, with three egress postures: full air-gap (no inbound or outbound), intake-only (no outbound), and limited outbound (allow-listed egress).

2. Customer Managed Cloud

You run kiLM yourself on any cloud IaaS — AWS, GCP, or Azure — in your own account or project. The same single-tenant product; you own the infrastructure and its operations.

3. POLYCRACY Managed Cloud

We operate a dedicated single-tenant instance for you on GCP — as a public-cloud footprint or an isolated private-cloud footprint. You get the product without standing up or running the infrastructure.

4. Hybrid

On-premise (customer managed) for the control plane and sensitive data, combined with POLYCRACY-managed GCP cloud for specific non-sensitive workloads. You draw the line; kiLM honours it.

Deployment Tiers

One product, four tiers. Trial and PoC run on CPU via Docker Compose; Evaluation and Production use an NVIDIA GPU for high-throughput inference, on Compose-GPU or Kubernetes.

Tier Inference runtime Substrate GPU Typical use
Trial Ollama (small Model) Docker Compose Not required Single host, fastest to stand up
PoC Ollama (real-quality Model) Docker Compose Optional Low-cost evaluation on real content
Evaluation vLLM (batched) Compose-GPU or K8s GPU pool Recommended Representative throughput for sign-off
Production vLLM (batched + prefix cache) Compose-GPU or K8s GPU pool Recommended Continuous load; HA via multi-node Kubernetes

On-Premise (Customer Managed) — Data-Flow Modes

A customer-managed on-premise install runs in one of three egress postures. All three keep the product inside your network; they differ in whether kiLM may reach out at all.

a. Full Air-Gap (No Inbound or Outbound)

Fully sealed both ways: kiLM accepts no inbound connections and makes no outbound connections of any kind. Content, models, and updates enter only through a reviewed import path (signed offline bundles); nothing leaves. The strictest posture — for classified or regulated enclaves. No internet connection is required at run time.

b. Intake-Only (No Outbound)

Sealed to outbound traffic — kiLM makes no outbound connections — but a reviewed inbound intake path is permitted (for example a one-way data feed or import gateway). Nothing leaves your perimeter. No outbound internet is required at run time.

c. Limited Outbound

Stays sealed to inbound traffic but permits controlled, allow-listed outbound egress — for example to a sanctioned update mirror or an approved enterprise data source — through your egress controls. You define the allow-list; everything else is denied.

POLYCRACY Managed Cloud (GCP)

The same kiLM, operated by us on GCP as a dedicated single-tenant footprint. GCP is the default; AWS or Azure can be supported on request. Prefer to run it yourself? See the Customer Managed Cloud pattern above.

a. GCP Public Cloud

A dedicated, single-tenant GCP project with managed compute, storage, and networking on Google's public cloud, sized to your tier — the fastest managed footprint to stand up.

b. GCP Private Cloud

A more isolated GCP footprint — private networking (VPC Service Controls / private endpoints, no public ingress) for stricter data-residency and network-isolation requirements, still operated by us.

Not Sure Which Fits?

Tell us your network constraints and data sensitivity in the quote form and we will recommend a pattern. Each pattern also carries a specific licensing posture — see the Licensing Policy.

Preferences saved on this device.