Hardware, OS & Infrastructure

Runs Where You Run

kiLM ships as a self-contained stack you host yourself — a single Linux box for a CPU-only trial, an NVIDIA GPU for production inference, or a Kubernetes GPU node-pool at scale. Same images on bare-metal, EKS / GKE / AKS, or fully air-gapped. Pick the tier; we size the rest with you.

Hardware Specifications

Minimum, sustained-performance specifications for serving Small, Medium and Large local LLMs. Treat these as planning floors — we confirm final sizing against your model, context length, concurrency and corpus before you buy.

ParameterSmall LLMMedium LLMLarge LLM
Model Class3B–8B14B–35B70B–72B~
Typical Quantization4-bit or 8-bit4-bit4-bit
Approx. Model-Weight Size2–8 GB9–24 GB40–55 GB
Typical useTrial, PoC, Small Department, Low ConcurrencyEnterprise Production, better Reasoning / Tool Use, Moderate ConcurrencyHigh-Quality Reasoning, Long-Context or Specialist Workloads
Processor16 physical CPU Cores, 3 GHz+ preferred24–32 physical CPU Cores48–64 physical CPU Cores per Compute / Data Node
Processor featuresModern Server / Workstation x86-64, AVX2; High Memory BandwidthServer-Grade x86-64, AVX2/AVX-512 where supportedDual-Socket or High-Core Server Platform; High Memory Bandwidth and PCIe lane Capacity
System RAM64 GB recommended; 32 GB absolute Platform floor128–256 GB ECC256–512 GB ECC per Node
GPUOptional for 3B; strongly recommended for 7B–8BRequired for satisfactory Production PerformanceRequired
Recommended VRAM16–24 GB (with GPU)32 GB minimum; 48 GB recommended80 GB allocatable per Model Replica
Example GPU classNVIDIA L4 24 GB, RTX 4090 24 GB or equivalentL40S/A6000/RTX 6000-class 48 GB; 32 GB card for constrained 30B servingA100/H100-class 80 GB or equivalent
Number of GPUs0–11 per inference replica1×80 GB per replica, or multiple GPUs only with validated tensor parallelism
Hot Storage1 TB enterprise NVMe SSD2–4 TB enterprise NVMe SSD4–8 TB NVMe per inference/data node
Model-Cache Volume250–500 GB NVMe500 GB–1 TB NVMe1–2 TB NVMe
Lake storage2 TB usable starting capacity4–10 TB usable10–50 TB+, Governed by Corpus and Retention
Backup StorageAt least 2× Protected Live DataAt least 2× protected live data, off-hostOff-host/off-site immutable backup, sized for retention and restore drills
HDD UsageCold Backup / Archive onlyCold Backup / Archive onlyCold Tier or Immutable Backup only
Network1 GbE minimum; 10 GbE preferred10 GbE25 GbE; 100 GbE for Distributed Inference / Training
Power/CoolingWorkstation/Server supply sized for one GPUDedicated Server Power and CoolingDatacenter-grade Redundant Power and Cooling
Recommended TopologyOne all-in-one HostSeparate GPU inference host and CPU/data-services host2+ GPU Replicas plus 3 Fault-Domain-separated Service / Data Nodes
Inference RuntimeOllama or vLLMvLLM for Evaluation / Production; Ollama for PoCvLLM with validated Tensor Parallelism and Replica Routing
Indicative WorkloadLow Concurrency, Shorter ContextsModerate Concurrency, 14B–32B Model, Multi-LoRA possibleHigh-Quality 70B serving, Long Context, High Concurrency or Multiple Models

Suggested Configurations

Concrete, purchasable starting points. The Medium configuration is our default recommended enterprise starting point; Large is a multi-node architecture, not simply a bigger single server.

ResourceSmall / PoCMedium / ProductionLarge / Enterprise (per inference node)
CPU cores1624–3248–64
RAM64 GB128–256 GB ECC256–512 GB ECC
GPU1×24 GB NVIDIA1×48 GB NVIDIA1×80 GB (2×80 GB only with validated tensor parallelism)
Hot storage (NVMe)1 TB2–4 TB enterprise4–8 TB local model/cache
Object / backup2 TB separate4–10 TBOff-host immutable
Network10 GbE10 GbE25 GbE or faster

A CPU-only Small install is possible, but response latency, VLM processing, embeddings and ingestion throughput will be limited — still provision 64 GB RAM if the full platform stack is enabled.

For better isolation at Medium and above, place vLLM and GPU workers on one GPU node and PostgreSQL, Qdrant, OpenSearch, Temporal and SeaweedFS on a separate CPU/data node.

Large is a topology, not a box: at least two inference replicas, three fault-domain-separated service/data nodes, off-host immutable backup, and N+1 capacity after losing one node. A single machine running several containers is process redundancy, not enterprise HA.

Sizing GPU VRAM Correctly

GPU VRAM is not a simple aggregate — a 48 GB model does not automatically fit on two 24 GB cards. Unless the exact model, quantization and runtime have passed a tensor-parallel benchmark, kiLM sizes the model against the largest single GPU.

Required VRAM =
    quantized model weights
  + KV cache  (context length × concurrent sequences)
  + inference-engine overhead
  + loaded LoRA adapters
  + 20–25% operational headroom

KV-cache use grows with both context length and simultaneous requests, so a 32B model that fits one 8K-context request may fail under several 32K-context requests. A reasonable production assumption is a 24B–35B 4-bit model on a 32 GB card — roughly 18–21 GB of weights, with the remainder for cache and adapters.

Capacity & scalability — VRAM, storage and throughput sizing as the deployment scales (illustrative).
Illustrative representation.

Storage & Training Sizing

Storage is sized against sustained IOPS and throughput, not just capacity. HDDs must not back PostgreSQL, OpenSearch, Qdrant, Temporal, active graph data or model loading — they remain fine for cold archive and backups only.

Storage classRead IOPSWrite IOPSThroughputMax p99 fsync
Trial / PoC hot stores10,0005,000200 MB/s5 ms
Evaluation / Production hot stores50,00025,000800 MB/s2 ms
Production WAL volume—10,000200 MB/s1 ms
Model / object volumeThroughput-oriented—400 MB/s seq.—
Backup / archiveNo strict floor—Capacity-oriented—

Fine-tuning and distillation need separate sizing — training must not share the production chat GPU, since it can exhaust VRAM and cause unpredictable latency. Indicative training starting points:

Training workloadPractical starting point
7B–8B LoRA / QLoRA1×24 GB GPU; 64–128 GB RAM
14B LoRA / QLoRA1×48 GB GPU; 128 GB RAM
30B–35B QLoRA1×80 GB GPU; 256 GB RAM
70B QLoRAMultiple 80 GB GPUs; benchmark exact topology
Full-parameter fine-tuningDedicated multi-GPU training cluster; outside normal install sizing
DistillationSeparate teacher-inference + student-training, or non-production windows

GPU & Model Dependencies

Beyond the host OS and container runtime, GPU serving and model onboarding carry a few customer-side prerequisites.

GPU software (BYOL)

A compatible NVIDIA host driver, the NVIDIA Container Toolkit, GPU access from Docker/Kubernetes, and — on K8s — the device plugin / GPU Operator (plus DCGM Exporter for telemetry). The NVIDIA driver is a customer/BYOL host dependency, not supplied by kiLM.

Model & AI dependencies

Customer-approved model weights with a completed license review per model; tokenizer, config and quantization files; embedding and reranker models; optional VLM/vision encoders for image search; a model registry with immutable version IDs; and golden evaluation data before promoting a model. Air-gap bundles ship all weights, manifests and license material inside the signed release.

Kubernetes (when scaling out)

Not required for Trial or PoC. Horizontally-scaled production adds a cluster, private registry, CNI/NetworkPolicy, CSI StorageClasses, ingress, GPU node pools, Helm and model-weight PVCs. The focused vLLM inference chart exists today; the full-platform umbrella chart is a larger validation milestone we call out explicitly in proposals.

Internet, Paid-Service & Licensing Posture

A fully packaged air-gap installation needs no internet during normal execution — only internal DNS, time sync, identity services and network access to your own data sources.

When internet is needed

Only when you choose to: pull images or model weights at install time, connect to cloud LLMs or SaaS, use public package/model repositories, send external email/webhook alerts, or use cloud-managed storage, identity or monitoring.

Licensing

Proprietary NVIDIA drivers/tooling are BYOL host dependencies. Every open-weight model's license is evaluated independently. Tools such as fio (GPL-2.0), considered for storage benchmarking, are not bundled without explicit approval.

TLS & secrets

kiLM does not terminate production TLS — front it with a permissively-licensed reverse proxy (Traefik, Caddy or Nginx), bring your own certificates/PKI, and provide secure secret generation and storage. Keycloak handles identity, with enterprise IdP federation where required.

What You Provide

Indicative starting points — not hard floors. Storage and memory scale with corpus size, concurrent users and the Model you choose, and we size against sustained storage performance, not just capacity. We confirm exact sizing for your workload before you buy.

Trial / PoC — CPU

A single Linux x86-64 host with Docker Engine + Compose. Indicatively 8–16 vCPU, 32–64 GB RAM, ~100 GB SSD. No GPU needed; the first CPU answer is slower while the Model warms.

Evaluation / Production — GPU

The above plus one or more NVIDIA GPUs (recent driver + CUDA). VRAM scales with the Model: roughly 16–24 GB for 7–14B-class Models, more for larger ones. NVMe storage recommended for weights + datastores.

OS & Runtime

Any current 64-bit Linux (e.g. Ubuntu 22.04+) with Docker Engine + Compose for the single-host tiers, or a Kubernetes cluster with a GPU node-pool for Eval/Production. Images are amd64.

Hard Dependencies

The non-negotiables the host must provide. Everything else kiLM needs ships inside the containers — these are the few things that live on your side of the line.

Host & OS

64-bit Linux with cgroup v2 (Ubuntu 22.04 / Debian 12 preferred), a synced clock (NTP), working DNS, a host firewall, and persistent volumes for the datastores.

Container Runtime

Docker Engine 24+ and Compose v2.20+ for the single-host tiers (or a Kubernetes cluster with a GPU node-pool for Evaluation / Production).

Deploy Tooling

Python 3.11, make, and jq on the host for the install / maintenance scripts. Containers bundle their own runtimes — these are only for the deploy tooling.

Area Dependencies
Core platform Docker / Compose, persistent volumes, host firewall + open ports, DNS, clock sync (NTP), and an external TLS reverse proxy.

TLS Termination (Yours)

kiLM expects TLS to terminate outside the stack — front production URLs with a reverse proxy such as Traefik, Caddy, or Nginx.

OS-Specific Connectors

The core stack is Linux-container only, but some BYOL engineering-tool connectors run where the tool runs: SOLIDWORKS and AutoCAD sidecars need Windows hosts; CATIA / Creo / NX may need vendor-specific Linux or Windows runtimes — usually on separate customer-side execution nodes.

You Operate, You Secure

A few things stay your responsibility: restrict admin & datastore ports, isolate BYOL sidecars, scope cloud-connector egress, apply a hardware-security baseline (Secure Boot, TPM, full-disk encryption) on production hosts, and hold the Licenses for third-party components (e.g. Ghostscript, Grafana, vendor CAD tooling).

Self-Contained Stack

kiLM bundles everything it needs — no external managed services to procure first. On Kubernetes you can optionally swap in managed datastores; on a single host it all runs in containers.

Bundled Data Plane

Knowledge Graph, vector search, relational + time-series store, object storage and the durable workflow engine all ship inside the deployment — wired together out of the box.

Identity & Observability

Bundled SSO/identity (OIDC + MFA, brokered to your IdP) and metrics/logging dashboards with GPU device-health alerting (Xid, ECC, thermal) — no separate auth or monitoring stack required to get started.

Models

Run the bundled open Models or bring your own. Ollama keeps trials CPU-friendly; vLLM gives GPU throughput at load. Model choice drives the VRAM you need.

Deployment Models

The same container images and feature set everywhere — environment differences live in overlay config, not in the product.

Single Host

Docker Compose on one Linux box (optionally GPU). The fastest path for trials, PoCs and smaller production sites.

Kubernetes

A GPU node-pool on EKS, GKE, AKS or bare-metal. Cloud-specific bits (storage class, ingress, secrets backend) are isolated to per-environment value overlays.

Air-Gapped

No internet connectivity is needed — kiLM runs fully offline.

No runtime internet: images mirrored to your private registry, Model weights pre-staged, egress denied by policy. The signed release bundle installs offline.

Tell us your corpus size, user count and whether you need GPU or air-gap, and we will recommend a tier and exact specs.

Ready to size in detail? We walk the full procurement questionnaire — target model and quantization, context length, concurrency, corpus size and ingestion rate, availability target, air-gap vs restricted egress, and backup RPO/RTO — and the complete dependency checklist with you on your quote.

Preferences saved on this device.