Models & LLM strategy

Any LLM. Your Choice, per Task.

kiLM is model-agnostic by design. Run open-source, open-weights, frontier, or paid commercial Models — whichever you are licensed for and trust — and map different Models to different tasks inside the same deployment. You own the Infrastructure that matches your Model choice; kiLM runs a basic alignment check so obvious mismatches surface before they bite.

kiLM lets you choose a different LLM per task, routing each responsibility to the model you select, in your own environment (illustrative).
Illustrative representation.

Bring Your Own Models

kiLM does not lock you to a single Model or vendor. Point it at the Model you want — running locally inside your deployment, or behind a hosted API you already pay for. The choice, the licensing, and the data-handling posture stay yours.

Model Category What it is Where it runs Why teams pick it
Open source Free Download Models with openly published weights and permissive Training/usage Terms. Locally, on your GPU/CPU — fully inside your boundary. Maximum control, no per-token cost, works air-gapped.
Open weights Free Download Downloadable weights under a model-specific license (review the Terms for your use). Locally, on your hardware. Strong capability you can self-host and Fine-Tune.
Frontier Your Purchase The latest top-capability Models, offered as a hosted API. Vendor-hosted — reached over an outbound network path you control. Best-in-class reasoning when the task warrants it.
Commercial Your Purchase Any subscription or metered API Model your organization already procures. Vendor-hosted API. Use existing contracts, SLAs and vendor relationships.

Local (open-source / open-weights) Models keep every byte inside your deployment and are the only option for fully air-gapped sites. Hosted (frontier / paid) Models require an outbound network path and are available only when your deployment posture permits outbound traffic.

Open-weight families are supported across tiers — including NVIDIA Nemotron (Nano for CPU / small-GPU sites, Super for a single GPU, and Ultra for multi-GPU installs), each with a built-in hardware-fit and license-review check. Nemotron ships under the NVIDIA Open Model License (or inherited Llama Terms), so kiLM flags it for license review rather than treating it as fully permissive, and enables it only when you opt in.

Open-weight coding models are supported the same way — including Muse Glimmer 30B (experimental, operator-supplied — pending license verification), offered in three serving profiles: a 4-bit K-quant with KV cache that runs on a single 24 GB card (32K context), a 4-bit build for the full 131K context on 32 GB+, and full-precision BF16 for a datacentre GPU (~80 GB). kiLM runs its hardware-fit check per profile and flags the model for license review; it stays out of automatic routing until an operator confirms the weights' license and specs and opts in.

Bring your own NVIDIA NIM. kiLM can connect to NVIDIA NIM inference microservices you run yourself — an opt-in, license-gated connector for Nemotron and other NIM-served Models over the standard OpenAI-compatible API. kiLM ships the connector, not the NIM containers: you supply your own NVIDIA entitlement, container, and hardware, on-prem or in your own cloud. Off by default; nothing NVIDIA is bundled or pulled unless you enable it.

Embedders (Retrieval Vectors)

Retrieval quality starts with the embedder — the Model that turns your text and images into vectors. kiLM ships permissive defaults, runs them locally (no external API, air-gap friendly), and lets you choose per deployment. The embedder is deliberately locked once you generate embeddings: the vector index is dimension-locked, so switching Model happens only through a controlled, no-downtime re-index — never a silent change.

Embedder License Dim Runs on Role in kiLM
multilingual-e5-large (default) MIT 1024 CPU or GPU Multilingual text Retrieval (JA/ZH/KO/AR + English)
SigLIP 2 Large Apache-2.0 1024 GPU recommended Image + figure/diagram Retrieval (same vector space)
FlashRank / BGE reranker MIT / Apache-2.0 — CPU (FlashRank) / GPU (BGE) Re-ranks retrieved results for precision
all-MiniLM-L6-v2 · BGE · Nomic · E5 MIT / Apache-2.0 384–1024 CPU Smaller / English-only alternatives (choose before ingesting)
Llama-Embed-Nemotron 1B / 8B (planned) NVIDIA Open Model / Llama (review) high GPU only Premium high-accuracy multilingual / enterprise Retrieval

Planned open-weight embedders (e.g. NVIDIA Llama-Embed-Nemotron) are treated as premium, opt-in, GPU-only options under license review — never the universal default. Most installs are best served by the permissive defaults above.

Usable LLMs by Tier

A quick orientation: representative Models kiLM customers run at each cost and hardware tier as of 2026. Use it to anchor sizing — not as a fixed menu. kiLM auto-applies the best free Model for a task and only uses a paid Model after an explicit admin opt-in.

Tier Example Models (2026) Runs on Typical use in kiLM
Free (open-weight, no license cost) Llama 3.2 1–3B, Qwen3 1.7–4B, Gemma 3 4B, Phi-4-mini, plus open embeddings (bge, nomic, all-MiniLM) Your own CPU or a small GPU — zero per-token cost Run kiLM end-to-end at no Model spend; extraction, classification, embeddings
Frontier (hosted API) Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, Grok 4.3 Vendor-hosted; an outbound path you control Deepest reasoning and synthesis for the few steps that need it
CPU only (no GPU) Gemma 3 4B, Qwen3 1.7–4B, Llama 3.2 1–3B, Phi-4-mini (quantised) Commodity servers or laptops — fully air-gappable Lightweight chat, routing and extraction where no GPU is available
GPU (small) — one card, ~8–24 GB Qwen3 8–14B, Gemma 4 12B, Mistral Small 4, Llama 4 8B-class, DeepSeek-R1-distill 14B, Muse Glimmer 30B (4-bit, experimental) A single consumer or pro GPU Solid local chat / RAG and a per-task workhorse
GPU (large) — multi-GPU or 48 GB+ Llama 4 Scout / Maverick, Qwen3.5 (122B), Mistral Large 3, DeepSeek V4, GLM-5.1, Kimi K2.6, Muse Glimmer 30B BF16 Multi-GPU or a large-VRAM server Top open-weight quality and long context, fully in-boundary

Examples, not limits. Model names and rankings move fast — treat this as a June 2026 snapshot. Choosing any other Model — a newer release, a different family, or a commercial API you already License — is entirely possible and is up to each customer.

Pick the Right Model for Each Task

One Model rarely fits every job. kiLM lets you assign a Model per task or pipeline stage — so a small, fast, local Model can chew through high-volume extraction while a frontier Model handles the few steps that genuinely need deeper reasoning. You tune the balance of cost, latency, privacy and quality task by task.

Task / pipeline stage A sensible Model choice The trade-off you are tuning
High-volume Ingestion & extraction Small/fast local Model Throughput & cost over peak reasoning
Embeddings & Retrieval A dedicated embedding Model Recall quality & vector dimensionality
Chat / RAG over your Knowledge A capable mid/large local or hosted Model Answer quality vs latency & privacy
Nuanced reasoning & synthesis A frontier Model (where posture allows) Top capability vs per-token cost & egress
Classification & routing A small specialised/local Model Speed & determinism at scale

Infrastructure Alignment: Yours to Own, Ours to Check

Choosing a Model implies an Infrastructure to match it — GPU memory, quantization, context window, throughput and concurrency for local Models; API keys, an outbound path and any data-residency obligations for hosted ones. Aligning your Infrastructure to the Model you select is the customer's responsibility — you size and provision it (our Infrastructure guidance helps), and we will size it with you.

You own kiLM checks (basic alignment)
Provisioning GPU memory / CPU for the chosen local Model. Flags when a Model's weights or context clearly won't fit the configured resources.
API keys, credentials and the outbound network path for hosted Models. Flags a missing key or a blocked egress path for a Model configured as hosted.
Choosing the embedding Model and its dimensionality. Flags an embedding-dimension mismatch against your vector store.
Capacity, throughput and performance tuning for your workload. Surfaces obvious misconfiguration at start-up; it does not auto-provision or guarantee performance.

The check is a guardrail, not a substitute for sizing: it catches the obvious mismatches early so a deployment fails fast and clearly rather than at runtime. Final capacity and performance remain a customer-side responsibility.

Beyond Text: Multilingual, Audio & Video

Your data isn't only English text. kiLM lets you select Models that handle multilingual corpora and multimodal inputs — audio (for example speech and transcription) and video — wherever the Model you choose and the Infrastructure you provide support it. Pick a multilingual or multimodal Model for the relevant pipeline, then align the Infrastructure it needs (typically more GPU memory, and specialised runtimes for audio or video).

The same principle applies: kiLM gives you the freedom to choose multilingual or multimodal Models and run a basic alignment check; matching the Infrastructure to those heavier workloads stays with you.

Choosing an Open-Source / Open-Weight Model for kiLM's Responsibilities

Across the whole lifecycle — Ingestion, KG & relation extraction, Ontology alignment, federated SQL planning, agentic tool-calling chat, compliance reasoning and multilingual response — kiLM leans on a specific set of Model capabilities, not raw size. Start from what each step needs, then pick a Model that clears the bar. You can still assign any Model per task; for high-risk steps, selection should fail closed.

What kiLM Asks of a Model

Capability Why kiLM needs it (and the practical bar)
Context length Long documents, multi-chunk synthesis and summarization. Practical bar: ~8k+ tokens for chat/extraction, ~32k for summarization.
Tool-calling Agentic chat, the MCP gateway and federated SQL planning are driven by tool calls. A Model without tool-calling cannot run these steps at all.
Structured output KG/relation extraction, Ontology alignment and schema/SQL generation need reliable JSON and schema adherence — not prose that merely looks structured.
Reasoning depth Compliance reasoning, relationship inference and Ontology convergence need mid-to-high reasoning, not just fluent Retrieval.
Multilingual Multilingual corpora and responses for organizations that Operate beyond English.

How Common Models Fit

Fit for kiLM's core role Example Models & guidance
Core workhorse qwen2.5:14b — clears context, structured output and multilingual for the main extraction / chat / medium-reasoning path.
Heavy reasoning mixtral:8x7b; qwen2.5:72b / llama-4-scout / approved frontier APIs — complex grounded RAG, compliance explanation, relation reasoning and the hardest paths (subject to license, hardware, outgoing-plan and cost approval).
Entry / fallback only llama3.1 (8B), qwen2.5:7b, phi-4-mini — fine for basic RAG, routing and simple structured tasks; too small alone for deep KG reasoning, long-context synthesis, compliance or federated SQL planning. Pair with a stronger Model on the heavy steps.
Specialist (supporting, not core) Vision: moondream, llava, qwen2.5vl:7b · Code: deepseek-coder:67b · Tiny: phi3:mini (3.8B/4k). Strong only in their niche; they lack the tool-calling / structured output the core role needs. Use for visual extraction, code, or classification/fallback — not as the main reasoning Model.
Unverified — gate before production Placeholder / unverified builds (e.g. gemma-4-31b, kimi-k2.6, qwen-3.6, deepseek-v4). Don't trust on a production decision path until operator-verified for the capabilities above.

Why it matters: a Model that misses a capability fails in practical ways — malformed JSON, missed tool calls, hallucinated schema mappings, unsafe SQL plans and bad Ontology updates. kiLM's guards help, but for high-risk steps Model selection should fail closed. Model names are a June 2026 snapshot — route per task and verify any new Model before trusting it on a decision path.

Preferences saved on this device.