loading experience
Design · The engine · DesireeIA

CORE — The proprietary inference engine: why code sovereignty is the key to AI execution

Not a black box: an engine we control.

Our local inference engine, written in C++17 with a C ABI and .NET and Python wrappers: complete control over the whole execution chain, from the model file to the token.

Scroll to explore
The principle

API consumers or owners of the engine?

In a tech landscape dominated by off-the-shelf services and black-box frameworks, most enterprises settle for acting as passive consumers of third-party APIs. The illusion of rapid initial deployment is paid for in the long run.

For us AI cannot be an opaque black box with zero tuning levers. So we chose the more demanding yet only sustainable path for enterprise: building and maintaining our own local inference engine, DesireeIA.

API consumer
  • Loss of process governance
  • Unpredictable latency spikes
  • Vendor lock-in on opaque runtimes
  • Ballooning operational costs
Proprietary engine
  • End-to-end control of the execution pipeline
  • Constant, predictable latency profile
  • Our own code, tunable at every level
  • Data and models on your own hardware
01 · Our vision

Why build a proprietary inference engine

It originates from a clear architectural priority: complete, end-to-end control over the execution pipeline.

A Hardware optimization and universal portability

CPUs with SIMD instructions

Native use of AVX2 and AVX-512 on x86_64, plus ARM NEON/SVE on Apple Silicon and edge hardware.

Discrete and embedded GPUs

Direct acceleration via modern APIs (CUDA, Vulkan, Metal): enterprise GPUs, cloud servers or low-power embedded devices.

Network efficiency

Minimal payloads between clients and nodes through optimized binary serialization, enabling edge-native and distributed topologies.

B No garbage collection: determinism

Runtimes built on garbage-collected languages introduce random pauses during inference, resulting in latency jitter. With the core in native C++, memory allocation is deterministic: zero unannounced stalls and a predictable latency profile.

02 · How it works

What is an inference engine

The runtime that loads a model's parameters (weights) into memory, accepts input tokens (prompts) and executes the tensor linear algebra required to yield output tokens sequentially.

Input textTokenizerToken IDsEmbedding + RoPE
Transformer · N layers
  1. RMSNorm / LayerNorm
  2. Self-Attention (Q, K, V)
  3. Residual connection
  4. Feed-Forward (SwiGLU)
KV Cache read and written at every layer
Final RMSNormOutput Head / LogitsSampler Temp · Top-PDetokenizer → text

the generated token is fed back as input for the next step

A

What are weights

Numerical parameters learned during training: the knowledge stored in attention matrices and feed-forward layers. In FP16 each weight takes 2 bytes. Quantization (GGUF in INT8, INT4 or K-quant) cuts it to 8, 4 or even 2 bits: less memory, faster bus transfers.

Mweights = P × b / 8
FP16
16.1 GB
INT8 (Q8_0)
8.5 GB
Q4_K_M
4.9 GB
Q2_K
2.6 GB

Example: Llama 3.1 8B (8.03 billion parameters); effective bits per weight, scales included.

B

The GGUF format and token generation

GGUF is a single-file binary container storing both the architecture metadata (layer count, attention heads, context length) and the quantized weights arranged in aligned blocks.

  1. Parsing & mmap the file is mapped into virtual address space.
  2. Tokenization text becomes a vector of Token IDs.
  3. Prefill all prompt tokens run through the layers in parallel and populate the KV Cache.
  4. Autoregressive generation the model computes Logits; the Sampler (temperature, Top-P, Top-K, Min-P) picks the token; the Detokenizer decodes it and feeds it back as input.
03 · Architecture

DesireeIA's three-tier architecture

A decoupled architecture that blends the execution speed of native C++ with the developer ergonomics of high-level languages.

Enterprise applications.NET / C# AppPython Script / Agent
Idiomatic wrappers.NET (native binding)Python
ABI / InteropC ABI · extern "C"
Native coreC++17 Engine
  • GGUF Parser & mmap Loader
  • Memory & Tensor Manager
  • KV Cache Controller
  • Hardware Offloader (CUDA / CPU)

C++17 native core

The mathematical and execution core: memory allocation, GGUF parsing, SIMD vectorization and direct hardware dispatch.

C ABI layer

A clean extern "C" interface removes C++ name mangling and enforces binary stability across compilers and runtimes, with zero-overhead interop.

Idiomatic .NET and Python wrappers

Object-oriented, asynchronous, idiomatic bindings over the C ABI, without losing underlying execution speed.

04 · Hardware resources

Disk, RAM and VRAM

A core requirement of local inference is managing how memory and compute assets are split between storage, host RAM and video VRAM.

  1. GGUF file on diskmmap zero-copy
  2. System RAMCPU layers · context
  3. Dynamic offloadingPCIe
  4. VRAM GPUGPU layers · KV cache
A

Disk and memory mapping (mmap)

Instead of reading the GGUF into host RAM through standard synchronous I/O, the engine uses mmap: the OS maps file blocks into virtual address space, enabling zero-copy page loading on demand and near-instant startup.

B

System RAM (host memory)

  • Context state and model metadata
  • Layers that don't fit in VRAM
  • KV Cache overflow on long contexts
C

VRAM and layer offloading

GPU VRAM offers orders of magnitude more memory bandwidth than system RAM. The engine can split layers: with 8 GB of VRAM and a 12 GB model, it allocates for example 20 layers to the GPU and keeps the other 12 on the CPU.

The CPU computes the initial layers, passes intermediate tensor state via PCIe, and the GPU completes the forward pass. A 12 GB model over 32 layers = 0.375 GB per layer.

VRAM
~1.000 GB/s
System RAM
60-100 GB/s
05 · Context and performance

Context, KV Cache and performance dynamics

Inference performance shifts during a session because of the interaction between the context window and the KV Cache.

A

The context window

The maximum sequence length (prompt + accumulated response tokens) the model can evaluate in a single pass. Because Self-Attention relates every token to all the preceding ones, naive complexity is quadratic:

O(N2)

N = context length.

B

KV Cache

To avoid recomputing attention for historic tokens at every step, the Key and Value tensors of every layer are stored in memory: for token N+1 the engine computes K and V only for the new token and reads the rest from cache. Its memory cost grows linearly with context:

MKV = 2 × L × hkv × d × N × b

Llama 3.1 8B: L = 32, h_kv = 8, d = 128, b = 2 bytes (FP16) → 128 KiB / token

KV Cache versus model weights (GB)

Q4_K_M weights
4.9 GB
KV · 4k tokens
0.5 GB
KV · 32k tokens
4.3 GB
KV · 128k tokens
17.2 GB

Beyond about 37,141 tokens the KV Cache outgrows the model weights themselves.

Why performance shifts: compute-bound vs memory-bound

Phase 1 · Prefill

Compute-bound

Prompt processing: hundreds or thousands of tokens evaluated concurrently (matrix-matrix multiplication, GEMM). It needs dense floating-point operations: the limit is raw compute. It dictates the Time To First Token (TTFT).

Compute cores100%
Memory busspare
Phase 2 · Generation

Memory-bandwidth bound

One token at a time (matrix-vector multiplication, GEMV): for every token the engine streams the whole weight matrix, gigabytes of data, from VRAM/RAM. The cores wait on the memory bus. It dictates throughput, in Tokens Per Second (TPS).

Compute coresidle
Memory bus100%
Upper bound on generation speed
TPSmax ≈ BW / (Mweights + MKV(N))
N = 0
205
16
N = 8k
168
13
N = 32k
109
9
N = 128k
45
4
VRAM 1,000 GB/sRAM 80 GB/s

Theoretical tokens per second for Llama 3.1 8B in Q4_K_M: a physical bandwidth bound, not a measurement. As context grows, the KV Cache adds data to read per token, TPS drops and VRAM/RAM usage climbs.

DesireeIA

Granular control over every phase of the pipeline.

From KV Cache allocation to hardware compute streams: enterprise reliability, security and predictable execution.