CPUs with SIMD instructions
Native use of AVX2 and AVX-512 on x86_64, plus ARM NEON/SVE on Apple Silicon and edge hardware.
Not a black box: an engine we control.
Our local inference engine, written in C++17 with a C ABI and .NET and Python wrappers: complete control over the whole execution chain, from the model file to the token.
In a tech landscape dominated by off-the-shelf services and black-box frameworks, most enterprises settle for acting as passive consumers of third-party APIs. The illusion of rapid initial deployment is paid for in the long run.
For us AI cannot be an opaque black box with zero tuning levers. So we chose the more demanding yet only sustainable path for enterprise: building and maintaining our own local inference engine, DesireeIA.
It originates from a clear architectural priority: complete, end-to-end control over the execution pipeline.
Native use of AVX2 and AVX-512 on x86_64, plus ARM NEON/SVE on Apple Silicon and edge hardware.
Direct acceleration via modern APIs (CUDA, Vulkan, Metal): enterprise GPUs, cloud servers or low-power embedded devices.
Minimal payloads between clients and nodes through optimized binary serialization, enabling edge-native and distributed topologies.
Runtimes built on garbage-collected languages introduce random pauses during inference, resulting in latency jitter. With the core in native C++, memory allocation is deterministic: zero unannounced stalls and a predictable latency profile.
The runtime that loads a model's parameters (weights) into memory, accepts input tokens (prompts) and executes the tensor linear algebra required to yield output tokens sequentially.
the generated token is fed back as input for the next step
Numerical parameters learned during training: the knowledge stored in attention matrices and feed-forward layers. In FP16 each weight takes 2 bytes. Quantization (GGUF in INT8, INT4 or K-quant) cuts it to 8, 4 or even 2 bits: less memory, faster bus transfers.
Example: Llama 3.1 8B (8.03 billion parameters); effective bits per weight, scales included.
GGUF is a single-file binary container storing both the architecture metadata (layer count, attention heads, context length) and the quantized weights arranged in aligned blocks.
A decoupled architecture that blends the execution speed of native C++ with the developer ergonomics of high-level languages.
The mathematical and execution core: memory allocation, GGUF parsing, SIMD vectorization and direct hardware dispatch.
A clean extern "C" interface removes C++ name mangling and enforces binary stability across compilers and runtimes, with zero-overhead interop.
Object-oriented, asynchronous, idiomatic bindings over the C ABI, without losing underlying execution speed.
A core requirement of local inference is managing how memory and compute assets are split between storage, host RAM and video VRAM.
Instead of reading the GGUF into host RAM through standard synchronous I/O, the engine uses mmap: the OS maps file blocks into virtual address space, enabling zero-copy page loading on demand and near-instant startup.
BGPU VRAM offers orders of magnitude more memory bandwidth than system RAM. The engine can split layers: with 8 GB of VRAM and a 12 GB model, it allocates for example 20 layers to the GPU and keeps the other 12 on the CPU.
The CPU computes the initial layers, passes intermediate tensor state via PCIe, and the GPU completes the forward pass. A 12 GB model over 32 layers = 0.375 GB per layer.
Inference performance shifts during a session because of the interaction between the context window and the KV Cache.
The maximum sequence length (prompt + accumulated response tokens) the model can evaluate in a single pass. Because Self-Attention relates every token to all the preceding ones, naive complexity is quadratic:
N = context length.
To avoid recomputing attention for historic tokens at every step, the Key and Value tensors of every layer are stored in memory: for token N+1 the engine computes K and V only for the new token and reads the rest from cache. Its memory cost grows linearly with context:
Llama 3.1 8B: L = 32, h_kv = 8, d = 128, b = 2 bytes (FP16) → 128 KiB / token
Beyond about 37,141 tokens the KV Cache outgrows the model weights themselves.
Prompt processing: hundreds or thousands of tokens evaluated concurrently (matrix-matrix multiplication, GEMM). It needs dense floating-point operations: the limit is raw compute. It dictates the Time To First Token (TTFT).
One token at a time (matrix-vector multiplication, GEMV): for every token the engine streams the whole weight matrix, gigabytes of data, from VRAM/RAM. The cores wait on the memory bus. It dictates throughput, in Tokens Per Second (TPS).
Theoretical tokens per second for Llama 3.1 8B in Q4_K_M: a physical bandwidth bound, not a measurement. As context grows, the KV Cache adds data to read per token, TPS drops and VRAM/RAM usage climbs.
From KV Cache allocation to hardware compute streams: enterprise reliability, security and predictable execution.
Talk to the community
See who is online, send private messages and exchange images and files (6 MB max). A free account is required.
Sign in Create an account