loading experience
C++ NVIDIA CUDA Python .NET Available now v0.1.4

DesireeIA

The local AI inference engine for .NET and Python.

Loads quantized models and runs them entirely on your own hardware: CPU or NVIDIA GPU picked automatically, GGUF and safetensors, an OpenAI-compatible server with a web UI.

0network calls
65,6tokens/s generating on a business laptop
4.092tokens/s prefill on the same 6 GB GPU
20+supported architecture families
Demo

Watch it work.

A real session with the Spark model, running locally.

What it does

An engine of your own, on your machine.

Born in 2023 from the need for a local inference engine we truly own, built for the .NET ecosystem and usable from other technologies too. The core is C++ for memory control and the best performance on ordinary hardware.

100% local and private

No network calls, no external services: prompts and answers stay on your machine.

Adapts to your hardware

Probes the machine and builds an execution plan: threads, memory, quantization and backend, CPU or NVIDIA GPU.

Native for .NET and Python

NuGet package with an async C# API, a PyPI package and a C++ core with a stable C ABI.

GGUF and safetensors

Loads safetensors checkpoints directly, Mixture-of-Experts included, with no prior conversion.

MoE streamed from SSD

Expert weights go through a tiered store, with router prediction, instead of all sitting in RAM.

Vision

If the model carries a vision encoder it is loaded automatically: you can ask questions about an image.

Tools and JSON

Tool calling, structured output, stop sequences and async streaming with cancellation.

Saved sessions

An agent's system prompt is restored from a saved session in well under a second.

Performance

Measured on an ordinary laptop.

Intel Core Ultra 7 155H, NVIDIA RTX 1000 Ada GPU with 6 GB, 32 GB RAM, Windows 11. Business hardware, not workstation hardware: GPUs with more SMs scale accordingly.

GPU (CUDA), tokens per second

ModelWeightsPrefill 512Prefill 2048Prefill 8192 GenerationAgent prompt, first runFrom saved session
Spark-X2.5 4BQ4_K_M1.9182.0731.94346,54,3 s1,0 s
Gemma 3 4BQ4_K_M2.6372.6242.74759,53,4 s0,9 s
Qwen2.5-Coder 3BQ8_03.0353.1333.02449,73,0 s0,3 s
MiniCPM5 2BQ8_04.0924.0073.67065,62,7 s0,3 s

CPU only, no GPU (the path every Mac and non-NVIDIA PC takes)

ModelPrefill 512 (tok/s)Generation (tok/s)
Spark-X2.5 4B Q4_K_M8917
Gemma 3 4B Q4_K_M10618
Qwen2.5-Coder 3B Q8_013413
MiniCPM5 2B Q8_018416

On the same model, GGUF file and GPU, DesireeIA is on par with or faster than a reference open source engine: from +0.1% to +13.3% in prefill and from +3.3% to +5.4% in generation, with more work per joule on the GPU. Laptop thermals make repeated runs vary by roughly ±10%.

Models

Many families, one engine.

The differences between families are handled by a per-model quirks system, not by one implementation each.

  • Gemma 1-3
  • Qwen 2 · 3 · MoE
  • Mistral
  • MiniCPM
  • InternLM2
  • Exaone
  • SmolLM3
  • Spark 2.5
  • DeepSeek2 (MLA + MoE)
  • Mamba2
  • BERT
  • Starcoder2
  • Nemotron
  • Olmo2
  • Cohere2
  • Falcon
  • GPT-2
  • BLOOM
  • MPT
  • StableLM
  • Orion
  • Arcee
Getting started

From the terminal to your code.

Pick the path you need: server with a web UI, Python library or C# library.

OpenAI-compatible server

pip install desireeia-server
desireeia-server --model modello.gguf --port 8080
# apri http://127.0.0.1:8080

/v1/chat/completions API, router mode for several models, tool calling, voice, Prometheus metrics.

C# / .NET

dotnet add package DesireeIA

await foreach (var piece in model.ChatStreamAsync(messages))
    Console.Write(piece);

Async API with cancellation, stop sequences and structured output.

Python

pip install desireeia

The same native engine, bundled in the package for the supported platforms.

Ready-to-run CLI

desireeia-cli chat modello.gguf

Self-contained bundles for Windows x64 and Linux x64, no .NET or CUDA install. On macOS, build from source.

License

Free to use, open to dialogue.

DesireeIA can be downloaded, installed and run for free, for both personal and commercial use, and shared in its original, unmodified form. Modifications, derivatives and reuse of parts of the code require the author's written permission. The code may not be used to train or replicate artificial intelligence systems without consent.

Ideas, collaborations and integrations into new architectures are welcome: get in touch.

Contact us
DesireeIA

Intelligence that stays at home.

Download DesireeIA, try it with your own models and tell us what you build with it.