100% local and private
No network calls, no external services: prompts and answers stay on your machine.
The local AI inference engine for .NET and Python.
Loads quantized models and runs them entirely on your own hardware: CPU or NVIDIA GPU picked automatically, GGUF and safetensors, an OpenAI-compatible server with a web UI.
A real session with the Spark model, running locally.
Born in 2023 from the need for a local inference engine we truly own, built for the .NET ecosystem and usable from other technologies too. The core is C++ for memory control and the best performance on ordinary hardware.
No network calls, no external services: prompts and answers stay on your machine.
Probes the machine and builds an execution plan: threads, memory, quantization and backend, CPU or NVIDIA GPU.
NuGet package with an async C# API, a PyPI package and a C++ core with a stable C ABI.
Loads safetensors checkpoints directly, Mixture-of-Experts included, with no prior conversion.
Expert weights go through a tiered store, with router prediction, instead of all sitting in RAM.
If the model carries a vision encoder it is loaded automatically: you can ask questions about an image.
Tool calling, structured output, stop sequences and async streaming with cancellation.
An agent's system prompt is restored from a saved session in well under a second.
Intel Core Ultra 7 155H, NVIDIA RTX 1000 Ada GPU with 6 GB, 32 GB RAM, Windows 11. Business hardware, not workstation hardware: GPUs with more SMs scale accordingly.
| Model | Weights | Prefill 512 | Prefill 2048 | Prefill 8192 | Generation | Agent prompt, first run | From saved session |
|---|---|---|---|---|---|---|---|
| Spark-X2.5 4B | Q4_K_M | 1.918 | 2.073 | 1.943 | 46,5 | 4,3 s | 1,0 s |
| Gemma 3 4B | Q4_K_M | 2.637 | 2.624 | 2.747 | 59,5 | 3,4 s | 0,9 s |
| Qwen2.5-Coder 3B | Q8_0 | 3.035 | 3.133 | 3.024 | 49,7 | 3,0 s | 0,3 s |
| MiniCPM5 2B | Q8_0 | 4.092 | 4.007 | 3.670 | 65,6 | 2,7 s | 0,3 s |
| Model | Prefill 512 (tok/s) | Generation (tok/s) |
|---|---|---|
| Spark-X2.5 4B Q4_K_M | 89 | 17 |
| Gemma 3 4B Q4_K_M | 106 | 18 |
| Qwen2.5-Coder 3B Q8_0 | 134 | 13 |
| MiniCPM5 2B Q8_0 | 184 | 16 |
On the same model, GGUF file and GPU, DesireeIA is on par with or faster than a reference open source engine: from +0.1% to +13.3% in prefill and from +3.3% to +5.4% in generation, with more work per joule on the GPU. Laptop thermals make repeated runs vary by roughly ±10%.
The differences between families are handled by a per-model quirks system, not by one implementation each.
Pick the path you need: server with a web UI, Python library or C# library.
pip install desireeia-server
desireeia-server --model modello.gguf --port 8080
# apri http://127.0.0.1:8080
/v1/chat/completions API, router mode for several models, tool calling, voice, Prometheus metrics.
dotnet add package DesireeIA
await foreach (var piece in model.ChatStreamAsync(messages))
Console.Write(piece);
Async API with cancellation, stop sequences and structured output.
pip install desireeia
The same native engine, bundled in the package for the supported platforms.
desireeia-cli chat modello.gguf
Self-contained bundles for Windows x64 and Linux x64, no .NET or CUDA install. On macOS, build from source.
DesireeIA can be downloaded, installed and run for free, for both personal and commercial use, and shared in its original, unmodified form. Modifications, derivatives and reuse of parts of the code require the author's written permission. The code may not be used to train or replicate artificial intelligence systems without consent.
Ideas, collaborations and integrations into new architectures are welcome: get in touch.
Contact usDownload DesireeIA, try it with your own models and tell us what you build with it.
Talk to the community
See who is online, send private messages and exchange images and files (6 MB max). A free account is required.
Sign in Create an account