loading experience
Design · Software design

STACK — Software design in the age of AI: architecture, polyglot tech stack and human-centricity

AI-native, not AI-bolted-on.

Designing in the age of AI isn't adding an endpoint to an LLM: it takes sizing, microservice Onion Architecture, security, reliability and cost control, a polyglot stack and an interface that adapts to people.

Scroll to explore
The principle

An AI-native architecture takes engineering, not an endpoint.

An enterprise AI-native architecture requires a rigorous approach to resource sizing, a microservice organization based on Onion Architecture, a balance between security, reliability and cost, and the selection of the polyglot tech stack best suited to combine performance, maintainability and innovation.

01 · AI Sizing

AI sizing and model selection

It starts by analyzing the workload to find the right balance between accuracy, latency and Total Cost of Ownership (TCO). Not every operation needs a model with hundreds of billions of parameters.

SLM1B – 8B

Classification, triage, entity extraction, structured parsing. On-premise or on modest hardware, latency under 50 ms.

LLM70B+

Complex synthesis, long-form content and advanced RAG (Retrieval-Augmented Generation).

ReasoningSystem Two

Multi-step planning, logical audit and articulated problems, where inference-time compute replaces raw speed.

Incoming request Classification & triageSLM
Structured task / Enum confidence ≥ 85%

Local SLM

1B-8B · On-Premise
Analysis / generation complex or ambiguous case

LLM / Reasoning

Private Cloud / API

structured output

Σ Quantitative parameters

Required VRAM
VRAM ≈ ( P × b / 8 + MKV(N) ) × 1,10

P = parameters · b = bits per weight (FP16 = 16, FP8 = 8, INT4 ≈ 4) · M_KV = KV cache for context N · 1.10 = headroom for buffers and activations.

VRAM needed with an 8k-token context (GB)

8B · FP16
18.8 GB 24 48 80
8B · FP8
10.0 GB 24 48 80
8B · Q4_K_M
6.5 GB 24 48 80
70B · FP16
158.3 GB 24 48 80
70B · FP8
80.6 GB 24 48 80
70B · Q4_K_M
50.0 GB 24 48 80

The badges show which GPU the model fits on (24, 48, 80 GB): green = fits. Quantizing from FP16 to INT4 cuts the weights by up to 75% (GGUF, AWQ, FP8 formats) with typically under 1% accuracy loss, to be verified on your own task. Llama 3.1: 8B = 32 layers, 70B = 80 layers, 8 KV heads, d = 128.

Cloud · API managed

OPEX

  • Immediate scalability
  • Zero hardware management
  • Variable costs and data-sovereignty constraints
On-Premise / Private Cloud

CAPEX

  • Full data governance and air-gapping
  • Fixed, predictable compute cost at large volumes
  • Capital cost and infrastructure management
02 · Onion Architecture

Microservice Onion Architecture

The architecture must isolate the application domain from fast-changing AI technology. Onion Architecture keeps business logic independent from models, vector databases and cloud providers: dependencies always point inward.

  1. Domain Layer invariant core

    Business entities, domain rules and system events. No dependency on external modules or AI frameworks.

  2. Application Layer Use Cases & CQRS

    Use cases, workflow orchestration, commands and queries (CQRS via MediatR/C#).

  3. Infrastructure Layer Adapter & Drivers

    Concrete adapters for AI models (Python/Rust clients), vector databases (Qdrant, Milvus), SQL/NoSQL storage and message brokers.

  4. Presentation / API REST · gRPC · WebSockets · SignalR

    REST and gRPC endpoints, and real-time connections.

Enterprise microservices topology

Client / Adaptive UIAPI GatewayOAuth2 · Rate-limit
Business Core ServiceC# .NET Core gRPC · RabbitMQ Relational databasePostgreSQL / SQL Server
AI Inference WorkerPython / Rust Inference · RAG Vector DatabaseQdrant / Milvus

The design triad: security, reliability and cost

Security · Zero Trust

Strict input sanitization before models (anonymization, PII masking), API key isolation, audit logging of every automated decision and Prompt Injection defense.

Reliability · Resilience

Circuit Breaker, Retry with exponential backoff and graded fallbacks: if the primary model is down, the system automatically falls back to a local SLM.

Cost · FinOps

Real-time token monitoring, semantic response caching (no duplicate inferences), budget capping and dynamic workload allocation.

Circuit Breaker

Three states

Retry · Exponential Backoff

Wait between attempts

tn = min( tcap , t0 × 2n )
attempt 1
200 ms
attempt 2
400 ms
attempt 3
800 ms
attempt 4
1.6 s
attempt 5
3.2 s
attempt 6
6.4 s
attempt 7
12.8 s

t₀ = 200 ms, t_cap = 30 s. Random jitter is added so that clients don't all retry at once.

Semantic caching: cost as the hit rate varies
C = N × (1 − h) × c
h = 0%
$15,000
h = 20%
$12,000
h = 40%
$9,000
h = 60%
$6,000

Illustrative example: 1,000,000 requests per month at $0.015 each. The real hit rate depends on how repetitive the questions are.

03 · Tech stack

The polyglot tech stack: C#, Python and Rust

No single language is the perfect solution for every component. A structured polyglot environment leverages each technology's strengths.

C# .NET Core

The Enterprise Backbone

  • RoleBusiness core, API Gateway, microservice orchestration, CQRS
  • StrengthsExtreme performance, strong typing, memory efficiency
Python

The AI Orchestration Engine

  • RoleML pipelines, PyTorch / Hugging Face, RAG, data science
  • StrengthsNative AI ecosystem, flexibility and integration speed
Rust

The Ultra-Fast System Core

  • RoleHigh-frequency tokenization, C/C++ bindings, JSON parsing
  • StrengthsZero-cost abstractions, memory safety without a Garbage Collector
C# .NET Core

Why it is the enterprise backbone

  • Open-source infrastructure cross-platform, mature, backed by the global community
  • Memory and concurrency Span<T>, Memory<T> and ValueTask: minimal allocations and high throughput
  • Clean patterns Dependency Injection, CQRS via MediatR and native gRPC for low-latency inter-service communication
Python & Rust

The role of the two specialists

  • Python remains the reference for the scientific ecosystem and Deep Learning frameworks (PyTorch, Transformers, LangChain); it is encapsulated in microservices dedicated to AI processing only
  • Rust in latency-critical system layers: with no Garbage Collector and memory safety it runs tokenization, vector preprocessing and data transformation at near-hardware speed
04 · Frontend

The frontend in the AI era: human-centricity and Adaptive UI

The frontend didn't fade with AI: it evolved beyond static interfaces with rigid forms. Technology must adapt to human cognition, not the other way around.

Generative & Adaptive UI

The interface is no longer a static grid of hard-wired elements: it is built or modified dynamically based on user intent, operating context and the system's confidence score.

Lower cognitive load

AI condenses complex information and presents only the key decisions to validate (Human-in-the-Loop), removing visual noise.

Streaming & micro-feedback

Server-Sent Events and WebSockets provide immediate visual feedback during generation, raising transparency and perceived trust.

05 · Summary

Architecture comparison

Traditional monolithic architecturePolyglot AI-native architecture
InfrastructureUnified servers / classic web appsDecoupled microservices (Cloud / On-Prem)
Tech stackSingle language (e.g. Python or Java only)C# .NET Core (business) + Python (AI) + Rust (performance)
Model managementDirect calls to external APIsOnion layer with SLM vs LLM triage and automatic fallback
User interfaceStatic pages and rigid formsAdaptive UI oriented to the human experience
Security and costsFixed pricing plans / basic criteriaFinOps, PII Masking, semantic caching and Zero Trust
DigitalSolutions

Architecture in service of people.

Bring us your project: we size it, isolate it from the domain and put it into production with the right stack for every component.