Classification, triage, entity extraction, structured parsing. On-premise or on modest hardware, latency under 50 ms.
STACK — Software design in the age of AI: architecture, polyglot tech stack and human-centricity
AI-native, not AI-bolted-on.
Designing in the age of AI isn't adding an endpoint to an LLM: it takes sizing, microservice Onion Architecture, security, reliability and cost control, a polyglot stack and an interface that adapts to people.
An AI-native architecture takes engineering, not an endpoint.
An enterprise AI-native architecture requires a rigorous approach to resource sizing, a microservice organization based on Onion Architecture, a balance between security, reliability and cost, and the selection of the polyglot tech stack best suited to combine performance, maintainability and innovation.
AI sizing and model selection
It starts by analyzing the workload to find the right balance between accuracy, latency and Total Cost of Ownership (TCO). Not every operation needs a model with hundreds of billions of parameters.
Complex synthesis, long-form content and advanced RAG (Retrieval-Augmented Generation).
Multi-step planning, logical audit and articulated problems, where inference-time compute replaces raw speed.
Local SLM
1B-8B · On-PremiseLLM / Reasoning
Private Cloud / APIstructured output
Σ Quantitative parameters
P = parameters · b = bits per weight (FP16 = 16, FP8 = 8, INT4 ≈ 4) · M_KV = KV cache for context N · 1.10 = headroom for buffers and activations.
VRAM needed with an 8k-token context (GB)
The badges show which GPU the model fits on (24, 48, 80 GB): green = fits. Quantizing from FP16 to INT4 cuts the weights by up to 75% (GGUF, AWQ, FP8 formats) with typically under 1% accuracy loss, to be verified on your own task. Llama 3.1: 8B = 32 layers, 70B = 80 layers, 8 KV heads, d = 128.
OPEX
- Immediate scalability
- Zero hardware management
- Variable costs and data-sovereignty constraints
CAPEX
- Full data governance and air-gapping
- Fixed, predictable compute cost at large volumes
- Capital cost and infrastructure management
Microservice Onion Architecture
The architecture must isolate the application domain from fast-changing AI technology. Onion Architecture keeps business logic independent from models, vector databases and cloud providers: dependencies always point inward.
- Domain Layer invariant core
Business entities, domain rules and system events. No dependency on external modules or AI frameworks.
- Application Layer Use Cases & CQRS
Use cases, workflow orchestration, commands and queries (CQRS via MediatR/C#).
- Infrastructure Layer Adapter & Drivers
Concrete adapters for AI models (Python/Rust clients), vector databases (Qdrant, Milvus), SQL/NoSQL storage and message brokers.
- Presentation / API REST · gRPC · WebSockets · SignalR
REST and gRPC endpoints, and real-time connections.
Enterprise microservices topology
The design triad: security, reliability and cost
Security · Zero Trust
Strict input sanitization before models (anonymization, PII masking), API key isolation, audit logging of every automated decision and Prompt Injection defense.
Reliability · Resilience
Circuit Breaker, Retry with exponential backoff and graded fallbacks: if the primary model is down, the system automatically falls back to a local SLM.
Cost · FinOps
Real-time token monitoring, semantic response caching (no duplicate inferences), budget capping and dynamic workload allocation.
Three states
Wait between attempts
t₀ = 200 ms, t_cap = 30 s. Random jitter is added so that clients don't all retry at once.
Illustrative example: 1,000,000 requests per month at $0.015 each. The real hit rate depends on how repetitive the questions are.
The polyglot tech stack: C#, Python and Rust
No single language is the perfect solution for every component. A structured polyglot environment leverages each technology's strengths.
The Enterprise Backbone
- RoleBusiness core, API Gateway, microservice orchestration, CQRS
- StrengthsExtreme performance, strong typing, memory efficiency
The AI Orchestration Engine
- RoleML pipelines, PyTorch / Hugging Face, RAG, data science
- StrengthsNative AI ecosystem, flexibility and integration speed
The Ultra-Fast System Core
- RoleHigh-frequency tokenization, C/C++ bindings, JSON parsing
- StrengthsZero-cost abstractions, memory safety without a Garbage Collector
Why it is the enterprise backbone
- Open-source infrastructure cross-platform, mature, backed by the global community
- Memory and concurrency Span<T>, Memory<T> and ValueTask: minimal allocations and high throughput
- Clean patterns Dependency Injection, CQRS via MediatR and native gRPC for low-latency inter-service communication
The role of the two specialists
- Python remains the reference for the scientific ecosystem and Deep Learning frameworks (PyTorch, Transformers, LangChain); it is encapsulated in microservices dedicated to AI processing only
- Rust in latency-critical system layers: with no Garbage Collector and memory safety it runs tokenization, vector preprocessing and data transformation at near-hardware speed
The frontend in the AI era: human-centricity and Adaptive UI
The frontend didn't fade with AI: it evolved beyond static interfaces with rigid forms. Technology must adapt to human cognition, not the other way around.
Generative & Adaptive UI
The interface is no longer a static grid of hard-wired elements: it is built or modified dynamically based on user intent, operating context and the system's confidence score.
Lower cognitive load
AI condenses complex information and presents only the key decisions to validate (Human-in-the-Loop), removing visual noise.
Streaming & micro-feedback
Server-Sent Events and WebSockets provide immediate visual feedback during generation, raising transparency and perceived trust.
Architecture comparison
| Traditional monolithic architecture | Polyglot AI-native architecture | |
|---|---|---|
| Infrastructure | Unified servers / classic web apps | Decoupled microservices (Cloud / On-Prem) |
| Tech stack | Single language (e.g. Python or Java only) | C# .NET Core (business) + Python (AI) + Rust (performance) |
| Model management | Direct calls to external APIs | Onion layer with SLM vs LLM triage and automatic fallback |
| User interface | Static pages and rigid forms | Adaptive UI oriented to the human experience |
| Security and costs | Fixed pricing plans / basic criteria | FinOps, PII Masking, semantic caching and Zero Trust |
Architecture in service of people.
Bring us your project: we size it, isolate it from the domain and put it into production with the right stack for every component.