loading experience
Design · Measuring AI

ROI — The failure of improvised AI: how to engineer and size an Artificial Intelligence process and save costs (and jobs)

AI is not magic: it has a precise cost.

Why cut staff when you can engineer the code, spend up to 80% less on cloud resources and get an incomparably better process?

Scroll to explore
The problem

The failure of improvised AI.

Over the last three years we have witnessed an unprecedented gold rush: hundreds of startups and enterprise projects born with the promise to "revolutionize an industry with Generative AI". The reality of the balance sheets tells a different story: an impressive number of them are failing or in serious financial distress.

The reason is neither a lack of market nor the quality of the models. It is the total absence of architecture and engineering sizing. Many teams treat AI as magic: they plug in an SDK, make a REST call to the provider of the day, pass the entire context as input and consider the job done.

We have analyzed real cases where this superficiality produced monthly API bills of tens of thousands of euros for trivial tasks. The most tragic consequence: companies that cut staff or reduced salaries to cover those costs.

Generative AI is not magic: it is a software component with a precise computational cost.

01 · The trap

The "REST call" trap

Naive chatbot versus multi-agent orchestrator: to see where the waste hides, we compare the amateur approach with the engineering one.

Naive approach
User Monolithic REST call Giant cloud LLM Huge cost on every message
Engineered approach · DigitalSolutions
User AI Proxy / Firewall Router / Classifier local SLM
Flow Manager / Classic algorithm$0.00
RAG / GraphRAG mid-size or local model$
Generative LLM only if indispensable$$$
Aggregator & output
Example A

The "naive" chatbot: the token black hole

In a customer-service bot, on every message the REST call sends the LLM a monumental system prompt with all company rules, the entire history growing every turn, and company documents pasted into the context.

  • At the tenth turn: 8,000 input tokens + 500 output tokens per answer
  • Per message: 8,000 × $2.50/1M + 500 × $10/1M = $0.025
  • × 10,000 messages a day = $250 per day (about $7,500 per month)
  • Answers take 4-6 seconds
Example B

The modular multi-agent orchestrator

A Semantic Router (lightweight classifier or local SLM) analyzes the intent using very few resources, or zero API cost if local:

  • Shipping status? A Flow Manager or a direct call to the logistics database. AI cost: $0.00
  • Frequent question? Optimized RAG with a compact model.
  • Complex reasoning? Only here do we escalate to a specialized agent with the suitable LLM.

Context Isolation: each agent receives only the information needed for its micro-task, keeping prompts compact.

02 · Token

Prompt engineering for financial efficiency

The prompt is not narrative text, it is a machine instruction. Writing the best prompt doesn't only mean getting the right answer, but optimizing information density to minimize tokens.

Prompt Caching (KV Cache Hits)

Place the static part of the prompt (system instructions, schemas) at the start of the message to exploit provider-side or local caching.

Savings of up to 50-80% on repeated input tokens and a drastic latency reduction.

Structured Outputs (JSON Schema)

Force the model to answer with a strict, minimal JSON schema instead of verbose prose.

Cuts output tokens by 40-60% and removes ambiguity of interpretation.

Context pruning & windowing

Keep a sliding window over the history and extract only key entities or incremental summaries instead of sending the whole chat history.

Avoids the exponential growth of per-message cost in long threads.

Semantic Compression

Remove stopwords, redundancy and decorative formatting from RAG-retrieved sources before injecting them into the context.

Reduces the size of the injected context by 30-50%.

Golden rule

In commercial models, output tokens cost 3 to 4 times more than input tokens and are the primary drivers of latency (Time to First Token and Tokens Per Second). Cutting the model's verbosity saves money and speeds up the user experience.

03 · Granular cost

Measuring the cost of every single component

There is no generic "AI cost" line item: there is a sum of transparent components.

Total cost per call
Ctot = Cembed + Cvector + Cfirewall + Cin + Cout + Cinfra
Cin = Tin × Pin Cout = Tout × Pout
Cembed

Embedding

Vectors for the user query via embedding models (e.g. $0.00002 per 1k tokens).

Cvector

Vector / graph DB

Read and search on the memory database (Qdrant, Pinecone, Neo4j), based on RCU/WCU or memory/CPU usage.

Cfirewall

AI Proxy & Guardrails

Computational cost of the security checks (PII Redaction, Prompt Injection check).

Cin

Input tokens

Input tokens × input price per token.

Cout

Output tokens

Output tokens × output price per token.

Cinfra

Orchestration

Share of the serverless or container infrastructure (RAM/vCPU) running the business logic.

04 · Sizing

A practical guide to sizing

Sizing a process means moving from abstract estimates to math applied to the workload.

1 Load metrics

NreqRequests per day / month
RPSpeakRequests per second at peak
SLAlat < 1,5 sMaximum response time

2 Model selection matrix (SLM vs LLM)

3 The cascade pattern: the 80/20 rule

In real processes 80% of requests are of low-to-medium complexity, while only 20% require advanced reasoning. Sizing correctly means implementing a Fallback / Cascade Pattern.

Costs for a call of 1,500 input and 500 output tokens, at list price: GPT-4o-mini ($0.15 / $0.60 per 1M tokens) versus GPT-4o ($2.50 / $10).

Blended cost per request
Cmix = 0.8 × $0.000525 + 0.2 × $0.00875 = $0.00217
All on Frontier
$0.00875
80/20 cascade
$0.00217

Reduction in average cost per request: −75.2%

05 · Case study

Economic model and human impact

We show with math how correct sizing saves the company's bottom line and avoids pointless layoffs.

100,000document cases per month
10operators, average cost €35,000/year each
€29,167per month (€350,000 per year) without AI
45 sof human work per case (triage and data extraction)
Sizing check: operators needed
nop = ⌈ Nreq × t / (3.600 × H) ⌉ = ⌈ 100,000 × 45 / (3.600 × 125) ⌉ = 10

H = 125 productive hours per month per operator (about 78% of 160 hours). By pre-filling the case, AI cuts t by 85%: from 45 s to 6.75 s.

Total monthly process cost (€/month)

Naive approach (API only, unoptimized cloud LLM)
€34,960
10 operators without AI
€29,167
Engineered: AI + 2 operators
€6,810

The naive approach costs more than the 10 operators; the engineered process costs 76.7% less than the manual process.

Option A · Naive approach

The startup failure

An agentic pipeline sends the whole 15,000-token document to a high-end cloud LLM at every step (extraction, classification, verification, summary…): 8 calls per case.

Ccall = 15,000 × $2.50 / 1M + 1,000 × $10.00 / 1M = $0.0475
Ccase = 8 × $0.0475 = $0.38
100,000 × $0.38 = $38,000 / month (≈ €34,960)

The company spends more on APIs than it used to spend on salaries (€34,960 vs €29,167). If it cuts staff to compensate, quality collapses: the unoptimized system hallucinates and nobody checks.

Option B · DigitalSolutions

Human Augmentation

  1. Local pre-processing OCR + heuristic extractor isolate the relevant section: input from 15,000 to 1,500 tokens.
  2. Smart routing 70,000 standard cases on a local SLM on a dedicated GPU (€600/month); 30,000 complex cases on a mid-tier cloud model with Prompt Caching, 3 calls of 1,500 tokens in and 500 out.
  3. Human-in-the-Loop The AI pre-fills the case and the operator validates: from 45 to 6.75 seconds (−85%).

Needed: 2 operators out of 10. The other 8 are not laid off: they are reassigned to higher-value work, business development and complex quality control.

Monthly cost of the engineered process, line by line

Cembed100,000 × 1,500 tokens × $0,02/1M × 0.92€2.76
CvectorManaged vector database (estimate)€100
CfirewallAI Proxy / guardrails on CPU (estimate)€80
Cin + Cout30,000 × 3 × $0.000525 × 0.92€43.47
CinfraDedicated GPU for the SLM €600 + orchestration €150€750
CtechTotal technology€976
Cop2 × €2,917€5,833
CprocessTotal process€6,810

Assumptions: $1 = €0.92; 125 productive hours/month; public provider list prices (GPT-4o, GPT-4o-mini, embedding small) to be rechecked at the time of reading; fixed costs for vector DB, firewall and orchestration are estimates to be replaced with the client's actual quotes.

−97.2%on AI costs: from €34,960 to €976 per month
€6,810total monthly process cost, versus €29,167
€22,357direct net savings per month (−76.7%)
€268,285per year, without laying anyone off
Conclusions

Engineering as respect for resources and for people.

AI failures are not caused by the limits of the models, but by the lack of engineering culture. AI is an infrastructure component: it must be analyzed, segmented, token-optimized, data-protected and cost-sized. You don't need to cut jobs to be profitable and innovative: just write better code, choose the right architecture and eliminate technological waste.