loading experience
Design · The future of AI

S1·S2 — The evolution of AI decision models: integrating fast judgment (System One) and slow reasoning (System Two)

Fast when it's enough, deep when it matters.

Inspired by Kahneman's two cognitive systems, we separate fast, calibrated judgment from slow, deliberate reasoning: the first resolves almost everything in milliseconds, the second steps in only for complex cases.

Scroll to explore
The principle

An architectural misconception: an LLM for every task.

For years, enterprise software engineering has deployed general-purpose generative LLMs for every task, including plain classification, triage and binary decisions.

Asking a 70-billion or 400-billion parameter model to generate free text just to evaluate whether a transaction is fraudulent or whether a ticket belongs to billing is the engineering equivalent of using a jet airliner to travel 100 meters.

Separate and orchestrate fast judgment and slow reasoning.

01 · Cognitive architecture

From the human mind to software engineering

The brain does not process every stimulus with deep analytical focus. On a familiar road or facing a known face we rely on System 1: immediate, energy-efficient, pattern-driven. For an equation or a chess move we engage System 2: slow, sequential, demanding.

System One · Fast Judgment

Fast, intuitive and calibrated judgment

  • ObjectiveInstant assessment and classification
  • OutputConstrained and typed: booleans, enums, scores
  • Latency10 – 50 ms
  • NatureDeterministic, probability-calibrated
  • ComputeMinimal: small SLMs and classifiers
System Two · Slow Reasoning

Slow, analytical and deliberate reasoning

  • ObjectiveComplex problems and multi-step planning
  • OutputGenerative: Chain-of-Thought, reasoning graphs
  • Latency2 – 60 s
  • NatureAutoregressive, test-time compute
  • ComputeHigh: frontier LLMs and reasoning models

p The mathematics of probability calibration

The problem with traditional generative LLMs is overconfidence: they can output a completely fabricated statement (hallucination) with apparent absolute certainty. System One models are trained or post-calibrated with dedicated techniques (for example reinforcement learning with calibration rewards, or temperature scaling) to align stated confidence with true empirical accuracy:

ℙ( Ŷ = Y | P̂ = p ) = p

If the model states p = 0.92, then across a large sample of cases with that score about 92 in 100 are correct. That makes confidence a reliable threshold on which software can act.

02 · Orchestration

The decision routing flow

In production, System One acts as a Decision Gateway and Security Router: it filters, classifies and resolves the vast majority of the workload, escalating to System Two only when genuine complexity or ambiguity is detected.

Incoming requestpayload · transaction · prompt AI Firewall / Proxysanitization & PII masking System Oneclassification + confidence
High confidence score ≥ threshold

Fast path

  • Dynamic forms generated
  • Direct backend API call
  • Automated execution
< 50 ms
Low confidence score < threshold

System Two

  • Chain-of-Thought (CoT)
  • Graph / tree search
  • Deep audit & synthesis
Human-in-the-Loop

response or action

The threshold

A decision rule, not an opinion

azione = { execute se p ≥ τ ; escalate se p < τ }

With calibrated probabilities the threshold τ is a measurable business lever: the higher it is, the fewer automatic errors and the more cases go to slow analysis.

The phases

Execution lifecycle

  1. Ingress and sanitization the AI Firewall scrubs sensitive PII and blocks Prompt Injection vectors.
  2. System One evaluation a single forward tensor pass returns intent, action and confidence.
  3. Fast path above the threshold (e.g. ≥ 85%) the action runs directly in under 50 ms.
  4. Slow path below the threshold or for complex edge cases: System Two multi-step reasoning or a human supervisor.
03 · Comparative matrix

Three model families compared

Typical latency (log scale)

System One
10-50 ms
Generative LLM
1-3 s
System Two
5-60 s
5 ms50 ms0,5 s5 s50 s

Cost per request (log scale)

System One
$0,0001
Generative LLM
$0,005
System Two
$0,03-0,15
$0,0001$0,001$0,01$0,1
System OneStandard generative LLMs (GPT-4o, Llama 3)System Two (o1, o3, Reasoning)
Primary objectiveInstant decision, classification, routingFree-form text generation and synthesisComplex problem solving and logical planning
Output typeStrict data types (Enums, Bools, JSON)Unstructured text, MarkdownChain-of-Thought + final answer
Typical latency10 – 50 ms1,0 – 3,0 s5,0 – 60,0 s
Confidence calibrationHigh, to be verified on real data (ECE)Uncalibrated / absentEmbedded in the search process
Hallucination riskLow (constrained output space, but not zero)Medium-highLow on logic, medium on raw facts
Compute costExtremely low (~$0.0001 / req)Medium (~$0.005 / req)High ($0.03 – $0.15 / req)
Hardware targetLocal edge, CPU or light SLM GPUData-center GPUsHigh-VRAM multi-GPU clusters

Orders of magnitude at full load. In the case study below, a €350/month local instance over 900,000 requests costs €0.00039 per request: real cost depends on volume.

04 · Enterprise use cases

Where fast judgment makes the difference

Real-time fraud detection

At high payment volumes per second, sub-second windows are required: a generative LLM is non-viable on latency and cost. System One inspects the transaction, evaluates behavioral graphs and outputs APPROVE / FLAG_FOR_REVIEW within tens of milliseconds with calibrated confidence.

AI guardrails and ingress firewalls

Before a query reaches downstream RAG or corporate databases, System One evaluates intent: bypass attempts, policy violations, out-of-domain requests. Threats are blocked at the perimeter.

Intelligent customer-service triage

Instead of a verbose dialogue, System One identifies the business process and renders a dynamic form for the missing parameters, routing the case to the right backend without wasting tokens.

05 · Economic sizing

Cost and financial sustainability

An enterprise processes 1,000,000 operational workflows per month for document classification and administrative routing.

Monthly technology cost comparison (€/month)

Monolithic generative LLM (all cloud)
€13,800
Hybrid System 1 + System 2
€902

Infrastructure cost reduction: −93.5%

Scenario A · Monolithic generative

Everything on a cloud LLM

cmono = 4.000 × $2.50 / 1M + 500 × $10.00 / 1M = $0.015
1,000,000 × $0.015 = $15,000 / month (≈ €13,800)

Average perceived latency: 2,5 s

Scenario B · DigitalSolutions

System One + escalation

90% (900,000) → System One local: €350 / month
c2 = 1.000 × $2.50 / 1M + 350 × $10.00 / 1M = $0.006
10% (100,000) × $0.006 = $600 ≈ €552
Chybrid = €350 + €552 = €902 (≈ $980)

Average latency: 0,9 × 30 ms + 0,1 × 2,5 s = 277 ms · for 90% of users: 30 ms

Cost as the share resolved by System One varies
Chybrid = F + N × (1 − a) × c2
a = 70%
$2,180
a = 80%
$1,580
a = 90%
$980
a = 95%
$680

F = local instance (€350 ≈ $380). Each extra percentage point resolved by System One removes $60 per month of escalation: router quality is worth money.

−93.5%on infrastructure costs
€12,898saved per month
€154,776saved per year
~277 msaverage latency, from 2.5 s

Assumptions: $1 = €0.92; GPT-4o list price ($2.50 / $10 per 1M tokens); monolith with 4,000 tokens in and 500 out per request, escalation with 1,000 in and 350 out thanks to the context trimmed by System One; estimated local instance cost. To be replaced with the client's actual data.

DigitalSolutions

The future is orchestrating the two systems.

Fast, calibrated judgment for almost everything, deep reasoning only where it matters: faster, cheaper, more reliable.