Fast, intuitive and calibrated judgment
- ObjectiveInstant assessment and classification
- OutputConstrained and typed: booleans, enums, scores
- Latency10 – 50 ms
- NatureDeterministic, probability-calibrated
- ComputeMinimal: small SLMs and classifiers
Fast when it's enough, deep when it matters.
Inspired by Kahneman's two cognitive systems, we separate fast, calibrated judgment from slow, deliberate reasoning: the first resolves almost everything in milliseconds, the second steps in only for complex cases.
For years, enterprise software engineering has deployed general-purpose generative LLMs for every task, including plain classification, triage and binary decisions.
Asking a 70-billion or 400-billion parameter model to generate free text just to evaluate whether a transaction is fraudulent or whether a ticket belongs to billing is the engineering equivalent of using a jet airliner to travel 100 meters.
Separate and orchestrate fast judgment and slow reasoning.
The brain does not process every stimulus with deep analytical focus. On a familiar road or facing a known face we rely on System 1: immediate, energy-efficient, pattern-driven. For an equation or a chess move we engage System 2: slow, sequential, demanding.
The problem with traditional generative LLMs is overconfidence: they can output a completely fabricated statement (hallucination) with apparent absolute certainty. System One models are trained or post-calibrated with dedicated techniques (for example reinforcement learning with calibration rewards, or temperature scaling) to align stated confidence with true empirical accuracy:
If the model states p = 0.92, then across a large sample of cases with that score about 92 in 100 are correct. That makes confidence a reliable threshold on which software can act.
In production, System One acts as a Decision Gateway and Security Router: it filters, classifies and resolves the vast majority of the workload, escalating to System Two only when genuine complexity or ambiguity is detected.
response or action
With calibrated probabilities the threshold τ is a measurable business lever: the higher it is, the fewer automatic errors and the more cases go to slow analysis.
| System One | Standard generative LLMs (GPT-4o, Llama 3) | System Two (o1, o3, Reasoning) | |
|---|---|---|---|
| Primary objective | Instant decision, classification, routing | Free-form text generation and synthesis | Complex problem solving and logical planning |
| Output type | Strict data types (Enums, Bools, JSON) | Unstructured text, Markdown | Chain-of-Thought + final answer |
| Typical latency | 10 – 50 ms | 1,0 – 3,0 s | 5,0 – 60,0 s |
| Confidence calibration | High, to be verified on real data (ECE) | Uncalibrated / absent | Embedded in the search process |
| Hallucination risk | Low (constrained output space, but not zero) | Medium-high | Low on logic, medium on raw facts |
| Compute cost | Extremely low (~$0.0001 / req) | Medium (~$0.005 / req) | High ($0.03 – $0.15 / req) |
| Hardware target | Local edge, CPU or light SLM GPU | Data-center GPUs | High-VRAM multi-GPU clusters |
Orders of magnitude at full load. In the case study below, a €350/month local instance over 900,000 requests costs €0.00039 per request: real cost depends on volume.
At high payment volumes per second, sub-second windows are required: a generative LLM is non-viable on latency and cost. System One inspects the transaction, evaluates behavioral graphs and outputs APPROVE / FLAG_FOR_REVIEW within tens of milliseconds with calibrated confidence.
Before a query reaches downstream RAG or corporate databases, System One evaluates intent: bypass attempts, policy violations, out-of-domain requests. Threats are blocked at the perimeter.
Instead of a verbose dialogue, System One identifies the business process and renders a dynamic form for the missing parameters, routing the case to the right backend without wasting tokens.
An enterprise processes 1,000,000 operational workflows per month for document classification and administrative routing.
Infrastructure cost reduction: −93.5%
Average perceived latency: 2,5 s
Average latency: 0,9 × 30 ms + 0,1 × 2,5 s = 277 ms · for 90% of users: 30 ms
F = local instance (€350 ≈ $380). Each extra percentage point resolved by System One removes $60 per month of escalation: router quality is worth money.
Assumptions: $1 = €0.92; GPT-4o list price ($2.50 / $10 per 1M tokens); monolith with 4,000 tokens in and 500 out per request, escalation with 1,000 in and 350 out thanks to the context trimmed by System One; estimated local instance cost. To be replaced with the client's actual data.
Fast, calibrated judgment for almost everything, deep reasoning only where it matters: faster, cheaper, more reliable.
Talk to the community
See who is online, send private messages and exchange images and files (6 MB max). A free account is required.
Sign in Create an account