Controlled measurement of operating cost per completed task in 3 real enterprise agent scenarios (document review, claims processing, customer support with backend). With public methodology and reproducible data.
Benchmark limitations: the 3 scenarios are representative of Spanish B2B enterprise cases but not exhaustive. Token prices fluctuate (those used are June 2026). Human evaluators are specialists, not end users. Reproducing on your specific case is always more informative than blindly trusting external benchmark.
| Model | €/task | Quality /100 | Latency | Escalation % |
|---|---|---|---|---|
| Claude Opus 5 | €0.18 | 91 | 18 s | 4 % |
| Claude Sonnet 4.6 | €0.07 | 84 | 11 s | 7 % |
| Claude Haiku 4.5 | €0.02 | 68 | 4 s | 18 % |
| GPT-5 | €0.16 | 87 | 21 s | 6 % |
| GPT-5 mini | €0.03 | 72 | 7 s | 14 % |
| Gemini 2.5 Ultra | €0.12 | 82 | 16 s | 9 % |
| Gemini 2.5 Flash | €0.02 | 67 | 5 s | 19 % |
| Llama 3.3 70B (on-prem) | €0.04 * | 73 | 28 s | 15 % |
* Cost amortised on A100 GPU rented at $2.4/hr · no per-token charge. Requires constant volume to be competitive.
Quality winner: Claude Opus 5 (91/100) · Best quality/cost ratio: Claude Sonnet 4.6 (84/€0.07)
| Model | €/task | Quality /100 | Latency | Escalation % |
|---|---|---|---|---|
| Claude Opus 5 | €0.22 | 88 | 24 s | 7 % |
| Claude Sonnet 4.6 | €0.09 | 86 | 14 s | 9 % |
| Claude Haiku 4.5 | €0.03 | 75 | 6 s | 16 % |
| GPT-5 | €0.19 | 90 | 19 s | 5 % |
| GPT-5 mini | €0.04 | 82 | 8 s | 11 % |
| Gemini 2.5 Ultra | €0.15 | 85 | 17 s | 8 % |
| Gemini 2.5 Flash | €0.03 | 77 | 6 s | 14 % |
| Llama 3.3 70B (on-prem) | €0.05 * | 76 | 32 s | 13 % |
Quality winner: GPT-5 (90/100) · Best ratio: GPT-5 mini (82/€0.04) · Note: GPT-5 wins especially in structured PDF extraction and classification.
| Model | €/task | Quality /100 | Latency | Escalation % |
|---|---|---|---|---|
| Claude Opus 5 | €0.14 | 89 | 15 s | 8 % |
| Claude Sonnet 4.6 | €0.05 | 85 | 9 s | 10 % |
| Claude Haiku 4.5 | €0.01 | 78 | 3 s | 15 % |
| GPT-5 | €0.12 | 87 | 13 s | 9 % |
| GPT-5 mini | €0.02 | 80 | 5 s | 13 % |
| Gemini 2.5 Ultra | €0.10 | 84 | 12 s | 10 % |
| Gemini 2.5 Flash | €0.01 | 79 | 4 s | 14 % |
| Llama 3.3 70B (on-prem) | €0.03 * | 77 | 21 s | 15 % |
Quality winner: Claude Opus 5 (89/100) · Best for high volume: Haiku 4.5 or Gemini Flash (78-79 quality at €0.01/task is sweet spot for high-volume chat)
We repeat the 3 scenarios with the recommended architecture: Claude Opus 5 only for hard steps + Claude Haiku 4.5 for 70 % of routine decisions + prompt cache active + selective RAG. These are the final results.
| Scenario | Cost "Opus 5 everywhere" | Tiered architecture cost | Reduction | Quality maintained |
|---|---|---|---|---|
| S1 · Legal review | €0.18/task | €0.06/task | −67 % | 91 → 88 (−3 pt) |
| S2 · Claims | €0.22/task | €0.05/task | −77 % | 88 → 86 (−2 pt) |
| S3 · Customer support | €0.14/task | €0.013/task | −91 % | 89 → 84 (−5 pt) |
Practical conclusion: "using the best model for everything" is the most expensive and least efficient way to operate agents in production. Tiered architecture + cache + selective RAG drops operating cost between 67 % and 91 % with quality loss of only 2-5 points (imperceptible in most enterprise cases).
We model real cost with your specific volume in the free diagnosis. And if you already have an agent in production, we do cost audit with the same techniques from this benchmark.