[ Specialised service ]

LLM fine-tuning for Spanish companies

Personalised models with your data · when RAG isn't enough

We design and implement LLM fine-tuning (Llama, Mistral, Claude, GPT) for enterprise cases where RAG and prompt engineering don't give the quality, latency or cost you need. With own datasets, rigorous evaluation and deployment in European cloud or on-premise.

6–12 weeks
typical project · dataset to production
−40 to −70 %
cost per token vs base model
€30–200k
investment range
[ Honest scope ]

When it fits · when it doesn't.

We publish this list to qualify together if we fit. If not, we save you time.

✓ FITS IF…
  • High volume (100k+ executions/month) where cost per token reduction is critical
  • Very specific task where RAG + prompting doesn't reach required quality (< 90 %)
  • Need for low latency with small models (fine-tuned performs like large)
  • Very specific domain with proprietary jargon or format (specialised legal, industrial technical)
  • On-premise requirements demanding tuned open-source model
  • You have clean dataset of 500-10,000 labelled examples
✗ DOESN'T FIT IF…
  • You've started with AI recently (try RAG and prompting first)
  • You don't have a clean dataset (fine-tuning with dirty data is worse than base model)
  • Low volume (<10k executions/month) — doesn't justify investment
  • The task changes a lot from month to month (maintaining fine-tunes is expensive)
  • Looking for "general AI" — fine-tuning is for very specific tasks
[ Process ]

How we do it.

01 Week 1-2

Baseline evaluation + dataset analysis

We measure current performance with base model + RAG + prompting. If it doesn't reach, we confirm fine-tuning helps. Analyse quality and size of available dataset.

02 Week 3-5

Dataset preparation + eval set

Cleaning, deduplication, augmentation if applicable. Train/val/test split. Evaluation rubric definition with domain experts. Objective baseline.

03 Week 6-9

Training + iteration

Fine-tuning on GPUs (cloud or own). Iteration with hyperparameters and dataset variants. Blind evaluation against baseline. Regression detection.

04 Week 10-12

Deployment + observability

Deployment in your infra (on-prem with Ollama/vLLM, or managed cloud). Production quality monitoring, drift alerts, retraining plan.

[ Tech stack ]

Technologies we use.

Llama 3.3 · Mistral Large · Qwen 2.5 Open-source base (on-premise)
Claude fine-tuning (Bedrock) Fine-tune on Bedrock (limited today)
GPT-5 fine-tuning Alternative via OpenAI API
Axolotl · Unsloth · Torchtune Fine-tuning frameworks
LoRA / QLoRA Efficient parameter techniques
vLLM · TGI · Ollama Inference servers
Weights & Biases · Langfuse Tracking and evaluation
[ Verisimilar cases ]

Figures · sector · result.

Invented cases with metrics consistent with our real ranges.

01
Legal · tax firm · 60 people

Fine-tune Llama 3.3 for tax opinion writing

Model specialised in Spanish tax jurisprudence with firm's own format. On-premise to comply with professional secrecy. Prior RAG reached 82; with fine-tuning we reach 94.

82 → 94 quality · sub-2s latency on-prem
02
Fintech · credit scoring · 30k applications/month

Fine-tune small model for application classification

GPT-5 API replaced by Mistral 7B fine-tuned with 8,000 historical labelled applications. Performance equivalent to 1/3 cost per token, sub-500ms latency.

−65 % cost vs GPT-5 · same accuracy
03
Industry · automotive components

Fine-tune for technical sheet extraction

Dataset of 3,500 technical sheets with engineer-labelled fields. Fine-tune Qwen 2.5 32B on-prem. Reduces human review and accelerates cataloguing.

−40 % extraction errors vs base model
[ FAQ ]

Technical decision-maker questions.

When does fine-tuning NOT make sense?

Short answer: almost always RAG + prompt engineering + few-shot reach sufficient results with much less cost and complexity. Fine-tuning only pays off when: (1) you've exhausted RAG + prompting and don't reach required quality, (2) volume is so high that reducing cost per token with smaller model pays back in months, (3) you have clean dataset of 500+ examples representing the real task well.

If you're in early AI phase in your company (pilot, first months), 99% of the time the answer is "not yet". Fine-tuning is an optimisation, not a starting point.

How much dataset is needed for effective fine-tune?

Depends on task and base model. Ranges we see working: 500-2,000 examples for format/style/tone tasks with LoRA on large models (Llama 70B). 3,000-10,000 examples for tasks where the model needs to "learn" new knowledge or complex schemas.

More important than volume is quality. 1,000 clean expert-labelled examples beat 20,000 noisy examples. Before increasing dataset, increase quality.

If you don't have labelled data, consider synthetic data: generating examples with large model and refining with human. Moderate cost and usually works well for bootstrap.

Fine-tune on Claude/GPT or only open-source models?

Both are possible. Fine-tune on open-source (Llama, Mistral, Qwen) has advantages: total control, on-premise deployment, portability, no usage limits. Predictable operating cost.

Fine-tune on Claude via Bedrock or GPT via OpenAI API has advantages: managed infrastructure, superior base quality on many tasks, simpler update to new versions. Trade-off: operating cost depends on provider, less control, more lock-in.

Our per-case recommendation: for tasks where frontier quality is critical and volume isn't massive, fine-tune the commercial model. For high volume or on-premise/regulatory requirements, fine-tune open-source.

What happens when a new better model comes out? Do we redo the fine-tune?

Yes, each base model is different and the fine-tune is tied to the specific model. When a new better base comes out, decide: keep current fine-tune (cheap but may fall behind), redo fine-tune on new base (moderate cost, better quality), or return to prompting + RAG on new base (often new base without fine-tune already matches old with fine-tune).

Strategic recommendation: fine-tune is a moment-specific optimisation. Every 12-18 months reassess if it still makes sense or if the new base model already justifies replacing. Document dataset and evaluation well — it'll save you months when it's time to retrain.

Start with an acotated pilot?

Free 30-min diagnosis with architecture, timing and budget estimate for your specific case.