Machine learning & custom LLMs

General-purpose models are impressive generalists and mediocre specialists. When the problem hinges on your data, your terminology and your edge cases, a model tuned on that data can be more accurate and cheaper to run. We build and fine-tune them - and benchmark honestly, so you know exactly what you got.

Service details

At a glance

  • Custom and fine-tuned models on your own data
  • A quality benchmark set before work starts
  • Limitations reported as plainly as strengths
  • Drift monitoring once deployed

When generic stops being good enough

Off-the-shelf models know language; they do not know your claims codes, your product taxonomy, or the strange way your industry abbreviates things. Fine-tuning on your own data closes that gap - often with less data than people expect - and can cut running costs at the same time. Your data stays under your control throughout.

Prompt, retrieve, fine-tune - in that order

There is a ladder of intervention, and cost climbs with every rung: better prompting and structured outputs, then retrieval over your own documents, then fine-tuning, then - rarely - training something custom. We start at the bottom and only climb when the benchmark says the current rung cannot reach. It is remarkable how often a well-built retrieval system over honest data beats an expensive fine-tune over messy data, and how much cheaper it is to discover that in this order.

  • Prompting and structured outputs first - often enough on their own
  • Retrieval (RAG) when the knowledge is yours and changes often
  • Fine-tuning when the task is stable and the examples exist
  • Custom models only where the economics genuinely demand them

The problems this is for

Classification and triage on your own categories. Extraction from the documents your industry writes and nobody else reads. Retrieval and question-answering over a document estate. Scoring and forecasting on your own history - the classic machine learning that predates the hype and still quietly earns its keep. If “our data, our terminology, our edge cases” describes the problem, it belongs here; if a general-purpose model with a good prompt would do, we will tell you so and charge you considerably less.

Benchmark first, then build

You cannot improve what you have not measured, so the first artefact we produce is a quality benchmark: what does good look like, on your problem, on your cases? The model is then iterated against that bar, and the final report covers the failures as candidly as the wins. A model trusted in the wrong places is worse than no model.

  • An agreed benchmark before any training
  • Iteration measured against it, release by release
  • Failure modes documented, not glossed over
See benchmark-first in a live engagement

How an engagement runs

First, the benchmark: a set of real cases, graded, that defines what good looks like on your problem - assembled with your domain experts, because they are the only people who can mark the answers. Then a baseline: the best general-purpose model, measured against that set, so every pound spent afterwards has to beat a number rather than a feeling. Then iteration up the ladder - retrieval, fine-tuning, whichever the evidence justifies - with results reported release by release, failures included. Most clients have their first honest read within a few weeks.

The benchmark outlives the model

A quietly valuable by-product of working this way: the evaluation set - your cases, graded by your experts - is an asset that outlasts any particular model. When a better model ships next year, you can measure it against your problem in an afternoon instead of re-running a pilot. When a vendor makes a claim, you can check it. Clients who have been through one engagement here stop asking “is the new model better?” and start asking “what does it score?” - which is the healthiest question in the field.

Yours: the model, the weights, the pipeline

Everything produced in the engagement is yours in the fullest sense - the fine-tuned weights, the training pipeline, the evaluation sets, the deployment configuration - documented and handed over like any other software we build. A model tuned on your data is a competitive asset, and an asset you rent is not one. If you leave, everything needed to keep the model alive leaves with you.

What it costs

A benchmark-and-baseline engagement is the usual starting point: a fixed-scope project from £8,000 that tells you what a general-purpose model already achieves on your problem and what the next rung of investment would buy. It is deliberately self-contained - if the baseline is good enough, you stop there, better informed. Ongoing model work runs as an embedded team from £4,500 a month per engineer, reported against the benchmark throughout, so you can always see what the spend bought.

Deployed like software, watched like software

Deployment comes with monitoring for drift and degradation, because model quality slips as the world changes around it. You get a model that measurably beats the generic option on your problem - and an honest account of where its edges are.

And when a model is the wrong tool

We once spent two weeks teaching a language model to read a file whose schema had never changed, because the flow was signed off before the engineers were asked. The parser that eventually replaced it took an afternoon and has never got a row wrong. The moral is baked into how we scope now: models where the input is genuinely unstructured and a good guess beats no answer - plain code everywhere else. If your problem is the second kind, expect to hear it early.

Read the feature that never needed a model

Frequently asked questions

Do we have enough data for a custom model?
Quite possibly - fine-tuning needs far less than training from scratch. We assess what you have first, and if the honest answer is no, that is what you will hear.
Will our data stay private?
Yes. Training and deployment are designed around your privacy and security requirements, and your data stays under your control.
How do we know the model is actually good?
Against the benchmark we agree up front. Results - limitations included - are reported in full, so you know where the model can be trusted and where it cannot.
What happens once it’s live?
It is monitored for drift and degradation, because quality erodes quietly when nobody is watching.
Should we fine-tune a model or use RAG?
Retrieval suits knowledge that is yours and changes often; fine-tuning suits stable tasks with plenty of examples - style, taxonomy, judgement. They also combine well. The benchmark settles it: we measure both routes against your cases and choose on evidence, not fashion.
Do you work with classic machine learning or just LLMs?
Both. Plenty of valuable problems - forecasting, scoring, anomaly detection - are classic ML with no language model required, and cheaper to run. We recommend by the shape of the problem, not by what is currently fashionable.
What does a custom model cost?
Less than most people fear, if the ladder is climbed in order. The benchmark-and-baseline project starts from £8,000 and tells you precisely whether the next rung is worth paying for - plenty of clients stop a rung earlier than they expected to.
Can models run on our own infrastructure?
Yes. Open-weight models have closed much of the gap, and for some data self-hosting is the only acceptable answer. It trades convenience for control; we lay that trade out honestly for your case and build whichever side of it you choose.
How long does a fine-tuning project take?
The benchmark and baseline usually land within a few weeks; a first fine-tuned model measured against them typically follows within a few more. The slow part is rarely the training - it is agreeing what good looks like, which is exactly why we do that part first.
Do we own the resulting model?
Yes - weights, training pipeline, evaluation sets and deployment configuration, handed over and documented like any other software we build. A model tuned on your data should be your asset, not a subscription.
Can you evaluate an AI product we are thinking of buying?
Yes - and it is one of the cheapest ways to use us. We build the benchmark from your cases and run the vendor’s product against it, which turns a sales conversation into a measurement. Vendors with good products tend not to mind; the other kind’s reaction is also informative.
What do we need to prepare before starting?
Less than you might think: examples of the task being done right - historical cases, decisions, documents - and access to the people who can judge an answer. The first weeks are designed to work with what exists; if something genuinely blocking is missing, you will hear it in week one, not month three.

Ready to talk through Machine learning & custom LLMs?

Book a free 30-minute consultation with a senior engineer to see how we can help.