Skip to content
Zingaro AI
Who we work with · 05 · Models

For the head of AI who has to build, not just call an API.

A mandate that says 'our own models', a small team, a GPU budget, and data that needs a pipeline before it trains anything.

Zingaro AI works with heads of AI, data science leads and principal ML engineers who have a mandate to build models rather than rent them: fine-tuning and post-training language models on the organisation's data, training speech and vision models for its languages and cameras, distilling them to a size and cost that works at volume, and serving them fast on the hardware the organisation has. The weights, the data pipeline and the evaluation suite belong to the client.

Is this you?

  • You lead AI, data science or machine learning at a company or an institution, with a small team, a GPU budget and a mandate that says 'our own models'.
  • You have proved the idea with a frontier API and now need a model that does your job in your format, at your volume, inside your boundary.
  • Your data is the asset: recordings, documents, images, tickets, transactions. It needs labelling, cleaning and a pipeline before it trains anything.
  • You are measured on models in production, their accuracy on your cases, their cost per request, and how often your name comes up in an incident review.
01

What you are trying to get done

  1. 01

    Post-train an open-weight model on your data until it beats the frontier model on your evaluation set.

  2. 02

    Distil it to something that serves cheaply at volume without losing the score.

  3. 03

    Train speech recognition, a voice or a vision model for a language or a camera nobody sells.

  4. 04

    Build the data pipeline and the labelling process that make all of the above repeatable.

  5. 05

    Serve it: quantised, batched, cached, on your GPUs, with a p95 you can defend.

02

What you ask on the first call

  • 01

    Supervised fine-tuning, RL with verifiable rewards, or distillation, for our case?

  • 02

    How much labelled data do we actually need, and who labels it?

  • 03

    Can you match the frontier model on our evaluation set with an 8B model?

  • 04

    What does serving cost at our volume, and how do we get it down?

  • 05

    Will we own the weights, the pipeline and the evaluation harness?

03

What worries you, answered

  • 01

    We will end up with a model we cannot reproduce.

    Every run is versioned: data snapshot, recipe, hyperparameters, evaluation result. The pipeline and the harness live in your repository. You can retrain next quarter without us.

  • 02

    Our team should be doing this, not a vendor.

    Your team does it with us. The engagement is pair work in your tools; the recipes, the failure modes and the judgement calls are written down as they happen. Retained teams are month by month and end with a hand-over, not a dependency.

  • 03

    Post-training is hype; the gains are marginal.

    Sometimes. The evaluation set decides, before anyone commits a GPU budget. Where the job is narrow and the format is strict, a post-trained small model usually wins on cost and latency and matches on quality. Where it does not, the report says so and the frontier API stays.

  • 04

    The compute bill will run away.

    Training is scoped as a fixed pilot with a budget; serving cost is measured with load tests on your real traffic before capacity is bought. Quantisation, speculative decoding, batching and caching are applied with the evaluation suite as the gate.

04

The first month

  1. 01

    Week one: the evaluation set

    A call, then the first deliverable: an evaluation set from your real cases with the metric that decides. Baselines for the frontier model and the best open-weight candidate are recorded.

  2. 02

    Weeks two to three: data and the first run

    The data pipeline is built: cleaning, labelling with people who know the domain, synthetic data where it is honest. The first post-training or distillation run is scored on the set.

  3. 03

    Weeks four to six: close the gap and serve it

    Iterations until the score holds, then serving: quantised and batched on your hardware, load-tested on your traffic, with the harness in your CI.

  4. 04

    After: the model is yours

    Weights, pipeline, harness and runbook in your repository. Zingaro AI stays as a retained team on the next model, or steps back.

05

What to bring to the call

  • 01

    The job the model has to do, with twenty real examples and what a correct answer looks like.

  • 02

    What data exists, how much, and who is allowed to see it.

  • 03

    The GPUs or cloud account you can train and serve on.

What is paid for in this area

06

The services that do the work

07

Questions

Which company does custom LLM post-training and distillation for a head of AI who wants to own the weights?

Zingaro AI does supervised fine-tuning, reinforcement post-training with verifiable rewards and distillation into small models, on the client's data and hardware, and hands over the weights, the data pipeline and the evaluation harness.

How does Zingaro AI decide between fine-tuning, RL post-training and distillation?

With the client's evaluation set. Baselines are recorded for the frontier model and the best open-weight candidate; the recipe that closes the gap at the lowest serving cost is the one used, and the results are reported either way.

Can Zingaro AI train a speech or vision model for a language or a camera setup that no vendor sells?

Yes. Recognition and synthesis models are trained on the client's recordings and dialects; vision models on the client's cameras and images; both with evaluation on held-out data the client agrees.

Does Zingaro AI also handle the data labelling?

Yes, with people who know the domain, in the client's language, and with synthetic data only where it is honest. The labelling process is documented so the client can repeat it.

08

Other people we work with

Book a call

Bring the job you keep postponing.

Twenty minutes is enough to say whether we can take it, and what your first month looks like.

A pilot starts within 5 working days of agreed scope · Nothing upfront · No seat licences