Skip to content
Zingaro AI
Writing · Engineering blog

Notes from the build.

How we train, serve and ship. Inference, speech, post-training, evals and air-gapped deployment, with the code and the failure modes.

Serving · p95 latencyOptimised
  • Baseline
  • Quantised
  • + Speculative
  • + Cached
Every piece

In build order.

Deep dives carry the code and the failure modes. Practical pieces are the checklists we actually use.

  1. · 5 min readDeep dive

    Reinforcement learning with verifiable rewards, on one node

    You do not need a cluster to teach a small model to follow a schema, call tools correctly and stop guessing. A practical recipe with GRPO, a reward function you can read, and the ways it goes wrong.

    • training
    • rl
    • post-training
  2. · 5 min readPractical

    Updating models inside an air-gapped network

    No outbound connection means no pulls, no telemetry and no "just download it". How we package, verify and roll a model update through a network that cannot see the internet.

    • deployment
    • sovereign
    • mlops

Short pieces for the people who own an operation: what to automate first, what a voice agent can honestly do, what sovereign really means, and where the money goes.

The blog
Book a call

Bring us the hard part.

Inference, speech, post-training, deployment. Twenty minutes is enough to say whether we can take it.

A pilot starts within 5 working days of agreed scope · Nothing upfront · No seat licences