Every operations leader we talk to has the same opening line: "We know we should be doing something with AI. We just do not know where to start." Then they show us a list. Sometimes it is a slide with forty ideas on it. Sometimes it is a spreadsheet a consultant left behind. The ideas are usually fine. The problem is that nobody has a way to choose between them, so nothing starts.
This is the method we use in the first week of every engagement. It is not clever. It works because it forces one decision instead of forty.
Make the real list, not the exciting one
The exciting list is the one with "AI copilot for the sales team" and "predict churn" on it. The real list is the one you get by asking a different question: what do people in this business do today by following a script or a template?
Walk through a normal week and write those down. Inbound calls that follow a flow. Documents that get read, checked and keyed into a system. Emails that get sorted into folders. Records that get tidied on a Friday afternoon. Reports assembled from the same four sources every Monday.
The real list is longer than the exciting one, and it is where the money is. A job that is done by script is a job a model can be trained on and measured against. A job that needs judgement every time is not a first job.
The three tests
Every candidate on the real list gets three tests. A job has to pass all three to be the first one.
1. There is a number
Not "it would save time". A number that exists today, or could be counted by Friday. Calls answered inside thirty seconds. Invoices keyed per day. Hours between a document arriving and a decision being made. Percentage of records with a missing field.
If there is no number, you cannot tell whether the pilot worked, and you will end up arguing about whether it "feels better". Pick something else.
2. It has edges
A first job starts somewhere and ends somewhere, and it touches one system of record. A call comes in and ends with a booking in the calendar. An invoice arrives and ends as a line in the ledger. A ticket opens and ends with a category and an owner.
Jobs without edges, such as "handle whatever customers ask", are the ones that produce the demos everyone loves and the production systems nobody trusts. Give the model edges first. Widen them later.
3. A wrong answer is recoverable
Somebody has to be able to catch a mistake before it costs money or trust. That means there is a review step, or an undo, or a hand-off to a person when the model is not sure. If a wrong answer goes straight out to a customer or a regulator with nobody in between, the job is not a first job, whatever the number says.
The first job is not the biggest one. It is the one you can measure, bound and recover from. Do that one well and the second one gets much easier to sell inside your own company.
Rank by pain, not by novelty
Once you have the jobs that pass all three tests, rank them by how much they hurt. Not by how impressive the demo would be. The job that keeps a manager late twice a week beats the one that would look good on a conference slide.
Pain is also what makes a pilot succeed politically. When the people doing the job today are relieved rather than threatened, they help. They tell you about the edge cases. They review the queue. They become the reason it works.
What a first job usually looks like
For the record, these are the shapes that keep passing the tests, across very different industries.
| Shape | Example | The number |
|---|---|---|
| A call with a flow | Appointment changes, order status, payment reminders | Calls handled without a transfer |
| A document with a destination | Invoices, claims, applications, KYC packs | Documents keyed per day, error rate |
| A queue with categories | Support tickets, shared inboxes, leads | Time to first correct routing |
| A record that drifts | CRM hygiene, product data, supplier lists | Records complete and current |
| A report from fixed sources | Weekly operations pack, daily exceptions | Hours to produce, errors found later |
What to avoid as a first job
Three things, and they are the same three every time.
- Open-ended chat. "Ask our AI anything" has no number, no edges and no recovery path. It is a feature, not a job.
- Anything without an owner. If nobody is accountable for the number today, nobody will be accountable for it after the pilot either.
- Anything that needs a policy that does not exist yet. A model cannot follow a refund rule the business has never written down. Write the rule first. Then automate it.
Write the number down before anything starts
Before we build a pilot, the number goes at the top of a one-page document, with today's value next to it. Everyone signs it. At the end of the pilot we put the new value next to the old one. That is the whole review.
It sounds obvious. Most AI projects skip it, and that is why most AI projects are hard to judge and easy to quietly stop. Bring us the list, the long one, and we will do the ranking with you on the first call.
