HomeBlogHow to Measure Vendor Reliability: A Simple Scoring System
Operations

How to Measure Vendor Reliability: A Simple Scoring System

Gut feel says the vendor is “usually fine”. The record says three late deliveries this month. A simple way to score vendors on delivery, accuracy and communication, using data you already have.

10 min read

Every owner already has a mental ranking of their vendors. The trouble is that mental rankings get built from the last two weeks and the loudest incident, so the vendor who quietly delivers on time for a year loses to the one who apologises beautifully. You measure vendor reliability for one reason: to let the whole record vote instead of just the recent memory.

It’s a small piece of work. Three measures, one number per vendor per month, computed from data you already produce every time you place an order. What you get back is the ability to walk into a price negotiation with facts, and to know, before it hurts, which vendor shouldn’t get the order that matters.

Why gut feel fails

Gut feel isn’t stupid. It’s a genuinely good summary of your experience, distorted in two predictable ways.

The first is recency. Ask yourself right now who your most reliable vendor is and the answer will be shaped almost entirely by the last fortnight. A vendor who was excellent for eleven months and shaky for three weeks feels unreliable. A vendor who was awful all year and flawless this month feels like he’s turned a corner. Neither read is right, and both are the read you’ll act on.

The second is the squeaky wheel. Vendors who communicate well about their failures score better in your head than vendors who quietly succeed. The one who calls at 7am to say the truck has broken down feels like a partner. The one whose truck simply arrives, every Tuesday, for two years, generates no memories at all. Attention follows drama, and reliability is the absence of drama.

There’s a third distortion nobody likes admitting to: the relationship. You like the salesman. He asks about your family. He’s been coming round for six years. None of that is irrelevant, and goodwill has real value the evening you need an emergency delivery at 6pm, but it should be a factor you apply consciously on top of the record rather than a fog that hides it.

Scoring doesn’t replace judgement. It makes sure judgement is applied to the whole year rather than to whatever happened on Thursday.

Three things worth measuring

You could measure fifteen things. You will maintain three. Pick these three, because each one maps to a distinct way a vendor can cost you money, and each comes out of records you already keep.

1. On-time delivery rate

Of the orders you placed, what share arrived by the date the vendor committed to?

The subtlety is whose date. Measure against the date the vendor confirmed, not the date you were hoping for. If you asked for Tuesday and he said Thursday and it came Thursday, that’s on time. What you have there is a lead-time problem, not a reliability problem, and the two need different fixes. Confusing them produces scores that make good vendors look bad.

Where no date was ever committed, the order drops out of the denominator. It counts somewhere else though, as measure three will show.

2. Order accuracy

Of the orders that arrived, what share arrived complete and correct, meaning right items, right quantities, undamaged?

This measure comes straight out of your delivery records. Every time someone counts a delivery against the order and writes down what actually came, they’re generating an accuracy data point. If you’re not recording received quantities per item then this measure simply isn’t available to you, and fixing that is worth doing for its own sake long before you care about scoring anything.

Count an order as inaccurate if anything was wrong: short, over, wrong grade, damaged, substituted without asking. Don’t try to weight severity at this stage. A yes or no per order is maintainable. A severity scale is not.

3. Confirmation speed

How long after you place an order does the vendor confirm it, with a yes, a price and a date?

This one surprises people, because it feels like a courtesy rather than a cost. It’s a cost. An order sitting unconfirmed for three days is an order you can’t plan around, and the vendor who takes three days to say yes is usually the same vendor who’ll take three days to tell you there’s a problem. Confirmation speed is the earliest signal you get of everything else, which makes it the most useful leading indicator on the list.

Measure it crudely. Same day, next day, or slower. Or just count the share of orders confirmed inside 24 hours. Precision adds nothing here.

Why stop at three

Price isn’t on the list, deliberately. You already track price obsessively and you don’t need a score to tell you who’s cheaper. The value of a reliability score is that it prices the things that never appear on an invoice: the emergency top-up bought at retail, the job postponed, the hour someone spent chasing a delivery. Keep the score about reliability and let price sit next to it as a separate fact.

Invoice accuracy makes a reasonable fourth if you already check invoices against orders and deliveries, because the data appears with no extra work. Add it only if that habit is already running.

A simple scoring method

Per order, record three yes/no facts at the moment you’re already touching the order anyway. Was it confirmed within 24 hours? That gets recorded when the vendor replies. Did it arrive by the committed date, and did it arrive complete and correct? Both get recorded when the delivery is received and counted, by the person doing the counting.

That’s the entire data collection burden. Three ticks per order, each entered by whoever was there. No vendor review meeting, no form to fill in at month end.

Once a month, per vendor, turn the ticks into three percentages. If you want a single number to sort by, average the three. An unweighted average is fine and much easier to explain than a weighting scheme your team will argue about for an hour.

What it looks like

Three vendors supplying the same small business over one month. Names are illustrative.

VendorOrdersConfirmed in 24hOn timeComplete & correctScore
Northline Traders1292%92%100%95%
Blue Harbour Supplies1050%60%90%67%
Ridgeway Wholesale8100%88%63%84%

Read across the rows rather than down the score column. The composite is only there for sorting, and these three vendors have three completely different problems.

Northline is doing what a good vendor does, which is nothing memorable. One late delivery in twelve, everything complete. You probably underrate them precisely because they never generate an incident worth remembering.

Blue Harbour is slow at both ends: half the orders sat unconfirmed and four in ten arrived late. That’s a planning problem more than a goods problem, since what turns up is broadly right, you just can’t rely on when. If they’re your cheapest source, the honest question is whether the saving covers the two emergency purchases you made last month.

Ridgeway is the interesting one. They confirm instantly and mostly deliver on time, but three of eight deliveries were short or wrong. Fast, responsive, sloppy in the warehouse. That’s a very fixable problem and exactly the kind you can raise productively, because the vendor is clearly paying attention. The failure is in picking, not in attitude.

Notice too that Ridgeway’s composite (84%) beats Blue Harbour’s (67%) while hiding the worst single number in the table. Composites do that. Use the score to sort the list and use the columns to decide what to say.

Using the score

A common mistake is treating a score as a verdict. It’s the opening line of a conversation, and it’s most valuable with vendors you intend to keep.

For conversations

“You’ve been unreliable lately” gets a defensive shrug. “Four of your last ten arrived late, what changed?” gets an actual answer, because it’s specific, checkable and asks a real question. Very often the answer is useful: a driver left, a route changed, they moved warehouses, your order goes out on a run that now leaves at a different time. Half of those are fixable by shifting when you place the order.

Deliver it as an observation, not an accusation. You’re not building a case for termination. You’re telling a supplier something about his own business that he probably doesn’t know, because nobody else is counting either.

For allocation

The most practical use of a score has nothing to do with vendor management. It’s deciding who gets the order that must not fail. The wedding function on Saturday, the site pour on Monday morning, the stock for festival week: those go to the highest accuracy and on-time vendor, even at a worse price. Routine, forgiving, easily-replaced orders can go to whoever is cheapest.

Most businesses do this instinctively and get it wrong roughly as often as memory is wrong.

For negotiation

When a vendor asks for a price increase, the record is your side of the table. If the service has been excellent, say so and pay: good suppliers are worth keeping, and telling them why you value them costs nothing. If the record is patchy, the response writes itself. You’re asking for more while a third of deliveries arrive short. Fix the second and let’s talk about the first.

What not to do with it

Don’t fire a vendor on one bad month, especially not on a handful of orders. Don’t rank vendors who supply different categories against each other, because a perishables supplier and a hardware supplier aren’t comparable and the composite will mislead you. And don’t publish an internal leaderboard. That turns a management tool into a scoreboard people game.

Where the data comes from

The part that makes this achievable is that you aren’t collecting new data. You’re writing down data you already generate.

Every order you place has a date. Every confirmation arrives with a timestamp, sitting in a chat thread. Every delivery gets counted by somebody. The only thing missing, in most businesses, is that those facts live in three different places, a phone and a delivery book and someone’s memory, and never get brought together per vendor.

So the prerequisite for scoring isn’t a scoring system. It’s tracking orders through their stages with real dates. Do that and the score is arithmetic. Skip it and no scoring method in the world will save you, because there’s nothing to compute from.

If your orders already live in one place with dates and quantities, this can be automatic. OrderBookApp computes a reliability score per vendor from exactly this data: confirmation times, promised versus actual delivery dates, ordered versus received quantities. The month-end arithmetic then happens on its own. A notebook works too, mind you. A tally of three ticks per order, added up monthly, beats a sophisticated system nobody feeds.

Starting from nothing

If you have no historical data, don’t try to reconstruct it. Start today, forward-looking, with your top five vendors by spend. In two months you’ll have enough to see a shape. In six you’ll have enough to be confident. That’s a short wait for information you’ll use for years.

Keep reading

Frequently asked questions

How many orders before the score means anything?
Roughly ten orders with a vendor before you draw any conclusion, and more before you act on a small difference. Below that, one bad week swings the percentage wildly: two late deliveries out of three orders reads as 33% on time, which sounds catastrophic and means almost nothing. For low-volume vendors, look at a rolling three or six months instead of a single month, and treat the direction of travel as more informative than the number itself.
Should I tell vendors their score?
Share the underlying facts, not the number. "Four of your last ten arrived late" is specific, verifiable and invites an explanation. "You scored 67%" invites an argument about the scoring method, which is a conversation with no useful outcome. The exception is a long-standing vendor you are actively trying to develop, where showing the trend over several months can be genuinely motivating, and even then you lead with the counts.
What is a good on-time rate?
It depends entirely on the category, and comparing across categories will mislead you. A daily perishables supplier on a fixed route should be near-perfect, because the delivery is routine. An importer, a fabricator, or anyone whose goods depend on production or shipping has a genuinely lower ceiling, and a rate that would be alarming for the first is normal for the second. What matters is the trend for that vendor, in that category, over time. A vendor slipping from his own normal tells you far more than any benchmark.
What if a vendor is unreliable but the cheapest?
Then price the unreliability instead of ignoring it. Add up what the failures actually cost last quarter: emergency purchases at retail rates, staff time spent chasing, jobs delayed, customers disappointed. Compare that with the saving. Often the cheap vendor is still cheaper, and the right answer is to keep them for routine orders while giving critical orders to someone else. Sometimes the arithmetic goes the other way and you can finally say so with numbers rather than frustration.
#Vendors#Reliability#Operations
Share