Part 2 — Speak Both Languages
The Metric-to-KPI Translator
If you cannot draw the chain from a technical metric to a business number, you don’t have a business case yet. You have an interesting technical project.
This is the piece most AI pitches skip, and it’s the one that decides whether a proposal gets approved or quietly shelved. Everything below exists to force one sentence into existence before any money moves.
Two languages, and why they don’t translate on their own
The technical metrics
| Metric | What it actually answers | Where it misleads you |
|---|---|---|
| Accuracy | Of all the predictions made, what share were right? | On rare events, it looks excellent while doing nothing. If 1% of transactions are fraudulent, a model that says “not fraud” to everything is 99% accurate and catches zero fraud. |
| Precision | When it says yes, how often is it right? | Can be pushed very high by making the model cautious — which quietly lets real cases through. High precision alone tells you nothing about what you’re missing. |
| Recall | Of all the real cases out there, how many did it find? | Can be pushed very high by making the model aggressive — which floods the process with false alarms people learn to ignore. |
| F1-score | A single number balancing precision and recall. | Useful shorthand, but it hides which of the two is actually weak — and in most businesses, one of them matters far more than the other. |
| Latency | How long it takes to respond. | Not a quality metric at all — increasingly an adoption one. Past a certain delay, the interaction stops feeling fluid and people stop using the tool. |
Which of precision and recall should you care about? That’s a business judgment, not a technical one, and it comes down to which mistake costs you more. Where a false alarm is expensive — blocking a legitimate customer, escalating a healthy account, stopping a production line — you care about precision. Where a missed case is expensive — a fraud that goes through, a customer who leaves without warning, a defect that ships — you care about recall. Decide which sentence describes your process before anyone shows you a number.
The business KPIs
No benchmark figures here, deliberately — percentage-improvement claims in vendor decks go stale fast and vary wildly by sector and starting point. The only baseline that survives a conversation with your CFO is your own, measured before you start.
| Area | KPI | AI typically moves it by |
|---|---|---|
| Sales & marketing | Conversion rate, churn, LTV, CAC | Better targeting, earlier risk flags, less wasted effort on leads that were never going to close |
| Operations | Cycle time, cost per transaction, error rate | Removing or shortening manual review steps; catching mistakes at the point they’re made |
| Customer support | Response time, tickets per agent, CSAT | Instant handling of routine cases, so people work on what actually needs a human |
The impact chain
A technical metric on its own means nothing. What matters is how a change in that number changes a behavior, a decision, or a process — and how that change lands on a number the business already tracks. Write it as one sentence, in this shape:
IF the model achieves [technical performance] THEN we can [do this differently — name the person whose work changes] WHICH IMPROVES [the business KPI, from this baseline to this target].
Before you present a chain, run it through three tests:
- The name test — can you name the actual person whose day-to-day work changes in the middle step? If not, the chain is theoretical. Go talk to whoever does that work today.
- The baseline test — does the business KPI have a number today, agreed by whoever owns it? If not, you can’t prove improvement later — establish the baseline before launch, not after.
- The counterfactual test — could that KPI have moved on its own, without this project? Seasonality, a pricing change, another initiative. Say so openly, before someone else raises it in the meeting.
Two worked chains
A back-office example — AP invoice matching in a 60-person distribution business. A finance team manually matches supplier invoices against purchase orders and delivery notes; roughly one in six needs a manual chase for a mismatch.
- Technical metric: recall on mismatched line items, 88%, with a confidence score per match.
- What changes in the process: only low-confidence matches reach the accounts payable clerk — routine matches clear automatically, and the clerk’s day shifts from checking everything to resolving the exceptions.
- Business KPI it lands on: cost per invoice processed, and days-to-close on month-end reconciliation. Baseline 45 seconds and 4.10€ per invoice fully loaded; target under half that.
A sales example — lead scoring for a B2B SaaS sales team.
- Technical metric: precision 85%, recall 70%.
- What changes in the process: reps trust the ranking and stop spending time on low-probability leads, concentrating effort on the ones flagged high quality.
- Business KPI it lands on: closing rate, and therefore revenue. Baseline 15%, target 22%.
Six more worked chains
The same shape applies to any AI use case: a technical metric, a change in the process, and the business KPI it moves.
- Product recommendations (e-commerce) — Technical metric: precision on “would buy” predictions. Process change: fewer irrelevant recommendations shown, more relevant ones surfaced first. Business KPI: conversion rate on recommended items, baseline vs. target.
- Fraud detection (payments) — Technical metric: recall on fraudulent transactions. Process change: more fraud caught before settlement, with a human review step on flagged cases. Business KPI: fraud loss as a percentage of revenue.
- Support ticket triage — Technical metric: classification accuracy on ticket category. Process change: tickets route to the right team on first pass instead of bouncing. Business KPI: first-response time and ticket reassignment rate.
- Document extraction (contracts/invoices) — Technical metric: field-level extraction accuracy. Process change: fewer manual re-keying steps, exceptions routed to a human. Business KPI: processing cost per document.
- Churn prediction (subscription business) — Technical metric: recall on customers who actually churn. Process change: earlier, targeted retention outreach on flagged accounts. Business KPI: churn rate, baseline vs. target.
- Predictive maintenance — Technical metric: precision on failure predictions. Process change: maintenance shifts from fixed schedule to flagged equipment. Business KPI: downtime hours per month, maintenance cost per unit.
Once you can write the sentence for your own project, you have the pattern for any of them.
The sentence to walk into the room with
Not “the model is accurate.” Instead:
“The model reaches this level of performance, which lets us do this differently, which moves this KPI from here to here — and here’s who agreed the starting number.”
That’s a complete business case in one breath. It’s the difference between a technical update and an investment decision.