Approach Engagements Work Writing About Contact Book a discovery call Chat on WhatsApp
Taking new engagements for Q3

hello@onetouchgrade.com

Insight

Scoring AI use cases before you buy anything

Score every candidate AI use case on five axes before choosing a tool: hours reclaimed, tolerance for error, run cost at real volume, data sensitivity, and time to a measurable result. Roughly a third of candidates do not survive the scoring, which is the point of doing it.

The failure is never the model

AI projects that quietly die share a shape. Someone saw a demo. A tool was bought under time pressure. It was pointed at a workflow nobody had mapped. And no one recorded what performance looked like beforehand.

Twelve months later there is a subscription nobody can defend and nobody can cancel, because cancelling would mean admitting it never worked, and without a baseline there is no evidence either way. Not one of those failures is a modelling problem. They are decision problems, made months before any model was involved.

The fix is unglamorous: score the candidates before choosing the tool.

The five axes

Every candidate use case gets scored on the same five dimensions, against your workflows rather than a vendor's case study.

Scoring axes for an AI use case
Axis The question Kills the use case when
Hours reclaimed How much time does this actually consume per month, measured? The honest measurement is small
Error tolerance What does a wrong answer cost, and who catches it? Wrong answers are expensive and nobody checks
Run cost What does this cost at real volume, including a spike? Cost per unit exceeds the value of the unit
Data sensitivity Can this data legally and safely leave your systems? It cannot, and no local option is viable
Time to signal How soon will we know whether it worked? The answer is more than a quarter

These are not weighted equally. Data sensitivity and error tolerance are gates: failing either removes the use case regardless of how well it scores elsewhere. The other three are trade-offs.

A worked example

A services business arrives wanting "AI for customer support", having seen a chatbot demo. Three candidates come out of the workflow mapping.

A customer-facing chatbot. Hours reclaimed: substantial. Error tolerance: low, because a wrong answer to a customer costs trust and nobody reviews replies before they send. Data sensitivity: moderate. Time to signal: fast. It fails the error gate as scoped, though a version that drafts replies for a human to send passes comfortably.

Triage and routing of inbound email. Hours reclaimed: moderate but constant. Error tolerance: high, because a misrouted email is noticed and forwarded in seconds. Run cost: trivial at their volume. Data sensitivity: low. Time to signal: two weeks. This one wins, and it was not on the original list.

Summarising call recordings into the CRM. Hours reclaimed: real, and hated by the team. Error tolerance: moderate. Run cost: meaningful, because audio is priced by duration and their calls are long. Time to signal: fast. Viable, and the run cost turns out to be the deciding variable rather than the capability.

The chatbot everyone arrived excited about was rescoped into something less impressive and more useful. That reordering is the normal outcome.

Then decide build, API, or buy

Only after scoring is it worth asking how to implement. Teams reach for the two extremes and both are usually wrong: training a model when a well-prompted API call and a fortnight of integration would have done it, or buying a platform whose one useful feature could have been shipped internally in a week.

Buying is right more often than engineers like. Building is right more often than vendors admit. The only way to know is to price all three paths over five years and read the result honestly, which is the same discipline as any other build-versus-buy decision.

Design the pilot so it can be killed

Every pilot ships with three things agreed in advance: a baseline recorded before launch, a success threshold, and a date on which it is judged.

The baseline is the part everyone skips and the only part that makes the rest possible. Without a measurement of how long the task took, how often it was wrong, or what it cost before the AI arrived, "it feels faster" is the strongest claim anyone will ever be able to make, and that claim can neither justify expansion nor support cancellation.

Naming the kill date in advance is what makes cancelling a normal outcome rather than an admission of failure. Roughly a third of use cases that reach scoring do not survive it, and a further portion do not survive the pilot. That is the process working, not the process failing.

Data governance, before rather than after

Which systems a model may read, which it may write to, what leaves your infrastructure, what is retained by whom, and who is accountable when an output is wrong. Agreed and documented before the first integration ships.

This is the section teams want to defer, and the one that ends AI programmes when it is deferred. In healthcare, insurance and financial services it is also the section a regulator will eventually ask about, and "we intended to write it down" is not an answer.

Frequently asked

How long does scoring AI use cases take?

Three to six weeks for a full engagement, most of which is mapping the workflows rather than evaluating tools. The scoring itself is quick once you know what the work actually is, which is precisely why so much of the effort sits before it.

What if we have already bought an AI tool nobody uses?

Score it as though you had not. Sometimes the tool is fine and was pointed at the wrong workflow; sometimes the workflow was never mapped. Occasionally the honest conclusion is to stop paying for it, which is worth reaching in week one rather than month six.

How do you measure whether an AI feature is working?

Against the baseline recorded before launch, on the metric agreed before launch. Usually one of: time per task, error rate, throughput, or cost per unit of work. If no baseline exists, the honest answer is that you cannot measure it, and the first step is establishing one.

Should we build our own AI models?

Almost never, for almost every business. The cases that justify it involve proprietary data at scale and a task no general model handles adequately. Everything else is better served by an API and good integration work, which is where the difficulty actually lives.

Is our data safe if we adopt AI tools?

That depends on design decisions, not on vendor marketing. Define what a model may read, what it may write, what leaves your infrastructure and what is retained, all before integrating. For some regulated workflows the correct conclusion is that a hosted model cannot be used at all.

Taking new engagements for Q3

Applying this to your own situation?

Thirty minutes, no pitch. Bring the decision you are actually weighing and I will tell you what I would do, including when the answer is to do nothing yet.

Chat on WhatsApp