Customer

Client voice
"They spent the first week watching us work instead of talking about models. The agent they built behaves like someone who has actually sat on the dispute desk."
Dana Whitlock
Head of Operations, Meridian Bank
How we work

Six weeks from
workshop to live.

The order matters. We will not build before we have watched the work done by hand, and we will not launch before the test set passes.

1

Watch the work

Two days sitting with the team that does the task today. We record the real decision points, the exceptions nobody wrote down, and where the cost actually sits.

Week 1 — process map & baseline
2

Build one narrow slice

A single workflow, running against real data in a sandbox. If it cannot beat the manual baseline on a labelled set, we tell you and we stop.

Week 2–4 — working agent
3

Harden, then operate

Guardrails, permissions, spend caps, rollback and escalation rules. Then we run it with you, on call, until your ops lead can ship a change alone.

Week 5–6 — live, then shared on-call
Selected work

Twelve months
of receipts.

Every number below is measured against the manual baseline we recorded in week one — not against a vendor benchmark.
Banking

Meridian — dispute triage

A card-dispute queue running nine days behind. The intake agent reads the claim, pulls the transaction trail and drafts the regulator-ready summary for an analyst to approve.

9d → 4h
Queue age
71%
Straight through
Healthcare

Kestrel — prior authorisation

Twelve staff assembling authorisation packets by hand. The agent gathers the clinical evidence, checks it against payer rules, and stops dead on anything ambiguous.

3.1×
Packets per day
0
Unreviewed sends
Logistics

Northwind — exception desk

Every delayed load used to generate four emails and a phone call. The agent chases the carrier, updates the customer, and wakes a coordinator only when the ETA slips twice.

4,200h
Saved per year
−38%
Inbound calls
Questions

Before you
get in touch

If yours is not here, ask it in the form below. We answer in writing before any call.

Do we need our own AI team first?
No. Most clients start with an operations lead and one engineer who can grant access. We supply the agent engineering; you supply the domain judgement about what a correct answer looks like. By month three your ops lead is usually shipping changes without us.
Into your infrastructure by default. On Scale we deploy into your cloud account; on Embedded we run fully on premise. Nothing trains a model, retention windows are set by you, and every third-party call is listed in the data map we hand over in week one.
It should stop rather than guess — that is what the confidence thresholds and refusal rules are for. When something still slips through, you replay the run step by step, see which tool call caused it, add the case to the test set, and the regression suite blocks it from happening twice.
Whichever wins on your evaluation set, re-tested quarterly. We are deliberately not tied to one provider — routing is a configuration file, not a rewrite. If compliance restricts you to a specific list, we fix routing to that list and show you the quality cost.
An agent running against real data in a sandbox by the end of week three. Production traffic usually lands in week six, gated behind the security review and a passing evaluation run. If we are going to miss that, you hear it in week two, not week five.
Yes, and it is written into the contract. You own the code, prompts, evaluation set and infrastructure definitions from day one. Handover is a four-week programme with shadowing, runbook review and a final on-call swap. No exit fee.
Two pilot slots open this quarter

Stop evaluating AI. Start operating it.

Ninety minutes with your operations lead and one of our engineers. You leave with a scoped workflow, an honest feasibility call and a number.