Every number below is measured against the manual baseline we recorded in week one — not against a vendor benchmark.
Banking
Meridian — dispute triage
A card-dispute queue running nine days behind. The intake agent reads the claim, pulls the transaction trail and drafts the regulator-ready summary for an analyst to approve.
9d → 4h
Queue age
71%
Straight through
Healthcare
Kestrel — prior authorisation
Twelve staff assembling authorisation packets by hand. The agent gathers the clinical evidence, checks it against payer rules, and stops dead on anything ambiguous.
3.1×
Packets per day
0
Unreviewed sends
Logistics
Northwind — exception desk
Every delayed load used to generate four emails and a phone call. The agent chases the carrier, updates the customer, and wakes a coordinator only when the ETA slips twice.
4,200h
Saved per year
−38%
Inbound calls
Logistics
Northwind — exception desk
Every delayed load used to generate four emails and a phone call. The agent chases the carrier, updates the customer, and wakes a coordinator only when the ETA slips twice.
4,200h
Saved per year
−38%
Inbound calls
Healthcare
Kestrel — prior authorisation
Twelve staff assembling authorisation packets by hand. The agent gathers the clinical evidence, checks it against payer rules, and stops dead on anything ambiguous.
3.1×
Packets per day
0
Unreviewed sends
Banking
Meridian — dispute triage
A card-dispute queue running nine days behind. The intake agent reads the claim, pulls the transaction trail and drafts the regulator-ready summary for an analyst to approve.
9d → 4h
Queue age
71%
Straight through
What we build
Five agents. One operating standard.
Each one ships with its own evaluation set, spend cap and escalation rule. You get the agent, the harness around it, and the screen your team runs it from.
Tier-1 resolution across email, chat and voice, grounded in your real help centre. Refuses anything outside scope and writes a clean handoff packet when a person is needed.
Invoice reconciliation, onboarding checks, refunds and renewals. The repeat work that quietly eats a third of your team’s week, done overnight with an audit trail.
The part most agencies skip. Labelled test sets, regression runs on every prompt change, and an alert the day quality starts to slip.
How it holds up
Built to be handed over, not rented.
You own the code, the prompts, the test set and the infrastructure definitions from day one. We are the team that runs it until your team wants to.
Runs in your cloud
Deployed into your account, or fully on premise for regulated work.
Spend capped per run
A hard budget per task. Past it, the work goes to a human queue instead.
Model-agnostic routing
Whichever model wins on your test set, re-checked every quarter.
400-day replay
Reconstruct any decision step by step, long after the invoice cleared.
Handover in 4 weeks
Client voice
"They spent the first week watching us work instead of talking about models. The agent they built behaves like someone who has actually sat on the dispute desk."