Case studies
Five AI agent engagements, end to end: the baseline that was measured before anything was built, what was built, and what the numbers said afterward. Client names are withheld by agreement.
Case study · Voice agent
A residential HVAC and plumbing contractor — $12M revenue, 34 employees, nine trucks.
The office answered the phone from 8am to 5pm on weekdays. Everything else went to voicemail. Instead of guessing what that cost, we pulled six months of call records: roughly 90 calls a month arrived outside office hours, about 22 of them genuine new-job inquiries. Next-morning callbacks recovered around five. The rest had already reached the next contractor on Google.
At the company's average job value and margin, that gap was worth roughly $18,400 a month in gross profit — going to voicemail.
A voice agent that answers every out-of-hours call, qualifies the job, checks live technician availability, offers two appointment slots, confirms by SMS, and writes the job into the field service system before anyone wakes up. It was scripted from transcripts of the office's own best coordinator, so it sounds like the company — not like a chatbot.
Before launch, the agent replayed thirty real recorded calls and its booking decisions were compared against what the office actually did — then the client's own staff spent a round of live calls actively trying to break it. It ran a week in shadow mode with every transcript emailed to the owner, then weekends only, then all out-of-hours. Nothing was trusted before it was measured.
After-hours jobs booked
5 / month
15–17 / month
Recovered gross profit
$0
~$13,000 / month
Annualized value
—
~$156,000
Running cost
—
$60–80 / month
The change, drawn to scale
After-hours jobs booked per month
Recovered gross profit per month
What keeps it honest: availability is always a live read — never a cached schedule — so the agent can't book work the trucks can't service. Some callers do hang up on an AI voice. The honest comparison is against voicemail, not against a human.
Case study · Operations agents
A freight brokerage — eighteen brokers, roughly 2,400 loads a month.
Two workflows consumed the working day. Brokers phoned and emailed carriers three to five times per load for status updates — about 600 hours a month spent chasing status. Meanwhile, inbound spot-quote requests took an average of 45 minutes to answer, in a market where the first credible quote usually wins the load.
A track-and-trace agent that contacts carriers at scheduled checkpoints — SMS first, a voice call if there's no reply — parses whatever comes back, updates the transport management system, and involves a broker only when something is off-plan. And a quoting agent that reads inbound requests, extracts the lane and requirements, prices against live market data and the firm's own margin rules, and returns a quote within minutes. The language model extracts and drafts; deterministic code does the pricing.
Five hundred historical quote emails were replayed and extraction was measured field by field before a single live quote went out. A month of past loads ran through the tracking logic to confirm it caught every exception the brokers had caught. Tracking launched on one broker's book first; quoting ran in approval-required mode for a month before any threshold was loosened.
Check-call hours
~600 / month
~90 / month
Quote response time
45 minutes
Under 5 minutes
Quote volume
Baseline
+40%, same headcount
Win rate
Baseline
+8–12%
The change, drawn to scale
Hours spent chasing load status per month
Time to answer a spot-quote request
Annualized value measured at over $700,000. What keeps it honest: approval bands stay tight permanently on non-standard loads, and the exception alerts belong to the brokers — not to a manager's dashboard.
Case study · Back-office agents
An accounting and bookkeeping practice — fourteen staff, 240 monthly bookkeeping clients.
Every month, staff emailed and called clients for bank statements, receipts and missing invoices — roughly 190 hours a month, and the work everyone deferred longest. Behind it, around 40,000 transactions a month needed categorizing, with about 15% requiring a manual touch or a query back to the client. Month-end close averaged eleven business days.
A document chaser that knows precisely what is outstanding for each client, chases on an escalating schedule across email and SMS, accepts uploads by reply — photographed receipts included — validates them and files them to the right client and period. And a reconciliation agent that categorizes transactions against each client's own history, never a generic chart of accounts, auto-posts only above a measured confidence threshold, and drafts plain-English queries for anything genuinely ambiguous.
Categorization was replayed against a full quarter of already-reconciled transactions, and the auto-post threshold was set where measured accuracy exceeded 98% — with a senior accountant reviewing 200 sample categorizations and every drafted query. The chaser launched on thirty clients with staff approving every message for the first week; reconciliation ran a full close cycle in review-everything mode before any threshold moved.
Document chase hours
~190 / month
~35 / month
Docs in by day 10
45%
80%
Manual-touch transactions
15%
4%
Month-end close
11 days
6 days
The change, drawn to scale
Document chase hours per month
Clients with documents in by day 10
Business days to close the month
Annualized value measured at over $190,000 — plus capacity for roughly forty more clients without hiring. What keeps it honest: the auto-post threshold stays conservative permanently, because a wrong posting surfaced in an audit costs far more than the minutes it saved, and per-client data isolation is verified, not assumed.
Case study · Intake agents
A personal injury law firm — six attorneys, roughly $4M in annual fees.
Paid advertising delivered roughly 400 inquiries a month. About 45 were viable. Paralegals spent around 90 minutes screening each one — approximately 530 hours a month, most of it spent ruling out the 355 cases the firm would never take. Worse, viable leads waited an average of four hours for a response, in a practice area where the firm that responds first usually signs the client.
An intake screening agent that engages every inquiry within sixty seconds by SMS, gathers the facts conversationally, scores viability against the firm's own criteria — with a written reason for every factor, not a bare number — and books qualified prospects directly into an attorney's calendar. Alongside it, a case summary agent that turns the intake record and uploaded documents into a one-page brief before the consultation, with every factual claim cited back to its source document. No citation, no claim.
Two hundred historical inquiries — one hundred accepted, one hundred declined, each with the recorded reason — became the gold-standard test set. The only score that truly mattered was false negatives: viable cases wrongly declined. The target was zero, and the threshold is deliberately tuned to over-refer to humans. The managing attorney reviewed and signed off every line of the script, two attorneys verified every citation in twenty sample briefs, and the agent ran two weeks in shadow mode with human review of every decision before going live on a single advertising source, then all inbound.
Response time
4 hours
Under 60 seconds
Screening hours
~530 / month
~120 / month
Signed cases
Baseline
+15–25%
Consultation prep
~40 minutes
~10 minutes
The change, drawn to scale
Paralegal screening hours per month
Response time to a new inquiry
Attorney prep per consultation
Annualized value measured at over $250,000. What keeps it honest: a distressed caller reaching a bot is the failure mode that matters most, so sentiment detection with immediate human handoff is mandatory — and transcripts are reviewed on a schedule, because without tight output constraints the language drifts toward sounding like advice.
Case study · Reporting & revenue agents
A full-service marketing agency — 60 staff, roughly $14M revenue, 45 retained clients.
Three structural drains. Account managers spent about fourteen hours per client per month assembling reports — pulling numbers from analytics, paid social, search and the CRM, then writing commentary — around 630 hours a month, performed by the most expensive client-facing people in the building. The fourteen-hour figure came from observation, not estimate. Meanwhile, client sales teams took hours to days to respond to the leads the agency generated — and renewals were judged on cost per acquisition, a step the agency didn't control. And approved creative briefs queued three to five days for a first draft.
A reporting agent that pulls platform data nightly, compares it against prior periods, cross-references the campaign changes logged in the project system, and drafts commentary explaining why numbers moved — producing a review-ready report for the account manager. A speed-to-lead agent deployed into client funnels that responds to form submissions within sixty seconds, qualifies, and books warm prospects for the client's sales team. And a brief-to-draft agent grounded in each client's approved past copy, so brand voice is inherited rather than invented.
The last three months of reports were regenerated for ten clients and put side by side with what was actually sent; account managers scored the commentary for accuracy, and anywhere the agent asserted a cause without evidence, the instructions were tightened. Three senior writers blind-reviewed drafts against a brief for voice consistency. Reporting launched on ten clients with full review, speed-to-lead with two willing clients, and drafting stayed internal until the writers trusted the output.
Reporting hours
~630 / month
~90 / month
Client lead response
Hours to days
Under 60 seconds
Brief to first draft
3–5 days
Same day
Reporting cost per client
~$1,050 / month
~$150 / month
The change, drawn to scale
Client reporting hours per month
Reporting cost per client per month
Approved brief to first draft
Annualized value measured at over $500,000 in recovered senior time, plus the retention benefit of faster client results. What keeps it honest: the single most important guardrail is admitting uncertainty — a report that confidently invents a cause is worse than one that says a movement is unexplained. And the drafting corpus contains only genuinely approved work, because writers reject drafts that miss the voice.
Every engagement follows the same discipline we apply everywhere: measure the baseline first, build against it, prove the change with a re-measurement, then control it so the gain holds. Results above were measured on these specific engagements — they are not a guarantee of outcomes.
Every engagement starts the same way these did: with a measured baseline. Book a call and we'll find the workflow where the numbers say automation pays back first.