Our work
Seven problems, and what actually changed.
Client names are withheld under NDA. The numbers are not. Every figure below comes from a real audit, gate run or deployment log on a live production system, and each page explains how it was measured.
A support inbox that answers first
A customer-support platform needed an agent that answers customers on its own without inventing anything, and a way for reps to work the inbox from inside Claude Desktop, Cursor or any AI assistant. We shipped a grounded support agent and a twelve-tool MCP server with permission tiers, suggest-only autonomy and bug routing to Linear, Jira and GitHub. Both live for paying customers.
MCP tools, 3 tiers
Autonomy modes
Bug trackers wired in
The meeting that books itself
Booking a call across Lahore and US time zones cost senior engineers four to six emails each. We built Geralt, an agent that lives in email: copy it on one message, it proposes slots, reads free-text replies, sends calendar invites and escalates to a human only when it must. No app, no login, no dashboard.
Email to book a call
Apps or logins
Runs on its own
Models that ship like software
A model in a notebook is an experiment. We built the operations around a marketplace's search models: a repeatable labelling pipeline, versioned fine-tuning runs, GPU serving with autoscaling, an LLM-as-judge evaluation loop and a regression gate that blocks bad releases before customers see them.
Median query time
Judge hit rate vs 72.4%
Fixtures in the gate
Search that stopped guessing
A UK life-sciences talent marketplace, operating across the UK, US and Australia, could not reliably match specialist briefs to the right people. The vocabulary of the industry defeated keyword search. We rebuilt the engine around embeddings, reranking and a domain taxonomy, then built the evaluation layer that proved it worked before anything shipped.
Exact-match accuracy
Regression fixtures
Byte-identical replay
A mobile release that would not build
A React Native app had never once produced a TestFlight build, and the store deadline was closing. Nine consecutive CI failures, from dependency locks and stale pods to signing keys and finally a C++ compiler error deep inside a formatting library, were each root-caused and fixed. We also rejected the expensive fix that would not have worked.
Build failures cleared
TestFlight build ever
Framework upgrades needed
A live site, compromised
A production marketing site went down to an SEO-spam compromise. We identified the entry vector, an unauthenticated API route, removed the web shells and scheduled-task persistence, restored files and database from a verified clean snapshot, rotated every credential, closed the hole and left a written incident report behind.
Diagnosed to restored
Backup tiers installed
Credentials rotated
A field platform, made trustworthy
A white-label survey platform used by field teams was silently failing on submission and hiding reference images from the people who needed them. We traced both to authorisation and data-integrity defects, fixed them without widening access, and rebuilt the entire interface against a new design system, verified screen by screen.
Silent prod bugs closed
Access-control flaw fixed
Regressions on rollout
Next step
Have a problem that looks like this?
Tell us what is actually going wrong. You will get a straight answer on whether we are the right people for it, and a fixed scope if we are.
