ServicesApplied AI

Applied AI that holds up in production.

Plenty of AI works in a demo. We make it a dependable part of a real workflow: tested on your own cases, with people approving the steps that carry risk, and monitored once it is doing real work.

  • Tested on your casesBefore launch and after every change
  • People approve risky stepsWith a record of what the system did and why
  • Runs in your cloud accountYour documents and records stay with you

Where it fits

Work where AI prepares and a person decides.

AI earns its place where people spend time reading, sorting and drafting from untidy material, and where the output can be checked. These are the problems we are asked about most.

01

Documents arrive faster than people can sort them

Invoices, claims, applications or orders arrive by email and upload in many formats, and someone reads each one to decide where it goes. AI classifies them, pulls out the key details and routes them, and anything it is unsure about goes to a person.

Typical output: an intake queue with classification, extracted fields and a review step
02

Replies and reports take hours to draft

Your team writes similar replies, summaries or reports from the same sources every day. AI prepares a draft from the case history and your templates, and a person edits it and sends it.

Typical output: drafts in the tool your team already uses, with their sources shown
03

Answers are buried in internal documents

Policies, contracts, tickets and wikis hold the answer, but finding it depends on a colleague who knows where to look. Search and question answering over your own documents, citing the passage it used and saying when it does not know.

Typical output: search with cited answers, limited to what each user is allowed to see
04

Useful data is locked in messy inputs

PDFs, scans, emails and spreadsheets in different layouts hold data that someone keys into a system by hand. AI extracts it into structured records, and rules check every record before it is saved.

Typical output: validated records in your database, with exceptions flagged for review
05

A step needs action, not just a suggestion

Some work means changing things in your systems: updating a record, raising a ticket, booking a slot. An agent can do this with a defined set of tools and limited permissions, and a person approves any step that carries risk.

Typical output: an agent with limited permissions, approval steps and a full action log
06

Your engineers want AI help without losing control

Coding agents are fast, but they do not know your team’s rules. We set them up with your conventions, checks and review steps, and measure whether they actually help.

Typical output: agent instructions and skills, automated checks and a measured comparison

From prototype to production

Five steps, in this order.

The order matters. We decide how the system will be judged before we build it, and keep judging it after it goes live.

  1. 01

    Frame the decision

    Agree which task or decision the AI supports, who is accountable for it and what a wrong answer costs: a quick correction, a delayed order or a problem with a regulator. That cost sets how much review the system needs.

    Outputscope, owner and the cost of a wrong answer, in writing

  2. 02

    Build the evaluation set first

    Before building anything, we collect real cases from your work, including the awkward ones, agree the correct answers with your team and write a checker that scores them. It runs before launch and after every change.

    Outputevaluation set and baseline results

  3. 03

    Design approval and fallbacks

    Decide which steps a person approves, what the reviewer sees and how easy it is to correct. Decide what happens when the model is unsure, slow or unavailable: send it to a person, fall back to a rule, or stop.

    Outputapproval points and fallback paths

  4. 04

    Build and integrate

    Retrieval over your documents, tools that act in your systems with limited permissions, and integration with the software your team already uses. It runs in your cloud account, and data leaves it only for the model calls you approve.

    Outputa working system in your environment

  5. 05

    Measure, monitor and control change

    In production we track quality through sampled checks and reviewer corrections, along with cost per task and response times, with alerts. Model versions are pinned, and a new model or prompt goes live only after it passes the evaluation set.

    Outputdashboards, alerts and a change log

What you get

A system your team can run.

The same handover standard as all our work: working software, the evidence that it works and the documents to keep it working.

  • A working system

    Deployed in your cloud account and connected to your existing software, handling a defined part of real work.

  • Evaluation set and results

    Your real cases with agreed answers, the checker that scores them and the results, so any later change can be measured the same way.

  • Decision records

    Why this model, this design and these approval points, written down so the choices can be revisited.

  • Runbook

    What to do when quality drops, costs rise, a provider fails or a model needs to change.

  • Monitoring dashboards

    Quality, cost per task, response times and failures, with alerts on the measures that matter.

  • Handover

    A walkthrough with the people who will run the system, so it does not depend on us.

Where AI isn’t the answer

Sometimes the right answer is no AI at all.

Part of the first conversation is checking whether AI is needed. If it is not, we will tell you, and suggest what would work instead.

  • The rules are already clearIf the logic can be written down, plain code or a workflow tool does the job more cheaply and predictably, and can be tested completely. We will say so.
  • Errors are unacceptable and cannot be reviewedIf a wrong answer would do serious harm and nobody can check the output in time, AI should not take that step.
  • The data does not exist yetWithout real cases to test against or documents to draw on, there is nothing to measure. The first job is to start capturing that data.

More on drawing the line: Choosing what to automate

How we test

Measured, not assumed.

In our Agent Skills note, we gave Claude Code the same 12 tasks in one repository with no skill, a short skill, a long skill and a skill it wrote for itself, and scored every run with an automated checker. On release tasks, Opus followed every house rule in 6 of 24 runs without a skill and in 24 of 24 with a short one. The skill the agent wrote for itself passed 0 of 30 release runs on Sonnet.

Bar chart from our Agent Skills note: share of Claude Code runs that followed every house rule, with and without an Agent Skill, on release and migration tasks, for Opus 5.5, Sonnet 5.5 and Haiku 5.5.

We hold client work to the same standard:

  • Question, conditions and analysis written down before the runs.
  • A run counts as a pass only if every check passes.
  • Results reported with their uncertainty and limitations.

Read the Agent Skills note

Questions

Questions buyers ask.

If your question is not here, include it in your enquiry and we will answer it directly.

Which AI models do you use?

We are not tied to one provider. We choose the model per task, using the evaluation set: Claude, GPT or Gemini models through their business APIs, or an open model hosted in your own cloud account where data rules require it. Often the smallest model that passes the evaluation is the right one.

What happens to our data?

The system runs in your cloud account, and your documents and records stay there. Model calls go only to providers you approve, under business terms that exclude training on your data, and we check those terms with you before anything is sent. We do not train models on your data.

How do you keep running costs under control?

We measure the cost per task on the evaluation set before launch, so you know what a typical month should cost. In production, cost is tracked alongside quality, with budgets and alerts, and we use cheaper models or fewer calls wherever the evaluation shows quality holds.

What happens when a provider changes or retires a model?

Model versions are pinned, so nothing changes silently. When a new version is due, we run it against the evaluation set and switch only if it passes. The result is recorded like any other change, and the previous version stays available until the switch is confirmed.

Can you take over a prototype we already have?

Yes. Many projects start with a prototype that works in a demo. We run it against real cases first, so we know what to keep, what to change and what is missing before it handles real work.

How long does a first version take?

It depends on the workflow, the data and the systems involved, so we agree it after the first call and set it out in the written proposal. We aim for a first version that handles a narrow slice of real work well, rather than a broad demo.

Contact

Tell us about the workflow.

Describe the work, where it slows down and what a wrong answer would cost. A few lines is enough, and a prototype is not needed.

Discuss a project
  • A reply within one working day. We read and answer every enquiry.
  • A 30-minute call. We talk through the problem and what you have tried so far.
  • A short written proposal. If it is a fit, setting out a first step, its scope and its cost.