Skip to content
Service

AI & Document Automation

AI applied to the boring, expensive parts of your operation — document extraction, classification and review workflows — with confidence scoring and a human in the loop.

What it covers

The work itself

Document extraction

Invoices, receipts, purchase orders, contracts, forms, statements — converted into structured, validated data with a confidence score on every field and a review queue for what falls below the threshold.

Classification and routing

Incoming documents, emails and tickets sorted, tagged and routed to the right queue or person, with the model's uncertainty made visible rather than hidden.

Retrieval over your own content

Search and question-answering across your policies, contracts, tickets and documentation — with citations back to the source, because an answer you cannot verify is not usable in an operational setting.

Agent workflows

Multi-step automations that call your systems, with explicit approval gates on any action that spends money, sends something externally, or is hard to reverse.

Evaluation and monitoring

A test set built from your real documents, accuracy measured against it, and regression checks that run when a prompt or model changes. Without this you are guessing, and the guess is usually optimistic.

Cost engineering

Per-document and per-request cost modelling, caching, model routing and hard spend ceilings. The economics have to work at your volume, not at a demo's volume.

Who it's for

And who it is not for

Teams where people currently retype, sort or look things up as a full-time activity, and where the volume is high enough that a few percentage points of automation are worth engineering properly.

Right fit if

  • Staff retype data from PDFs or emails into another system
  • Rule-based OCR works for your top vendors and fails on the long tail
  • Documents queue up and the backlog is managed by hiring
  • You need an audit trail on automated decisions
  • You have tried a generic AI tool and could not get it reliable enough to trust

Not the right fit if

  • Low-volume work where a person doing it manually is genuinely cheaper — we will say so
  • Projects that need an accuracy guarantee no honest vendor can give
  • Requests to 'add AI to it' with no specific, measurable task behind them

We would rather lose the project than take one we are the wrong firm for.

What you get

Deliverables, in writing

This list goes into the proposal. If something is not on it, it is not in scope — and we would rather argue about that now than in month three.

  • A working pipeline deployed to your infrastructure
  • Evaluation set built from your real documents, with measured baseline accuracy
  • Confidence scoring and a human review interface
  • Validation rules independent of the model
  • Audit trail on every automated decision
  • Cost-per-document instrumentation with a hard ceiling
  • Runbook covering failure modes and fallback behaviour

Timeline & engagement

What it costs you in time

Typical timelines

Proof of concept on your data

1–2 weeks

We run a sample of your real documents through a pipeline and give you the measured accuracy before you commit to anything.

Production pipeline

6–10 weeks

Ingestion, extraction, validation, review queue, integration and monitoring.

Full workflow platform

3–6 months

Multi-tenant, approval chains, accounting integration, analytics — the shape of the invoice extraction system in our solutions.

Engagement models

Paid proof of concept

A fixed-price two-week engagement producing measured accuracy on your own documents and an honest assessment of whether the economics work. Roughly a third of these conclude that they do not, and we say so.

Best for: Everyone, before committing to a build.

Fixed scope build

A defined pipeline at a defined price, once the proof of concept has established what is achievable.

Best for: Well-defined document types and volumes.

Retainer

Ongoing accuracy improvement, new document types, model updates and cost tuning.

Best for: Live pipelines, where accuracy work is continuous rather than one-off.

FAQ

AI Automation — questions we get asked

  • Will AI actually solve our problem?

    Sometimes not, and we would rather establish that in a two-week paid proof of concept than in a six-month build. The cases where it does not work are usually low volume where a person is cheaper, tasks needing an accuracy guarantee no model can give, or problems that turn out to be data-quality problems wearing an AI costume. We will tell you which one you have.

  • How do you stop the model from making things up?

    By not relying on the model to police itself. Output is constrained to a strict schema; every extracted value must be traceable to a location in the source; and arithmetic and business rules are validated by ordinary code that has no knowledge of what the model produced. Anything that fails those checks goes to a human, not to your database.

  • Do you use our data to train models?

    No. Your documents are not used for training. Where accuracy improves with use, it does so through retrieval — your own accepted examples supplied as context for future documents — and that context stays inside your tenant.

  • What does it cost to run?

    Model spend is usually small per document and it is always instrumented, so you can see it per tenant and per document type. We build in a hard spend ceiling with automatic cutoff, because an unbounded API bill is a real operational risk and pretending otherwise is negligent.

  • Can this run entirely on our own infrastructure?

    The pipeline, storage and interfaces can. The model is the question — self-hosted open-weight models trade some quality for full control, and whether that trade works depends on your documents. The proof of concept measures both options on your data so the decision is made on numbers.

  • Which model do you use?

    Whichever measures best on your task, and we re-evaluate as models change. Model choice is a configuration decision in the pipelines we build, not an architectural commitment — being locked to one provider is a risk we design out.

Tell us what your operation actually does.

A 30-minute call, no slide deck. We will tell you whether this is worth building, whether an off-the-shelf product would serve you better, and roughly what it would take.

30 minutes · We reply within one business day.

WhatsApp Book a call