AI StudioOpen for new builds

AI is not a product we sell you.
It is a layer we wire in.

Your ERP, your document store, your ticketing system, your data warehouse — they already hold the answers. We build the retrieval, orchestration and evaluation layer that lets a model use them safely, and we ship it into the product your customers already use.

3
modules live in production
6 wks
from kickoff to a pilot users can touch
100%
of retrieval scoped to user permissions
0
vendor lock-ins — models are configuration
The wiring board

Left: what you already own. Right: where it has to show up.

The interesting engineering is in the middle column. Anyone can call a model API — the work is making it answer from your data, under your permissions, with a way to prove it was right.

Sources
Systems of record
SAP · Dynamics · NetSuite · Salesforce
Document stores
SharePoint · S3 · Confluence · Drive
Operational data
Postgres · Snowflake · BigQuery
Ticketing & comms
Zendesk · ServiceNow · Jira · Outlook
The layer we build
01
Retrieval
Chunking, embeddings, hybrid search, re-ranking — scoped to the signed-in user's permissions.
02
Orchestration
Routing between models, tool calls, guardrails, fallbacks when a provider degrades.
03
Evaluation
A graded test set that runs on every change. Regressions block the release, like any other test.
04
Observability
Every prompt, retrieval and tool call traced, costed and replayable.

Runs in your cloud tenancy

Surfaces
Your product
In-app assistant, search, summarisation
Your agents
Draft replies, triage, next-best-action
Your back office
Extraction, reconciliation, routing
Your customers
Web chat, WhatsApp, voice, email
The board

Where each module actually stands

Ordered by maturity, not by how impressive it sounds. “Live” means it is running against real data for a paying client. Every row also states what it is still bad at — because that second column is how you judge the first.

Live in production

Grounded document assistant

Answers questions over your contracts, policies and runbooks, citing the paragraph it used. Refuses when retrieval comes back empty.

What it is still bad at

Strong on finding the deciding clause. Still weak on documents that contradict each other — it flags the conflict rather than resolving it.

Live in production

Service desk triage

Reads an inbound ticket, classifies it, attaches the customer's entitlement and history, and drafts the first reply for an agent to approve.

What it is still bad at

Reliable on the top twenty intents. We deliberately route the long tail to a human instead of guessing.

Live in production

Structured extraction

Turns invoices, purchase orders and scanned forms into validated records, with confidence scores and a review queue for anything below threshold.

What it is still bad at

Accuracy is good on clean scans and honest about the rest — low-confidence items go to a person, not into your ledger.

In client pilot

Product imagery generation

Generates on-brand product and lifestyle imagery from a controlled style reference, with a human approval step before anything publishes.

What it is still bad at

Not usable yet for products with fine text or precise logos on the item itself.

In client pilot

Agentic back-office workflows

Multi-step internal processes where an agent takes permissioned actions across systems, with every step logged and reversible.

What it is still bad at

The interesting part is not the agent, it is the authorisation model. We scope each agent to a short, explicit list of write actions.

In design

Voice intake

Handles routine inbound calls — order status, appointment changes — and hands off with the transcript attached.

What it is still bad at

Still in design. We are not shipping a voice agent until barge-in and interruption handling feel genuinely natural.

How we build it

Five rules that do not bend

These apply to every module on the board, including the ones still in design. They are the reason we can put an AI system in front of an auditor.

Retrieval before generation

If a claim can be looked up, it gets looked up. The model composes the sentence; it does not invent the fact.

Permissions are the user's, not ours

Retrieval runs as the signed-in user. There is no service account that can see everything.

A short list of write actions

Anything that changes your data is an explicit, authorised, logged action — never an inference.

Evaluation is a build artefact

A graded test set ships with the system and runs on every change. Regressions block the release.

Model-portable by default

Providers are configuration. Swapping a model is a deploy, not a rewrite — so you are never hostage to one vendor's pricing.

Three ways in

Start small enough that being wrong is cheap

01Fixed fee

AI readiness review

Two weeks. We inventory your data, permissions and processes, then tell you which three use-cases are worth funding — and which ones are not.

02Fixed fee

Six-week pilot

One intent, your data, behind a feature flag, with an evaluation set you keep. Cancellable at the halfway point.

03Monthly

Embedded AI engineers

GenAI and ML engineers on our payroll, inside your sprints, reporting to your lead. Scale up or down monthly.

Bring us the process, not the prompt

Tell us which workflow is slow, expensive or error-prone. We will tell you whether AI is the right tool for it — including when the answer is no.