A model that does a hard thing badly is still a hard thing. Use cases should be selected for the value of the decision, not the novelty of the technology.
Applied AI
AI that earns its place inside the organisation — by being governable, observable, and tied to outcomes worth measuring.
On this page
§ 01 ·
It is rarely the right answer on its own. Most of the work is the surrounding plumbing: the data, the integration, the governance, the change management, and the model of how a decision moves from an AI recommendation to a trusted action.
The four phases the firm uses to do this work are on the How we work page.
§ 02 ·
Patterns we design our practice to avoid
Six failure modes we have seen repeatedly, across AI engagements of every size. Each one motivates a specific choice in how we work. The page opens with the failure modes because that is where the work is hardest to do well.
Most AI failures are data failures. If the data pipeline cannot produce a reliable, fresh, labelled signal, the model will not work in production regardless of its quality on the test set.
Most production AI systems include a person somewhere in the loop, even when the architecture does not show one. The handoff is the difference between a system people use and a system people route around.
A model that improves accuracy on a metric that does not matter is a model that has been optimised in the wrong direction. Outcome metrics are the ones that justify the investment.
When governance reviews happen after the system is built, the only available lever is approval or rejection. Effective governance sets the paved roads up front so teams move quickly within them.
An LLM is an untrusted component that produces plausible output. Treating it as authoritative, without an explicit boundary between its output and the action it triggers, is the most common path to a production incident.
§ 03 ·
How we approach AI work
The common approach treats the model as the deliverable. Our approach treats the system around the model: the data, the integration, the governance, the operating model, as the deliverable, and the model as one component of it.
The common approach
A model is trained on historical data, wrapped in a thin API, handed to delivery, and "released". Success is measured on accuracy against a holdout set. Failures are investigated by the data team alone.
- Model accuracy is the primary success metric.
- Governance is a checklist at the end.
- Production behaviour is observed only when something breaks.
- The model is the artefact; the system is an implementation detail.
What we do
A model is one component of a system designed for a specific decision or operation. Readiness across data, workflow, governance, and operations is assessed before the model is built. The evaluation harness and the operating model are part of the deliverable.
- Outcome and trust metrics, not just model metrics.
- Governance paved roads up front; the review board sees exceptions.
- Continuous monitoring of model behaviour and business outcome.
- The system around the model is the artefact; the model is a component of it.
What you walk away with
A flat list of the artefacts an AI engagement produces. The full shape of the engagement is on the [How we work](/services/methodology/) page.
A use-case shortlist
A short list of use cases the organisation is willing to pursue, scored on value, feasibility, and risk. The output is not a roadmap; it is a small set of bets.
A readiness assessment and integration design
A readiness assessment across the four axes (data, workflow, governance, operations) for each selected use case, and a design for where the model sits in the decision flow, what its inputs and outputs are, and how its output is reviewed and overridden.
A working system with its evaluation harness
The model in production, the surrounding integration built, and the evaluation harness in place. The harness is part of the deliverable, not a side artefact.
Operating-model documents and monitoring
An operating model: who monitors the model, who escalates, who retrains, who reviews. Monitoring in production for drift, output sampling, and outcome metrics. Governance review for significant changes.
The LLM safety pattern (if the use case involves an LLM)
A defined contract for how the LLM is invoked, what its inputs are validated against, what its outputs are checked against before they reach the user or the system, and what actions it never has authority to take without human review.
§ 05 ·
What each phase produces in an AI context is on the How we work page.
§ 06 ·
Common questions from clients
The questions we hear most often, in the order they tend to come up.
We start from the decision or operation, not from the technology. An AI investment is justified when the underlying task is one a machine can credibly do, where the value of doing it better is large, and where the cost of being wrong is bounded and recoverable.
AI is rarely the right answer when the value of being right exceeds the cost of being wrong by an order of magnitude, when the data does not exist, or when the simpler alternative: process change, a rules engine, a better interface, would close most of the gap.
Most AI engagements run between three and nine months. The framing phase is six to twelve weeks; the readiness assessment and integration design is the bulk of the early work; the build phase is sized to the use case; the operation phase is ongoing. The total horizon is set in the framing session and revised only by mutual agreement.
Engagements are priced against the use-case shortlist produced in the framing phase. The cost depends on the use case, the readiness of the data, the depth of the integration, and the operating model the work needs. A typical range for a multi-phase engagement sits in the low-to-mid six figures; the final figure is set in the scoping call after the framing phase.
The first 30 days are the framing phase: an enumeration of candidate use cases with the leadership team, structured conversations with the people who will own the engagement, and the production of the use-case shortlist. By the end of the first 30 days, the client has a small set of bets it is willing to make.
An LLM in production is treated as an untrusted component with a defined contract. Inputs are validated and bounded. Outputs are checked against the contract before they reach the user or the system. The LLM never has authority it has not been explicitly given, and every action it influences is logged.
This is unglamorous work and most of it. The model is the visible part; the safety, evaluation, and observability layers around it are what make the system production-grade.
§ 07 ·
Evidence & references
Public frameworks and writing that inform the firm's practice.
The Govern-Map-Measure-Manage structure is the vocabulary for AI governance. We use it as a checklist against specific use cases, not as a process.
The end-to-end treatment of the operational side of ML — data, feature stores, monitoring, continual learning.
The pattern language for operationalising models: deployment, monitoring, drift detection, retraining cadence.
The data-pipeline framing: data flow, data lineage, the contracts between systems. Most of an AI engagement's hard work lives at this layer.
The technology-and-society framing. The choice of which AI work to do is a choice the firm and the client are making about what the technology is for, not a technical question with one right answer.
Want to talk AI?
If you are evaluating a use case, building a model into a production system, or strengthening an existing AI practice, the firm is useful at that boundary. A short conversation is the right next step.
Start a conversation