AI features that earn their keep. Data foundations that don't fall over.

One feature, six weeks, real evals — picked because it earns its keep, not because there's a board slide called AI. We tell you which features shouldn't be built, and ship the ones that should — with evals, observability, and a fallback path.

Engagement process diagram

The kind of AI and data work we take on

01.

AI features inside products that already ship

One production AI feature — search, support, classification, generation, or whatever the product actually needs — built with evals, observability, and a fallback path from day one. Six-week typical engagement, fixed price. Includes a written 'kill it or keep it' recommendation at the end, with the evidence you'd need to defend either choice to your board.

02.

Data foundations that make AI possible

If your AI feature is going to live, the data layer has to be honest. We build the pipelines, schemas, and observability that turn the data you already have into something a model can actually learn from. Retrieval, structured extraction, evaluation harnesses — built so your team can keep running them after we leave.

03.

AI roadmaps that survive a board meeting

A written, prioritised list of the AI features worth building, the ones that should wait, and the ones that shouldn't be built at all — with the cost, risk, and engineering shape of each. Usually two-to-four weeks. Built for CTOs, COOs, and CIOs who need a defensible answer to 'where's our AI strategy?' that isn't just hype.

01.

AI features inside products that already ship

One production AI feature — search, support, classification, generation, or whatever the product actually needs — built with evals, observability, and a fallback path from day one. Six-week typical engagement, fixed price. Includes a written 'kill it or keep it' recommendation at the end, with the evidence you'd need to defend either choice to your board.

02.

Data foundations that make AI possible

If your AI feature is going to live, the data layer has to be honest. We build the pipelines, schemas, and observability that turn the data you already have into something a model can actually learn from. Retrieval, structured extraction, evaluation harnesses — built so your team can keep running them after we leave.

03.

AI roadmaps that survive a board meeting

A written, prioritised list of the AI features worth building, the ones that should wait, and the ones that shouldn't be built at all — with the cost, risk, and engineering shape of each. Usually two-to-four weeks. Built for CTOs, COOs, and CIOs who need a defensible answer to 'where's our AI strategy?' that isn't just hype.

How our AI work is different

We say no a lot

About one in three AI features we evaluate, we recommend killing — before the build budget is spent, not after.

We'd rather lose the build than ship a feature that gets quietly turned off in three months. That's what 'production-grade' actually requires.

We say no a lot

Evals before features

Before we build, we agree on how we'll know it's working.

Eval sets, accuracy thresholds, latency budgets, failure modes — written down, owned by you, run on every change. No 'looks good in the demo, broken in production' surprises.

Evals before features

Fallbacks, not faith

Every AI feature ships with a deterministic fallback path and clear observability for when the model is wrong.

The point of production AI is that the rest of the product keeps working when the AI doesn't. Most teams skip this; we make it the default.

Fallbacks, not faith

Built so your team can keep running it

We hand over runbooks, prompts, eval pipelines, and the cost model.

Your team can change a prompt, add an eval, or swap a model without calling us. Lock-in is bad engineering — and we're not in the business of selling it.

Built so your team can keep running it

We say no a lot

About one in three AI features we evaluate, we recommend killing — before the build budget is spent, not after.

We'd rather lose the build than ship a feature that gets quietly turned off in three months. That's what 'production-grade' actually requires.

We say no a lot

Evals before features

Before we build, we agree on how we'll know it's working.

Eval sets, accuracy thresholds, latency budgets, failure modes — written down, owned by you, run on every change. No 'looks good in the demo, broken in production' surprises.

Evals before features

Fallbacks, not faith

Every AI feature ships with a deterministic fallback path and clear observability for when the model is wrong.

The point of production AI is that the rest of the product keeps working when the AI doesn't. Most teams skip this; we make it the default.

Fallbacks, not faith

Built so your team can keep running it

We hand over runbooks, prompts, eval pipelines, and the cost model.

Your team can change a prompt, add an eval, or swap a model without calling us. Lock-in is bad engineering — and we're not in the business of selling it.

Built so your team can keep running it

How we work on AI

AI features we evaluate, we recommend not building
NOVA
1 in 3
Typical engagement for one production AI feature with evals
NOVA
6 wks
For a written, prioritised AI roadmap your board can defend
NOVA
2-4 wks
Model lock-in. Yours to swap, ours to make easy.
NOVA
0

Questions a CTO usually asks

We evaluate against three honest questions: does it solve a real problem the user has today, can it ship with a fallback that doesn't break the rest of the product, and is the cost-per-call sustainable at the scale you actually run at. If the answer to any of those is no — or unclear — we say so.

We don't replace data scientists; we ship the engineering layer they need to get models into production. Pipelines, retrieval, evals, observability, model versioning, on-call patterns. Most data science teams have a model that works in a notebook and no clear path to making it the thing customers depend on. That's the gap we fill.

Model-agnostic by design. Frontier models from the major providers, open-weight models on your own infrastructure, retrieval against your data — we pick what fits the eval and the cost envelope. We bias toward stacks your team can keep running without us, not toward whoever pays us a referral fee.

Six weeks for one production AI feature with evals and observability. Two-to-four weeks for an AI roadmap. Eight-to-twelve weeks for a data-layer rebuild. We don't do open-ended retainers for AI work — every engagement has a written end date and a defined hand-over.

We agree on the eval thresholds before we build. If we don't hit them, we tell you in writing — with the evidence — and we don't pretend otherwise. Often the right call is to kill the feature; sometimes it's to keep iterating; occasionally it's to change the underlying problem we're solving. We don't ship features the evals say aren't working.

Your data stays on your infrastructure unless we have a written reason to move it. We sign DPAs and NDAs by default, work to your regulatory environment, and don't train on your data for anyone else. For regulated industries (healthcare, financial services) we work to your audit standards from day one.

Bring us the AI feature you're not sure about.

Not the one that's already in the press release. The one where you're not sure if it should be built, or whether it'll survive contact with real users. We'll leave you with a one-page memo on whether to ship it, kill it, or keep evaluating.