Skip to content
Delta AI Engineering

Approach

Measurable, inspectable, and under human control.

The Delta Method keeps every AI engagement honest: agree the change, build the simplest system that can produce it, prove it, and keep proving it.

Stages

Stage 1 of 5

What change should this system produce, and how will we measure it?

We map the workflow, the people in it and the data behind it, then agree a baseline and a success metric before any model is chosen.

Leaves behind

Outcome brief with baseline metric

Release gates

What has to be true before anything ships.

These gates are set per engagement. The numbers come from your baseline, never from a generic benchmark.

Release gates and how each is measured
GateHow it is measured
Task qualityGraded evaluation set built from your real cases, compared with the agreed baseline.
GroundednessShare of answers fully supported by cited sources; unsupported claims block release.
Latencyp50 and p95 end-to-end response time against the agreed budget.
Cost per taskModel, retrieval and infrastructure cost per completed task, tracked per release.
SafetyRed-team prompts, permission checks on tools and human approval for consequential actions.
OperabilityTraces, alerts and a runbook in place, with a named owner, before go-live.

Budgets

Quality, latency and cost are design inputs.

Drag the marker, or choose a priority below.

Priority
Model
Mid-tier model with routing to a larger model for hard cases
Retrieval
Hybrid retrieval (≈25 candidates) with reranking
Caching
Prompt caching on shared context
Oversight
Sampled human review with drift alerts

Illustrative design heuristics. Real choices are set by your evaluation results and budgets.

Principles

How we make decisions.

  • Measured, not assumed

    Every engagement starts with a baseline and ends with a comparison against it.

  • People stay in control

    Consequential actions pause for human approval. Automation earns autonomy gradually.

  • Explicit over clever

    State machines, typed contracts and written decisions make systems easier to trust and to change.

  • Budgets are features

    Quality, latency and cost targets are agreed up front and tracked like any other requirement.

Toolbox

Tools we work with.

Chosen per project against your constraints. We are not resellers or certified partners of any vendor listed.

  • LangGraph
  • Model Context Protocol
  • pgvector
  • Hybrid search + rerankers
  • LangSmith
  • Docker
  • Kubernetes
  • GitHub Actions
  • LoRA / QLoRA fine-tuning
  • Next.js
  • Connect RPC
  • XState

Tell us the change you need. We will tell you how we would measure it.

A short first conversation, a written summary afterwards, no obligation.