How It Works

From a Repetitive Task to a Monitored Agent

Four stages, in the same order every time: discovery, design & build, integrate & test, launch & monitor. This page goes into what actually happens at each one, and what we need from you to keep it moving.

Stage One

Discovery: Finding the Actual Task

Before anything gets designed, we need to understand the task as it's actually done today — not the idealised version, the real one with its exceptions.

What We Ask About

The current steps in the task, who does them today, which systems get touched, and — the part people often skip — the judgement calls a person makes along the way that a generic description would miss.

Where a Person Has to Stay in Charge

We ask directly which decisions you're comfortable handing to an agent and which ones should never leave a human's hands, rather than assuming and finding out later.

Discovery ends with a written scope: the systems involved, the decision boundaries, what counts as done, and an indicative price band. Nothing moves to design without that document existing first.

Stage Two

Design & Build: The Boundaries Come First

We design the agent's decision boundaries before writing the build that implements them — not the other way around.

Decision Design

What the agent decides alone, what it drafts for a human to approve, and what always escalates — written down explicitly, in your terms, before a line of the build depends on it.

Build Against Real Systems

We build against your actual tools wherever it's safely possible, connected through credentials you create and scope — not a stand-in environment that hides the integration problems that matter later.

This is the longest stage on most builds, and the one where scope questions genuinely come up. When they do, we raise them plainly rather than quietly absorbing or ignoring them.

Stage Three

Integrate & Test: Before It Touches a Real Case

An agent gets tested against real and edge cases before it's allowed near a live customer or record — not treated as done once it works on the obvious examples.

Real Cases

Drawn from the actual examples you shared during discovery, so the first real test is close to what the agent will actually face — not a clean hypothetical.

Edge Cases

The unusual, ambiguous or conflicting inputs that a generic test suite tends to skip — specifically to confirm the agent escalates them instead of guessing.

We show you what failed, what escalated correctly, and what needed adjusting — this stage is where the design gets refined against reality, not where it's assumed to already be right.

Stage Four

Launch & Monitor: Visible, Not Silent

Going live isn't the end of the engagement — it's when logging and review start mattering for real.

What You Can See

A record of what the agent did and why, so your team can spot-check its behaviour instead of trusting it blindly from day one.

Adjusting the Boundaries

Real usage surfaces cases discovery didn't anticipate. Monitoring is what lets the escalation rules get tightened or loosened deliberately, based on what actually happened.

From here, ongoing support covers what changes next — a connected system shifting shape, a new edge case, a boundary that needs revisiting. See pricing for how that's structured.

Frequently Asked

Process Questions

How long does discovery actually take?

Usually a single focused call plus a short follow-up, not a multi-week phase. It's long enough to map the task and the systems involved properly, short enough that it doesn't become a project of its own.

Do we need to know exactly what we want before the first call?

No. Most clients arrive with a specific, annoying, repetitive task rather than a finished spec. Discovery exists precisely to turn that into something buildable — arriving with a polished requirements document isn't a prerequisite.

What if requirements change partway through the build?

They often do, once the agent's decision boundaries get tested against real cases. Small adjustments happen inside the build. A genuine change in scope — a new system to connect, a new decision the agent needs to make — gets discussed and priced honestly rather than absorbed silently.

Do you build and test against our real systems, or a separate sandbox?

Your real systems, wherever that's safely possible — a generic demo environment tends to hide the exact integration problems that matter. Where a system is sensitive enough to warrant a staging copy first, we say so and build the testing plan around it.

What does "monitored" actually mean after launch?

Logging of what the agent did and why, visible to your team, plus a straightforward way to pause it if something looks wrong. Monitoring is what lets us and you adjust the boundaries as real usage reveals cases discovery didn't anticipate.

Start With Discovery, Not a Spec Document

Bring the task as it actually works today. We'll do the rest of the mapping on the call.