From a Repetitive Task to a Monitored Agent
Four stages, in the same order every time: discovery, design & build, integrate & test, launch & monitor. This page goes into what actually happens at each one, and what we need from you to keep it moving.
Discovery: Finding the Actual Task
Before anything gets designed, we need to understand the task as it's actually done today — not the idealised version, the real one with its exceptions.
What We Ask About
The current steps in the task, who does them today, which systems get touched, and — the part people often skip — the judgement calls a person makes along the way that a generic description would miss.
Where a Person Has to Stay in Charge
We ask directly which decisions you're comfortable handing to an agent and which ones should never leave a human's hands, rather than assuming and finding out later.
Discovery ends with a written scope: the systems involved, the decision boundaries, what counts as done, and an indicative price band. Nothing moves to design without that document existing first.
Design & Build: The Boundaries Come First
We design the agent's decision boundaries before writing the build that implements them — not the other way around.
Decision Design
What the agent decides alone, what it drafts for a human to approve, and what always escalates — written down explicitly, in your terms, before a line of the build depends on it.
Build Against Real Systems
We build against your actual tools wherever it's safely possible, connected through credentials you create and scope — not a stand-in environment that hides the integration problems that matter later.
This is the longest stage on most builds, and the one where scope questions genuinely come up. When they do, we raise them plainly rather than quietly absorbing or ignoring them.
Integrate & Test: Before It Touches a Real Case
An agent gets tested against real and edge cases before it's allowed near a live customer or record — not treated as done once it works on the obvious examples.
Real Cases
Drawn from the actual examples you shared during discovery, so the first real test is close to what the agent will actually face — not a clean hypothetical.
Edge Cases
The unusual, ambiguous or conflicting inputs that a generic test suite tends to skip — specifically to confirm the agent escalates them instead of guessing.
We show you what failed, what escalated correctly, and what needed adjusting — this stage is where the design gets refined against reality, not where it's assumed to already be right.
Launch & Monitor: Visible, Not Silent
Going live isn't the end of the engagement — it's when logging and review start mattering for real.
What You Can See
A record of what the agent did and why, so your team can spot-check its behaviour instead of trusting it blindly from day one.
Adjusting the Boundaries
Real usage surfaces cases discovery didn't anticipate. Monitoring is what lets the escalation rules get tightened or loosened deliberately, based on what actually happened.
From here, ongoing support covers what changes next — a connected system shifting shape, a new edge case, a boundary that needs revisiting. See pricing for how that's structured.
Process Questions
How long does discovery actually take?
Usually a single focused call plus a short follow-up, not a multi-week phase. It's long enough to map the task and the systems involved properly, short enough that it doesn't become a project of its own.
Do we need to know exactly what we want before the first call?
No. Most clients arrive with a specific, annoying, repetitive task rather than a finished spec. Discovery exists precisely to turn that into something buildable — arriving with a polished requirements document isn't a prerequisite.
What if requirements change partway through the build?
They often do, once the agent's decision boundaries get tested against real cases. Small adjustments happen inside the build. A genuine change in scope — a new system to connect, a new decision the agent needs to make — gets discussed and priced honestly rather than absorbed silently.
Do you build and test against our real systems, or a separate sandbox?
Your real systems, wherever that's safely possible — a generic demo environment tends to hide the exact integration problems that matter. Where a system is sensitive enough to warrant a staging copy first, we say so and build the testing plan around it.
What does "monitored" actually mean after launch?
Logging of what the agent did and why, visible to your team, plus a straightforward way to pause it if something looks wrong. Monitoring is what lets us and you adjust the boundaries as real usage reveals cases discovery didn't anticipate.
Start With Discovery, Not a Spec Document
Bring the task as it actually works today. We'll do the rest of the mapping on the call.