Product owners, engineers, and domain experts repeat reviews every release
Your team repeatedly inspect outputs without turning their judgment into a reusable capability.
AI quality systems for product and engineering leaders
Give product and engineering teams a shared, executable definition of quality, so leaders can make release decisions with evidence instead of demos, opinions, and repeated manual review.
A practical half-day workshop
~4-hours · next workshop date October 21st 9AM PT · $400 CAD per person
Product owners define the outcome users need. Engineers make behaviours testable. Domain experts define what good looks like in practice. Together, they agree on expected behaviours; observable checks provide shared evidence for release decisions.
The leadership risk
When expectations live in the heads of product owners, engineers, and domain experts, every release creates another round of manual review. Product and engineering debate whether the experience is good enough, while leaders lack consistent evidence for deciding whether to ship.
Your team repeatedly inspect outputs without turning their judgment into a reusable capability.
Words such as helpful, accurate, and safe hide different expectations until the team makes them observable and shared.
The team ships without a common view of which customer-critical behaviors are defined, tested, or still exposed.
Start with the user’s goal and the situations your feature needs to handle. Define what a successful response must do, which variations are acceptable, and what must never happen. Those expectations become specifications for repeatable tests.
Different valid wording can satisfy the same requirements. The goal is preserving important behaviours under relevant variations—not identical text or a guarantee of every future outcome.
What your team takes away
Capture expert judgment in a quality system the team can reuse, inspect, and improve.
Book the workshopThis practical working session produces implementation-ready specifications and a plan, not a fully integrated harness or completed CI setup.
Illustrative example: a scheduling assistant helping a user arrange an appointment.
Fuzzy expectation: “The assistant should be helpful and book meetings effectively.”
One scheduling feature → three required behaviours
Ask for required information before recommending an action.
Expectation definedUse current tool results, not invented availability. If the tool fails, explain that availability could not be checked.
Examples preparedObtain the required confirmation before committing.
Automated check pendingIllustrative pass: asks for missing account type, then uses supplied constraints.
Illustrative fail: invents missing context and recommends an irreversible action.
A defined expectation is not a passing automated test. Reproduce a newly discovered production failure with its context and add it to regression coverage.
~4-hours · next workshop date October 21st 9AM PT · $400 CAD per person
Visible input, hidden context, state, tools, and model variation. Input map.
User goals, expected behaviour, boundaries, and acceptable variation. Behavioral Map v1.
Select 3–5 high-value behaviours and define scenarios, fixtures/mocks, expected actions, and evaluator choices.
Determine where tests and production learning fit the delivery process.
Prioritize gaps, owners, and the next 30 days.
Includes a 10-minute break. Approximately four hours overall.
Bring one AI-enabled feature, a few examples of useful or problematic behaviour, and the context behind them. Product and engineering colleagues working together are a good fit. Sanitized or synthetic examples are welcome.
Book the workshop
Why Evan · Startup 0-to-1 & Fortune 50 AI
After his HealthTech startup was acquired by Virgin Pulse in 2019, Evan went on to work on agile delivery at enterprise scale and, in the Bay Area, on scaling the reliability of consumer-facing AI experiences.
That experience shapes his approach to a familiar leadership problem: how do you keep moving when product owners, engineers, and domain experts each have a different view of “good enough”? His focus is turning those perspectives into explicit behaviours and repeatable evidence, so release confidence does not depend on another round of subjective review.
The workshop brings that approach to one real feature: agree on what good looks like, define how to evaluate it, and prioritize implementation. When hands-on delivery is needed, Prototype Devshop can help turn that plan into an executable harness and regression workflow.
Half-day workshop
$400 CAD per person
Work through one AI feature with Evan. Agree on the behaviours that matter, turn 3-5 priorities into evaluation specifications, and leave with a Behavior Harness Starter Pack your team can use to begin implementation.
Book the workshopDefine the behaviours that matter, then build repeatable checks that make release decisions clearer.
Get hands-on implementation of an AI Behavior Harness for your feature. Done-for-you engagements start at $15,000 CAD.
Explore done-for-you implementationUse three-line sticky notes to define who the user is, where they are starting, and what a useful response must do. Turn one broad expectation into testable behaviours and see where evidence is still missing.
Download the free starter guideFree four-page introduction, worked example, and sticky-note worksheet. No email required. The paid workbook takes those notes onto a simple canvas and through a baseline harness implementation.
Product and engineering teams building or improving an AI feature. Product owners, engineers, and domain experts each bring a useful perspective on what the feature should do and how to judge its behaviour. A feature in development and a clear user goal are enough to work from.
A Behavior Harness Starter Pack for one feature: a map of its inputs and expected behaviours, 3-5 evaluation specifications, and priorities for implementation. It gives your team a concrete starting point for turning quality expectations into repeatable checks.
You can contribute to defining behaviours and deciding what good looks like without writing code. Engineering knowledge helps when planning how to capture outputs, implement checks, and fit them into your team's workflow. No specific evaluation platform is required.
Choose one AI feature or workflow you want to evaluate. Bring a short description of who uses it, what they are trying to accomplish, and a few examples of the behaviour you expect. Use examples you can share with the group.
The half-day workshop gives you the specifications and implementation priorities for your first harness. You will define what to check, the scenarios to recreate, and what passing looks like. Your team can then use that work to implement the checks in your own system.
Your team has a clear starting point for implementation. Use the Starter Pack to build the highest-priority checks, compare results after changes, and add new scenarios as you learn from real usage. Each specification connects a behaviour to the situation being tested and what passing looks like, so engineering can move from discussion to implementation. If you would like us to handle that work, explore our done-for-you implementation service.
The $16 CAD workbook lets you work through the canvas and baseline harness example at your own pace. The workshop adds Evan's guidance as you apply the method to your feature, work through ambiguous expectations, and decide which behaviours to prioritize.
Work through one feature and leave with the behaviours to protect, the first eval specifications, and a practical implementation plan.
Book the workshop