Playbook · Version 0.6 · June 2026

Commitment Flow Architecture — Playbook

Operating tools for diagnosing and designing conversion flows.

Companion to the Paper. The Paper makes the argument; the Playbook is the working toolkit — the diagnostic protocol, the mechanism library, audit and design templates, the automation layer, and a model-ready prompt. It assumes the vocabulary defined in the Paper; the canonical terms are collected in Section 1.1 so that humans and models use the same names.


1. How to use this playbook

The Playbook turns the framework into repeatable work. There are two entry points.

Auditing a flow that leaks — start with the diagnostic protocol (Section 2) and the per-step audit template (Section 4). You have analytics telling you where users stop; the protocol tells you how to form a hypothesis about why.

Designing a new flow — start with the design template (Section 5). You are building the sequence of commitments before there is data, so the work is to anticipate the blocker at each step and decide what the step must do to earn it.

The mechanism library (Section 3) is the lookup you reach for after diagnosis, never before. The automation layer (Section 6) and the model-ready prompt (Section 7) let an AI model run the protocol alongside you.

One rule sits above all the tools: diagnose the commitment problem before reaching for a mechanism. Most failed optimization is a correct mechanism applied to the wrong blocker.

1.1 Canonical vocabulary

These are the canonical names. Both documents and any model run defer to them.


2. The diagnostic protocol

Use this to audit an existing product. Six steps.

Step 1 — Map every commitment event. List every moment the user is asked to give something: attention, click, scroll, answer, data, email, phone, account creation, confirmation, payment, activation, message, upload, invite, renewal, upgrade, repeat purchase. For each, ask: what commitment is being requested here?

Step 2 — Estimate commitment load. For each commitment, ask how heavy the ask feels: how much time, effort, money; how sensitive the data; how reversible the action; what the user risks; what uncertainty remains. Score the ask against the load ladder (Paper §4) — attention → tap → answer → email → sensitive data → payment → recurring → B2B budget — so estimates are reproducible across auditors and runs, not a matter of intuition. The heavier the ask, the more relevance, desire, trust, and ability must precede it. If the load exceeds what any single intervention can realistically cover, flag the step for decomposition — can this one heavy ask be split into a sequence of lighter ones?

Step 3 — Identify available data. What do we know about this user at this moment — from targeting, behavior, user input, and previous steps? Is the current experience using it to make the next commitment more relevant, and is it doing so responsibly? Apply the boundary: would this use of data make the user feel understood or exposed? Intrusiveness is observable before it is fatal: opt-out or unsubscribe spikes after a personalized touchpoint, “how did you know that?” support tickets, and “creepy” appearing in reviews are the leading indicators that helpfulness no longer exceeds intrusiveness. Unused data sitting at a high-friction step is one of the most common high-leverage opportunities.

Step 4 — Diagnose the dominant blocker, in order. Work the conditions in sequence — relevance → desire → trust → ability — and treat the earliest unmet one as the real blocker. Behavioral signals:

Two disambiguations matter, and they are separate moves.

4a — Desire or resistance. When the signal points to desire, separate true desire from resistance: is the user not wanting the outcome, or wanting it but stalling on price, switching cost, or “later”? What they read or do before bouncing usually tells you — pricing and comparison behavior points to price; FAQ and refund behavior points to trust.

4b — The ordering override. The diagnosis is a structured hypothesis, not proof. In sensitive categories — health, dating, finance — trust can fail before desire forms, overriding the default order. Treat the ordering as the first hypothesis and let category and evidence beat it.

Step 5 — Match mechanism to blocker. Do not use a desire mechanism when the problem is trust, more proof when the problem is ability, more explanation when the problem is desire, or more personalization when the problem is intrusiveness. The earliest-unmet rule from Step 4 is what stops you intervening on a condition that was never the blocker.

Step 6 — Define the experiment. Each intervention becomes a testable claim with a primary metric, a guardrail metric, and a named baseline. Examples: if the blocker is relevance, a sharper problem reframe should beat generic benefit copy on progression; if the blocker is ability, friction removal should beat more persuasion on completion. State what you are beating — unstructured A/B testing, or the team’s prior experiment hit-rate — and the downstream guardrail (refunds, churn, retention, complaints) that protects against a false win. The framework does not replace experimentation; it improves the quality of the hypothesis.


3. The mechanism library

Mechanisms are tools for moving users through commitments. They are not interchangeable. Match the mechanism to the diagnosed blocker. The library is the superset — it contains every mechanism named in the Paper plus operational ones that exist only here. The families split along the gap/gate line from the Paper: the first three families add relevance, desire, or proof; the fourth removes what blocks completion.

3.1 Relevance mechanisms — used when the user does not recognize why this matters to them. Goal: make the user recognize themselves.

3.2 Desire mechanisms — used when the user understands relevance but does not want the next action enough. Goal: make the user want to continue now. First confirm it is desire and not resistance — none of these fixes a price objection or status-quo inertia.

3.3 Trust mechanisms — used when the user wants the outcome but does not believe or feel safe enough. Goal: make the user feel safe and confident enough to proceed.

3.4 Ability mechanisms — used when the user has intent but the action is too hard, unclear, or inconvenient. Goal: make the action easy enough to complete now.

3.5 Structural moves — when the ask is simply too heavy. Some asks will not clear the threshold no matter which mechanism above you apply, because the load is inherently high. These are structural, not per-blocker.

Goal: change the shape of the sequence so the heavy commitment becomes reachable.

An illustrative B2B decomposition, since B2B budget is the heaviest rung on the ladder: the contract decomposes into demo → scoped pilot with its own success metric → contract. Each rung carries its own micro-reward — a demo built on the prospect’s data, a measured pilot result — and each makes the next threshold reachable. No copy makes the contract light; the sequence does. (Constructed example, not from the case history that grounds the Paper.)

3.6 Misdiagnosis pairs. The library’s failure mode is mismatching, so here is what mismatching looks like: the same screen fixed for the wrong blocker, then the right one. All four pairs are constructed teaching examples, not cases from the Paper’s operating history.


4. Audit template

For each step in an existing flow, answer the following.

Step identity

User state

Commitment load

Data proximity

Blocker diagnosis

Loop quality

Intervention

4.1 Worked example (illustrative)

An archetypal quiz-to-subscription flow. The numbers are invented for demonstration and the case is constructed, not drawn from the Paper’s operating history; it exists so that humans and models have one fully filled template to copy the format from.

Step identity. Step 9 of 14 in onboarding: the user is asked for their date of birth. Desired next commitment: complete registration. 10,000 users reach the step weekly; 6,200 complete it.

User state. Knows the product builds a personalized plan; wants the result — eight questions already invested; doubts why birth data is needed for it; patience declining, curiosity intact.

Commitment load. Sensitive-data rung on the ladder: privacy exposure, reversibility unclear to the user, no visible justification at the ask. Decomposition flag: yes — not by splitting the ask, but by pairing it with immediate repayment.

Data proximity. Eight answers are available (goal, current habits, schedule); the current step uses none of them. The next step could become specific. Specificity tied to the stated goal would feel helpful; inferring something unstated would feel exposed.

Blocker diagnosis. In order: relevance met — the user self-selected through eight questions. Desire met — progression to step 9 with completion intent. Trust unmet — the drop concentrates at the sensitive ask, and exit sessions disproportionately open the privacy link first. Diagnosis: Trust Gap. Not resistance: no price has been shown yet.

Loop quality. The step asks and gives nothing back; the loop the quiz opened (“your plan is being built”) does not advance here. The value that would close it: immediate, visible use of the data just given.

Intervention. Transparent data exchange plus an immediate micro-reward: one line of justification at the ask (“we use this to calibrate your plan’s pacing”), and a short personalized reading generated from the date of birth and prior answers, delivered before the next ask. Hypothesis: justification plus immediate repayment beats the bare ask on step completion. Primary metric: step conversion. Guardrails: downstream paywall conversion and early refund rate — a forced completion that erodes trust would surface there. Baseline: the team’s prior onboarding experiment hit-rate.


5. Design template

Use this when building a new growth flow.

Product and user

Commercial architecture

Step architecture — for every step:

Testing architecture — for every major intervention:


6. The automation layer

The framework is well suited to AI-assisted growth work because it turns vague optimization into structured diagnosis and generation. The core model capabilities map directly to the protocol: funnel and commitment-event mapping, per-step blocker diagnosis, commitment-load scoring against the ladder, data-proximity analysis with the intrusiveness boundary, variant generation matched to the diagnosed blocker, and experiment design with guardrails and named baselines.

The operating constraint is the same as for a human: the model should not simply make copy more persuasive. It should diagnose the commitment problem first. The prompt in Section 7 enforces that order. The worked audit in Section 4.1 is the few-shot example for the long form; the condensed run below shows the short form.

A condensed run of the Section 7 prompt, on a described screen rather than a live one:

Input: “Quiz app, results paywall. The user finishes a 12-question assessment, sees a blurred results graphic under the headline ‘Your profile is ready,’ and a $39/year subscribe button. Drop-off at this screen is 91%.”

Output (abbreviated):

  1. Commitment requested: annual payment.
  2. Load: high — money compounded by a year of time; recurring rung on the ladder.
  3. Data available: twelve answers; the screen uses none of them — the blur communicates withholding, not insight.
  4. Earliest unmet condition: trust is the working hypothesis. Desire is evidenced by quiz completion, but nothing on the screen demonstrates the result is real or specific to this user. Check whether resistance (price) competes, via exit reads.
  5. Evidence: hypothesis, not proof — if exits read pricing or comparison content, re-diagnose toward price.
  6. Minimum necessary intervention: deliver one real, specific insight from the user’s answers before the ask — a partial reveal, not a bolder headline. 7–8. Variants: one-insight reveal vs. mechanism explanation vs. monthly-entry decomposition of the annual ask. Hypothesis: real partial value beats withheld value on paid conversion.
  7. Primary metric: paywall conversion. Guardrail: 30-day refund rate. Baseline: the current screen and the team’s prior hit-rate.
  8. Risks: the blur is a fake-loop pattern if the “profile” is not real — loop-must-be-real applies.
  9. Product type: self-insight.

7. Model-ready skill prompt

Use the following as the starting prompt for applying Commitment Flow Architecture with an AI model. It is the working spec for the skill; it assumes the canonical vocabulary in Section 1.1, and Section 4.1 is its long-form few-shot example.

Activation prompt

You are a growth strategist using Commitment Flow Architecture. Analyze the provided funnel step, screen, copy, flow, or screenshot. Do not simply make the copy more persuasive — diagnose the commitment problem first.

For each step:

  1. Identify the user commitment being requested.
  2. Estimate its commitment load against the load ladder: attention, effort, data sensitivity, money, emotional risk, reversibility, uncertainty. If the load is high, note whether the ask should be decomposed into a sequence of lighter commitments.
  3. Identify the data available about the user at this moment, and whether the experience uses it responsibly — helpful, not intrusive.
  4. Diagnose the dominant blocker by working the conditions in order — Relevance Gap, Desire Gap, Trust Gap, Ability Gate — and treating the earliest unmet one as the real blocker. If the apparent blocker is desire, separate true desire from resistance (price, switching cost, “later”).
  5. State the behavioral evidence for the diagnosis, and flag that it is a hypothesis. Note any sensitive-category reason the default order might not hold.
  6. Select the minimum necessary intervention — sometimes a mechanism, sometimes removing a step, sometimes resequencing the ask. If removal resolves the blocker, removal beats addition.
  7. Suggest 3 to 5 variants that match the diagnosed blocker.
  8. Define the experiment hypothesis.
  9. Recommend the primary metric, a guardrail metric, and the baseline the test must beat (unstructured A/B, or the team’s prior hit-rate).
  10. Flag ethical or trust risks: over-personalization, fake loops, sensitive data, cyclical or compulsion-loop dynamics, and short-term conversion wins that may damage retention.
  11. If useful, map the step to a product type — self-insight, dating / social discovery, conversational AI, AI output tool, fitness / wellness, B2B SaaS, or another relevant category.

This Playbook is the operational companion to the Paper. The product-type walkthroughs (self-insight, dating / social discovery, conversational AI, AI output tool, fitness / wellness, B2B SaaS) can be added here as application templates once the Paper is locked; conversational AI and AI output tools are first in the queue — most searched, least served by the adjacent literature, and the place where cyclical-loop ethics bite hardest.