Frameworks

Frameworks I build with

Two sets. AI-native methods from products I built, where agents do the work and people keep the judgment, and the product management frameworks I rely on, each shown through a real company that used it well. Every one ships with a template you can copy into a doc and use today.

AI-native

AI-native frameworks

Methods from products I built, for teams where agents do the work and people keep the judgment.

The problem

Without a gate, one bad draft reaches a customer before anyone sees it, and trust in the whole system goes with it. With a gate that never loosens, the reviewer becomes the bottleneck and people start approving without reading.

How to run it

  1. Split reads from writes

    List every action each agent can take. Reads run freely, and every write (send, post, pay, change a record) becomes a proposal that waits in a queue.

  2. Put the evidence on the card

    Each proposal shows the draft, the facts it used with links to their source, and approve, edit and reject in one click. Nothing is written until the person confirms.

  3. Commit exactly once

    Claim the approval inside a database transaction, so a double click or a retried request fires the effect one time. A rejected proposal never changes state.

  4. Start every write on approve to send

    Give each action type a rung: draft only, approve to send, auto after 24 hours unless cancelled, or autonomous. No write agent starts autonomous.

  5. Demote fast, promote on purpose

    Two consecutive rejects step that action type down one rung. Approvals never promote on their own; moving up is a decision the owner makes and logs.

  6. Keep some actions off the ladder

    Name the actions that always need a person however well the agent does, such as clearing a conflict of interest or giving advice that needs a license.

Done when

  • no code path lets an agent write, send or pay without a committed proposal.
  • a double click and a replayed request on one approval each produce exactly one effect, proven by a test.
  • every action type has a recorded rung, and a test shows two consecutive rejects demote it.
  • each proposal card links every figure in the draft to the record it came from.

How teams get it wrong

  • Teams gate the whole agent instead of each action type, so one risky action keeps every safe one stuck in review; give each action type its own rung.
  • The approve click and the send run as two separate steps, so a retry sends twice; claim the decision and fire the effect in one transaction.
  • Approvals quietly promote the agent until it runs unattended; make every promotion a logged decision by the owner.
Template Markdown, ready to paste into a doc
# Agent approval register

Agent: ____
Approver (named person): ____
Review queue lives in: ____
Date set up: ____

## Action types
| Action | Read or write | Starting rung | Always needs a person | Sources shown on the card |
|---|---|---|---|---|
| ____ | ____ | approve to send | yes / no | ____ |
| ____ | ____ | approve to send | yes / no | ____ |
| ____ | ____ | draft only | yes / no | ____ |

Rungs: draft only, approve to send, auto after 24 hours unless cancelled, autonomous.

## Rules
- Demote one rung after ____ consecutive rejects (default 2).
- Promote only when the unedited approval rate holds above ____% for ____ days. Log who decided.
- Actions that never leave approve to send: ____

## The proposal card shows
- [ ] The full draft, editable
- [ ] Every figure linked to its source record
- [ ] Approve, edit and reject in one click
- [ ] One line on why the agent proposed it

## Tests before launch
- [ ] A double click on approve fires the effect once
- [ ] A replayed approval request fires the effect once
- [ ] A rejected proposal changes nothing
- [ ] Two rejects in a row demote the rung

## Promotion log
| Date | Action | From rung | To rung | Decided by | Evidence |
|---|---|---|---|---|---|
| ____ | ____ | ____ | ____ | ____ | ____ |

The problem

A model asked to decide gives different answers to the same input, and nobody can explain why a lead was rejected or a number moved. A rule written only in a prompt is a request, and an odd input or a model update can talk past it.

How to run it

  1. Split the feature into words and decisions

    Write two lists. Language work covers understanding the caller, extracting fields and drafting prose. Decisions cover who qualifies, what it costs, which week runs short and what gets sent.

  2. Move every decision into a pure function

    Each decision becomes code with no I/O and a test table of hand-calculated cases, edge cases included. The server action only reads, calls the function and writes.

  3. Check the model at the boundary

    Give the model a schema and validate its output before use. Malformed output is retried or dropped, and never saved.

  4. Hand the model finished numbers

    The model explains figures the code already computed. Scan its output for any figure that is not in its inputs and reject the draft if one appears.

  5. Write hard rules as code

    Rules such as never giving legal advice or never clearing a conflict live in code paths with tests, so no prompt or input can switch them off.

  6. Ship a fallback with every model call

    Each call has a timeout and a deterministic path, such as a template, so the feature still does its core job with no API key or during a provider outage.

Done when

  • the same inputs produce the same decisions on ten repeat runs.
  • every decision function has a fixture test with hand-calculated expected values.
  • a test shows malformed model output is never written to the database.
  • the feature completes its core job with the model API key removed.

How teams get it wrong

  • The prompt says to qualify only certain callers and the team treats that as enforcement; move the rule into code and test it.
  • The model writes fresh test questions or criteria on every run, so results drift for no real reason; fix the inputs with rules and run extraction at zero temperature with a fixed seed.
  • The fallback is written last and never exercised; force a provider outage in CI and assert the output.
Template Markdown, ready to paste into a doc
# Words and decisions split: ____ (feature)

## Language work the model does
- ____ (for example: understand the caller)
- ____ (for example: extract the intake fields)
- ____ (for example: draft the reply)

## Decisions code makes
| Decision | Function | Hand-calculated fixtures | Owner |
|---|---|---|---|
| ____ | ____ | ____ cases | ____ |
| ____ | ____ | ____ cases | ____ |
| ____ | ____ | ____ cases | ____ |

## Model boundary
Output schema: ____
On malformed output: retry ____ times, then ____ (never save it)
Figures the model may mention: only ____
Temperature and seed for extraction: ____

## Hard rules in code
- [ ] ____ (for example: never gives legal advice)
- [ ] ____ (for example: flags a conflict, never clears it)

## Fallbacks
| Model call | Timeout | Deterministic fallback | Tested with key removed |
|---|---|---|---|
| ____ | ____ s | ____ | yes / no |
| ____ | ____ s | ____ | yes / no |

## Done checks
- [ ] Same inputs, ten runs, same decisions
- [ ] Malformed output never written, proven by a test
- [ ] Core job works with the API key removed

The problem

Models write fluent claims that no record supports, and a reader cannot tell the invented line from the true one. One wrong dollar figure in a customer email costs more trust than a hundred correct ones earn.

How to run it

  1. Define what counts as a source

    Name the record types a claim may cite, such as a commit, an invoice, a collected signal or a document section, and give each a stable id.

  2. Store source ids with every claim

    Each drafted claim carries the ids it rests on. A claim with an empty list is dropped before review.

  3. Rebind after drafting

    Look up every cited id again in the set you actually collected. Remove any claim whose id does not resolve, and recompute scores from the surviving evidence only.

  4. Check every figure

    Pull every number and amount out of the draft and confirm each one appears in the source facts, ignoring formatting. If one does not, reject the draft and fall back to a plain template.

  5. Label how strong the evidence is

    Show a tier beside each claim: verified from a primary record, corroborated by a second source, or inferred. Never upgrade an unknown.

  6. Show the receipt

    Link each claim to its source in the interface, and let every total open into the line items that add up to it.

Done when

  • a test that injects a fake source id shows the claim removed before review.
  • a test shows a draft with an invented dollar figure is rejected, including figures written in words.
  • every published claim shows its tier and opens its source in one click.
  • each headline total reconciles exactly to its listed line items.

How teams get it wrong

  • Teams ask the model to include sources and trust the links it writes; resolve every id against records you collected yourself.
  • The figure check only catches amounts with a dollar sign, and the model writes 500 dollars; test every way a figure can be written.
  • Thin evidence is shown with the same confidence as strong evidence; show the tier and pull thin evidence toward the base rate.
Template Markdown, ready to paste into a doc
# Source register: ____ (feature)

## What counts as a source
| Source type | Id format | Collected by | Default tier |
|---|---|---|---|
| ____ | ____ | ____ | verified / corroborated / inferred |
| ____ | ____ | ____ | ____ |
| ____ | ____ | ____ | ____ |

## Claim rules
- Every claim stores its text, source ids, tier and drafted time.
- A claim with no source id is dropped before review.
- Cited ids are looked up again in: ____ (the store you collected)
- Scores are recomputed from surviving evidence only.

## Figure check
Figures a draft may contain come from: ____
Formats the check must catch:
- [ ] $1,234.56 and $ 1234
- [ ] 500 dollars and USD 2,000
- [ ] Percentages and dates: ____
On a failed check, reject the draft and use: ____ (template)

## Reader view
- [ ] Each claim links to its source
- [ ] Each claim shows its tier
- [ ] Each total opens into its line items and reconciles exactly

## Tests
- [ ] A fake source id is removed
- [ ] An invented figure rejects the draft
- [ ] Grounding rate on the eval set is 100%

The problem

Without a fixed eval set, a prompt tweak that fixes one case quietly breaks five others, and nobody notices until a user does. Teams also test only the cases where the model should produce something, so it learns to always produce something.

How to run it

  1. Write the eval set in week one

    Hand-label cases before you tune a prompt. Cover both sides of the boundary, including cases where the correct output is nothing.

  2. Mark each metric as gate, floor or report

    Gates hold exactly, such as 100% grounding and zero banned phrases. Floors hold a minimum, such as an 85% clean-draft rate. Everything else is reported.

  3. Stand in for the human metric

    Your real target is something like drafts approved without edits. Write a deterministic proxy CI can check: a length band, a concrete token from the evidence, no banned phrase and no near-duplicates.

  4. Make it run anywhere

    Run the rails against a deterministic path so the eval runs offline and identically in CI, and run the live model path on a schedule.

  5. Fail the build

    The eval exits non-zero on any gate breach, any floor miss, or any case that behaves differently from its label, and CI blocks the merge.

  6. Grow it from real misses

    Every bad output a user reports becomes a labeled case. Raise the floors as the set grows.

Done when

  • the eval set includes labeled cases where the correct output is nothing.
  • a planted regression, such as a banned phrase in a template, fails CI.
  • the eval runs in CI with no model API key.
  • each metric is written down as a gate, a floor or a report, with its number.

How teams get it wrong

  • The eval set holds only happy cases, so the model passes by always answering; label silent cases on purpose.
  • The eval prints a score but never fails the build; make it exit non-zero and wire it into CI.
  • Teams tune the prompt until the set passes and then stop adding cases; add every production miss as a new case.
Template Markdown, ready to paste into a doc
# Eval gate: ____ (AI feature)

Owner: ____
Runs on every pull request, plus a live model run every ____

## Cases
| Group | Count | Expected behavior |
|---|---|---|
| Must produce output | ____ | grounded, above the threshold |
| Must stay silent | ____ | no output |
| Edge and abuse cases | ____ | ____ |
| From production misses | ____ | ____ |

## Metrics
| Metric | Type | Bar |
|---|---|---|
| Grounding rate | gate | 100% |
| Banned-phrase leaks | gate | 0 |
| Cases matching their label | gate | all |
| Clean-draft rate | floor | ____% |
| ____ | report | ____ |

## Clean-draft checks (proxy for approved unedited)
- [ ] Length between ____ and ____ characters
- [ ] Contains a concrete token from its evidence
- [ ] No banned phrase
- [ ] Not a near-duplicate of another draft

## Wiring
- [ ] Runs with no API key on a deterministic path
- [ ] Exits non-zero on any gate or floor miss
- [ ] CI blocks the merge on failure
- [ ] A planted regression proved it fails

## Human target
Drafts approved without edits: ____%, measured from ____

The problem

Many flows send the email or push to the CRM first and save later, so when a provider is down the lead disappears and nobody knows it existed. Retries then send the same confirmation twice or create a duplicate record.

How to run it

  1. Map what happens after capture

    List every downstream step: notification email, CRM or case-system export, calendar booking, text message and model call. Treat each one as best effort.

  2. Persist before anything else

    Save the record, even a partial one, to one ledger with status new before the first downstream call. A hang-up halfway through still leaves a record.

  3. Key every side effect

    Each email, booking, export and payment carries an idempotency key, so a replayed webhook or retried job does nothing the second time.

  4. Give each integration a fallback

    If the export fails, email the record; if that fails too, park it in a retry queue that shows on the dashboard.

  5. Test with each dependency down

    Switch off each provider in a test and assert that the record is saved and surfaced and that nothing was sent twice.

  6. Refuse instead of pretending

    If the record cannot be stored, tell the person it did not go through. Never show success for a submission you did not keep.

Done when

  • a test with every downstream provider failing still finds the record saved and visible.
  • replaying the same webhook twice produces one record and one message.
  • every integration has a written fallback and a test that takes it.
  • the form shows an error if the save fails.

How teams get it wrong

  • The record is written when the call or form ends, so a hang-up or timeout loses it; save the partial record first and update it as you go.
  • Email is the only copy of the lead; keep the database as the source of truth and treat email as a notice.
  • Retries go in without idempotency keys, so an outage turns into duplicate texts; key every side effect before you add retries.
Template Markdown, ready to paste into a doc
# Capture flow audit: ____ (form, call or webhook)

## Capture
Record saved to: ____ (table or ledger)
Saved at step: ____ (before any downstream call)
Partial records kept: yes / no
If the save fails, the person sees: ____

## Downstream steps
| Step | Provider | Idempotency key | Fallback | Retry queue |
|---|---|---|---|---|
| Notification email | ____ | ____ | ____ | ____ |
| CRM or case export | ____ | ____ | email with the record | ____ |
| Booking | ____ | ____ | ____ | ____ |
| Confirmation text | ____ | ____ | ____ | ____ |
| Model call | ____ | ____ | ____ | ____ |

## Failure tests
- [ ] Each provider down alone: record saved and visible
- [ ] All providers down: record saved and visible
- [ ] Same webhook replayed twice: one record, one message
- [ ] Save fails: the person sees an error, and no success screen
- [ ] Failed items show on the dashboard with a retry action

## Weekly check
Inbound events in provider logs: ____
Records saved: ____
Gap: ____ (target 0)

The problem

An answer engine often names two or three companies and sends no click to anyone else, and a brand that gets skipped never hears about it. Without a measure per question, teams publish more content at random and cannot tell which question they lost, or to whom.

How to run it

  1. Write the buyer question set

    List the questions buyers ask before they choose, in their words, built by fixed rules from your own pages and category. Keep the set versioned so runs compare.

  2. Ask the real engines

    Send each question to each engine you track and store the full answer. Never let one model play the part of another engine.

  3. Record three things per answer

    Whether you were named by brand or domain, who was named instead, and which pages the engine cited.

  4. Score it honestly

    Divide citations by the question and engine pairs that returned an answer, times 100. If an engine is down, shrink the denominator rather than counting a miss.

  5. Ship one fix per lost question

    Add a short direct answer to the question on your own page, earn mentions on the third-party pages engines already cite, or refresh dates on pages that go stale, like pricing.

  6. Prove the lift

    Re-run the same questions after a fix ships and log the before and after for that question next to the edit behind it.

Done when

  • the question set is versioned and repeat runs on unchanged inputs give the same score.
  • every result row stores the engine's raw answer and the pages it cited.
  • an engine outage lowers the denominator and leaves the score honest.
  • each shipped fix is linked to the question it targets and a before and after reading.

How teams get it wrong

  • A model writes fresh test questions on every run, so the score moves for no reason; build questions by fixed rules and cache results for the day.
  • Teams simulate engine answers to save money and get a score that swings from run to run; query the real engines and cap the spend instead.
  • Teams chase the overall number; work the single biggest lost question each week.
Template Markdown, ready to paste into a doc
# Share of Answer tracker: ____ (brand)

Domain: ____
Engines tracked: ____
Run time each day: ____

## Buyer question set (version ____)
| # | Question in the buyer's words | Source page or rule | Stage (compare, choose, price) |
|---|---|---|---|
| 1 | ____ | ____ | ____ |
| 2 | ____ | ____ | ____ |
| 3 | ____ | ____ | ____ |

## Daily result per question and engine
| Question | Engine | Named (yes / no) | Named instead | Pages cited |
|---|---|---|---|---|
| ____ | ____ | ____ | ____ | ____ |
| ____ | ____ | ____ | ____ | ____ |

Share of Answer = citations / pairs that returned an answer x 100 = ____
Engines unavailable today (removed from the denominator): ____

## This week's fix
Lost question: ____
Named instead: ____
Fix type:
- [ ] Direct answer passage on our page: ____ (URL)
- [ ] Mention on a page the engine already cites: ____
- [ ] Refresh a stale page: ____
Shipped on: ____
Before: ____  After: ____

The problem

Point-based scoring adds up activity, so a company with eight duplicate posts outranks one with two real buying triggers, and reps spend their week on noise. Scores also hide what they cannot see, so a confident number built on one signal reads like one built on ten.

How to run it

  1. Keep fit and intent apart

    Score fit (whether the account could buy) and intent (whether it is moving now) as two numbers. Decide with both in hand and never add them into one total first.

  2. Weight by cost to the buyer

    For each signal type, rate what it cost to produce: a senior person's time counts more than an anonymous visit, and a paid evaluation more than a blog read. Cheap traces get close to zero weight.

  3. Saturate repeats and decay with time

    The first instance of a signal type moves the score most, and each repeat adds less. Give each type a half-life, so an account that goes quiet drifts back to the base rate.

  4. Let bad news subtract

    Count disconfirming signals, such as a closed evaluation, a competitor adopted or a hiring freeze, as negative evidence.

  5. Qualify before a rep sees it

    Grade each account on five questions: right buyer, fresh trigger, able to pay, reachable, and a concrete need you meet. Hard disqualifiers drop it, and only A and B grades reach the list.

  6. Ship the blind spots with the score

    Show the evidence tier, the age of the newest signal, and a flag when the score rests on one person. Openers cite public signals only, never tracked behavior.

Done when

  • a test shows eight duplicate signals of one type scoring below two different strong triggers.
  • fit and intent are stored and shown as separate fields.
  • every score on the list links to the signals behind it, with source and date.
  • a test shows an account with no new signal for one half-life scoring lower than before.

How teams get it wrong

  • Teams count signals, so volume beats quality; saturate each type and weight by cost.
  • Low-fit, high-intent accounts get thrown away; route them to a lighter play instead of deleting them.
  • Outreach mentions what the scorer saw, like repeat pricing-page visits; reference only public signals the buyer chose to publish.
Template Markdown, ready to paste into a doc
# Signal weight sheet: ____ (ICP)

Fit in one sentence: ____

## Signal types
| Signal | Public source | Cost to the buyer (low / medium / high) | Half-life (days) | Confirms or disconfirms |
|---|---|---|---|---|
| Funding filing | ____ | ____ | ____ | confirms |
| Role posted that needs what we sell | ____ | ____ | ____ | confirms |
| Public request for a provider | ____ | high | ____ | confirms |
| Competitor adopted | ____ | ____ | ____ | disconfirms |
| ____ | ____ | ____ | ____ | ____ |

Repeats of one type: the first ____ count fully, then each adds less.

## Five-question qualification
- [ ] Right buyer: ____
- [ ] Fresh trigger (a change, not a state): ____
- [ ] Able to pay: ____
- [ ] Reachable (verified domain and a named buyer): ____
- [ ] Concrete need we meet: ____
Hard disqualifiers: ____
Grade: A / B / C / D (only A and B reach the weekly list)

## Shown with every score
- [ ] Evidence tier: verified, corroborated or inferred
- [ ] Date of the newest signal
- [ ] Flag if one person or one signal carries the score
- [ ] Opener cites a public signal only

The problem

Repeat work hides inside everyone's week, so nobody sees its total cost and it never gets fixed. Teams that automate too early build workflows for tasks that never come back, and teams that wait too long spend months on chores a tool could draft.

How to run it

  1. Log repeats for two weeks

    Each time someone does a task they have done before, add a line with the task, its trigger, the minutes spent and the tools touched.

  2. Count to three

    A task that shows up three times qualifies. Rank the qualifiers by minutes per month.

  3. Write the steps as done today

    Record the inputs, the outputs, and the one or two points where judgment is needed. Those points stay with a person.

  4. Use the lightest tool that fits

    Use a text expander such as Text Blaze for typed replies, Zapier or n8n for handoffs between apps, and Claude for drafting and summarizing.

  5. Keep a person on the output

    The workflow drafts or prepares, and a named person confirms before anything is sent or changed. Nothing is deleted automatically.

  6. Review monthly

    Compare minutes saved with time spent fixing the workflow, and retire any that cost more than they return.

Done when

  • the repeat log exists and every entry has a trigger and minutes spent.
  • each automated task has a named owner who confirms its output.
  • each workflow records how often it ran and how often a person had to fix it.
  • every task that hit three repeats has a recorded decision to automate it or keep it manual.

How teams get it wrong

  • Teams automate the judgment step along with the busywork; automate the preparation and keep the decision with a person.
  • A workflow breaks quietly and nobody owns it; give every workflow a named owner and an alert on failure.
  • Teams buy a new platform for one chore; start with the tools already in use.
Template Markdown, ready to paste into a doc
# Repeat work log: ____ (team)

Log period: ____ to ____

| Task | Trigger | Who | Minutes | Tools touched | Times seen |
|---|---|---|---|---|---|
| ____ | ____ | ____ | ____ | ____ | ____ |
| ____ | ____ | ____ | ____ | ____ | ____ |
| ____ | ____ | ____ | ____ | ____ | ____ |

## Qualified (seen three times or more)
| Task | Minutes per month | Judgment step that stays with a person | Tool | Owner |
|---|---|---|---|---|
| ____ | ____ | ____ | Text Blaze / Zapier / n8n / Claude | ____ |
| ____ | ____ | ____ | ____ | ____ |

## Workflow spec: ____
Trigger: ____
Inputs: ____
Steps the workflow does: ____
Output it drafts: ____
Person who confirms: ____
On failure, alert: ____

## Monthly review
- [ ] Runs this month: ____
- [ ] Times a person had to fix it: ____
- [ ] Minutes saved: ____
- [ ] Keep, change or retire: ____
Product management

Product management frameworks

The classics I rely on, each shown through a real company that used it well, with sources.

Case study · Superhuman

Two years of building, and only 22% of users would be very disappointed to lose it.

The problem
Superhuman started coding its email client in 2015, and in summer 2017 founder Rahul Vohra went looking for a way to measure product-market fit. He asked users how they would feel if they could no longer use Superhuman, and 22% said very disappointed. That was well short of the 40% bar Sean Ellis had found by benchmarking nearly 100 startups.
What they did
Vohra found the very disappointed group was mostly founders, managers, executives and business development people, and focusing on them alone lifted the score to 33%. The team described that customer as one persona who handles 100 to 200 emails a day, politely set aside users who would not be disappointed, and split the roadmap between doubling down on speed and removing blockers such as the missing mobile app, integrations and calendaring.
What happened
Three quarters later the score had reached 58%, and the team tracked it weekly and quarterly as its primary OKR.

Lesson for your teamMeasure fit inside the segment that loves you most, and spend half your roadmap on the people who almost do.

Sources: First Round Review: How Superhuman built an engine to find product/market fit Superhuman blog: the same essay by Rahul Vohra

How to run it

  1. Survey people who have used it for real

    Send four questions to users who have used the product recently and more than once: how they would feel without it, who would benefit most, the main benefit they get, and how to improve it.

  2. Score the very disappointed share

    Divide very disappointed answers by all answers to the first question. Sean Ellis set the bar at 40%, and below it you should treat fit as unproven.

  3. Find the segment that loves it

    Tag each respondent by role or use case and recompute the score for each group. Describe the highest-scoring group as one specific person, your high-expectation customer.

  4. Listen to the right near-fans

    Set aside people who would not be disappointed. Among the somewhat disappointed, keep those whose main benefit matches what your fans love, and read what holds them back.

  5. Split the roadmap in half

    Spend half the work deepening the main benefit fans named and half removing what blocks the near-fans. Write both lists down before you start.

  6. Track it as the main goal

    Rerun the survey on a fixed cadence with fresh respondents, and report the score for the target segment every time.

Done when

  • the very disappointed share is computed for all respondents and for each tagged segment, with the response count beside it.
  • the high-expectation customer is written as one paragraph the whole team can quote.
  • every roadmap item is labeled as deepening the main benefit or removing a blocker for near-fans.
  • the next survey run has a date and an owner.
Template Markdown, ready to paste into a doc
# Product-market fit survey: ____ (product)

Owner: ____
Sent to: users active in the last ____ days who have used it at least ____ times
Responses collected: ____
Next run: ____

## The four questions
1. How would you feel if you could no longer use ____? (Very disappointed / Somewhat disappointed / Not disappointed)
2. What type of people do you think would most benefit from ____?
3. What is the main benefit you receive from ____?
4. How can we improve ____ for you?

## Scores (bar: 40% very disappointed)
| Segment | Responses | Very disappointed | Score |
|---|---|---|---|
| All respondents | ____ | ____ | ____% |
| ____ | ____ | ____ | ____% |
| ____ | ____ | ____ | ____% |

## High-expectation customer
One paragraph, in plain words: ____

## Main benefit fans named (top three, in their words)
1. ____
2. ____
3. ____

## Roadmap split
| Deepen the main benefit (for fans) | Remove a blocker (for near-fans with the same main benefit) |
|---|---|
| ____ | ____ |
| ____ | ____ |

Set aside: answers from people who would not be disappointed.
Case study · Amazon (Kindle)

Ebooks in 2004 meant reading on a PC, with a thin catalog and high prices.

The problem
As Amazon veterans Colin Bryar and Bill Carr tell it, ebooks were close to a nothing business in 2004. You could read one only on a PC or a Mac, prices were high, and about 15,000 titles were available against hundreds of thousands in print, so there was no good reason to buy one.
What they did
Bezos had teams write the press release at the start of a project instead of the end, and the Kindle team worked backwards from the reading experience. That fixed details such as a device that arrives already linked to your account with past purchases downloaded, and an always-on connection so you could shop, download and read on one device.
What happened
Kindle launched in November 2007 after more than three years of work, with more than 90,000 titles, no computer needed, no monthly wireless fee, and books that downloaded in under 60 seconds.

Lesson for your teamWrite down the customer experience you would be proud to announce, then hold engineering to it.

Sources: First Round Review: lessons from Working Backwards by Colin Bryar and Bill Carr Amazon press release: Introducing Amazon Kindle (2007) About Amazon: an insider look at Working Backwards

How to run it

  1. Name the customer and the problem

    In one sentence each, say who the customer is and what is hard for them today, in words they would use.

  2. Write the press release

    Keep it under one page: headline, problem, solution, a customer quote and how to get started. If it does not describe something meaningfully faster, easier or cheaper than what exists, stop here.

  3. Write the FAQ

    Answer the questions a customer would ask, then the internal ones on cost, dependencies, risks and what has to be true. Keep it to five pages or less.

  4. Read it with skeptics

    Have everyone read the document in the meeting before anyone speaks, then challenge it. Rewrite it until the hard questions have plain answers.

  5. Turn the release into requirements

    Every promise in the press release becomes a requirement the build must meet. Anything the release does not promise waits.

  6. Shelve most of them

    Expect most PR/FAQs never to launch. Keep the ones that hold up under questioning and fund the best of them.

Done when

  • the press release fits on one page and names a specific customer and problem.
  • the FAQ answers price, the biggest risk and the dependencies in five pages or less.
  • a reader outside the team can say in one sentence why a customer would switch.
  • every requirement in the build plan traces to a line in the press release.
Template Markdown, ready to paste into a doc
# PR/FAQ: ____ (product)

Author: ____
Review meeting: ____

## Press release (one page or less)
Headline: ____ launches ____ so that ____ can ____
Subhead (who it is for and the one benefit): ____
The problem today, in the customer's words: ____
How the product removes it: ____
Quote from us (why we built it): ____
Quote from a customer (what changed for them): ____
How to get started: ____

## Customer FAQ
- What does it cost? ____
- Why is this better than what I use now? ____
- What does it not do? ____
- ____

## Internal FAQ (whole FAQ five pages or less)
- Who exactly is the customer, and how many are there? ____
- What has to be true for this to work? ____
- What are the biggest risks, and how do we retire each one? ____
- What does it depend on that we do not control? ____
- What will it cost to build and to run? ____

## Decision
Faster, easier or cheaper than today, and how: ____
Build / rewrite / shelve: ____
Requirements taken from the press release:
- [ ] ____
- [ ] ____
Case study · Southern New Hampshire University

Working adults asked about a degree and got a form letter and a mailed packet.

The problem
Southern New Hampshire University treated online students much like the high school graduates it had recruited the same way for decades. A prospective student who asked for information got a boilerplate reply within 24 hours and, a week or so later, the same mailed packet everyone received. The average online student was 30, juggling work and family, and needed a credential to improve their career prospects.
What they did
SNHU moved its small online team to separate offices in a former mill yard in Manchester and mapped every hurdle from first inquiry to first class for that job. It took over fetching transcripts from students' previous colleges, settled the financial aid conversation in a few days, and gave every new online student a personal adviser who stays in constant contact.
What happened
By the end of fiscal 2016, SNHU was closing in on $535 million in revenue, a 34% compound annual growth rate over five years, and its online college served more than 75,000 students.

Lesson for your teamSegment by the job and the situation, then fix the process the customer actually goes through.

Sources: Boston Globe: excerpt from Competing Against Luck GovTech: how jobs to be done applies to online education

How to run it

  1. Interview recent switchers

    Talk to people who recently started or stopped using a product like yours. Ask about the day they decided, and leave features out of it.

  2. Rebuild the moment of decision

    Ask when the need came up, what they were doing, what they tried first and what nearly stopped them. Write the timeline in their words.

  3. Write the job statement

    Use the form: when I am in this situation, I want to make this progress, so I can reach this outcome. Add how they want to feel and how they want others to see them.

  4. Map every hurdle

    List each step from first contact to the job being done, and circle every place where your process slows, confuses or ignores the customer.

  5. Rebuild around the job

    Change the offer, the process and the support so each circled hurdle is gone for people with this job, even if other customers keep the old path.

Done when

  • the team has at least eight switch interviews, each with a written timeline of the decision.
  • the job statement names a situation, the progress wanted and the outcome, and every interviewee fits it or is listed as an exception.
  • the hurdle map covers first contact through first value, with an owner for each circled hurdle.
Template Markdown, ready to paste into a doc
# Jobs to Be Done: ____ (product)

## Switch interviews
| # | Person | Switched from | Switched to | When they decided |
|---|---|---|---|---|
| 1 | ____ | ____ | ____ | ____ |
| 2 | ____ | ____ | ____ | ____ |
| 3 | ____ | ____ | ____ | ____ |

## Timeline of one decision (repeat per interview)
- First thought: when did the need come up, and what were they doing? ____
- First look: what did they notice or try first? ____
- Comparing: what else did they consider? ____
- The push: what made them act that day? ____
- What almost stopped them: ____

## Job statement
When I am ____ (situation),
I want to ____ (progress),
so I can ____ (outcome).
How I want to feel: ____
How I want others to see me: ____

## Hurdle map: first contact to first value
| Step | What the customer does | Wait or confusion | Hurdle (yes / no) | Owner | Fix |
|---|---|---|---|---|---|
| ____ | ____ | ____ | ____ | ____ | ____ |
| ____ | ____ | ____ | ____ | ____ | ____ |

## Exceptions
People who did not fit this job, and the job they had instead: ____
Case study · Intercom

A backlog full of good ideas, and no consistent way to compare any two.

The problem
Intercom's product team found the prioritization methods it looked at did not compare different ideas in a consistent way. The usual pulls kept showing up: pet ideas the team would use themselves, clever ideas over ones tied to goals, new ideas over ones they were confident in, and extra effort that was easy to discount.
What they did
The team built RICE, which multiplies reach per period, impact on a five-step scale and confidence as a percentage, then divides by effort in person-months to get total impact per time worked. Scores live in a shared spreadsheet that sorts the backlog automatically.
What happened
Intercom published the method and its spreadsheet for other teams to copy, with one standing rule: a dependency or a strategic need can justify building a lower score first.

Lesson for your teamPut every estimate in the open, so the debate is about the inputs instead of whose idea it is.

Sources: Intercom blog: RICE, simple prioritization for product managers

How to run it

  1. Fix the goal and the period

    Pick one goal and one time window, such as a quarter, so every estimate counts the same thing.

  2. Estimate reach in real numbers

    Count the people or events each idea touches in the period, from your own data, such as customers per quarter or signups per month.

  3. Score impact on a fixed scale

    Use 3 for massive, 2 for high, 1 for medium, 0.5 for low and 0.25 for minimal, judged against the goal you picked.

  4. Discount for confidence

    Use 100% when data backs the estimates, 80% when some of it does and 50% when it is mostly a hunch. Treat anything lower as a moonshot and scope it separately.

  5. Estimate effort in person-months

    Count the total work across product, design and engineering, in whole numbers, or 0.5 for anything well under a month.

  6. Rank, then decide

    Sort by the score and use the order to start the conversation. When you overrule it, write down why.

Done when

  • every idea on the list has all four inputs filled and a score from the same formula.
  • each reach figure names its data source.
  • every lower-scoring idea built ahead of a higher one has a written reason, such as a dependency.
Template Markdown, ready to paste into a doc
# RICE scoring: ____ (team)

Goal: ____
Period: ____ (for example, next quarter)
Reach unit: ____ (for example, customers per quarter)
Scored by: ____   Date: ____

## Scores
Score = Reach x Impact x Confidence / Effort
Impact: 3 massive, 2 high, 1 medium, 0.5 low, 0.25 minimal
Confidence: 100% high, 80% medium, 50% low

| Idea | Reach | Reach source | Impact | Confidence | Effort (person-months) | Score |
|---|---|---|---|---|---|---|
| ____ | ____ | ____ | ____ | ____% | ____ | ____ |
| ____ | ____ | ____ | ____ | ____% | ____ | ____ |
| ____ | ____ | ____ | ____ | ____% | ____ | ____ |
| ____ | ____ | ____ | ____ | ____% | ____ | ____ |

## Below 50% confidence (moonshots, scoped separately)
- ____
- ____

## Overrides
| Built ahead of a higher score | Reason (dependency, strategy, commitment) | Decided by |
|---|---|---|
| ____ | ____ | ____ |

## Next rescore
Date: ____
Trigger for an early rescore: ____
Case study · Basecamp

Customers kept asking for a calendar that could take six months to build.

The problem
Soon after Basecamp 3 launched, customers started asking the team to add a calendar. A proper calendar can easily take six months or more, and past versions of Basecamp had calendars that only about 10% of customers used.
What they did
Basecamp set an appetite of one six-week cycle and called a customer who had asked, to learn when she wanted a calendar rather than what it should look like. She needed to see free spaces on a shared schedule, so the team shaped a Dot Grid: a two-month, read-only grid with a dot for each event, and no dragging, no multi-day spans and no color coding.
What happened
The team shaped the Dot Grid as a six-week project and launched it at the end of that cycle, in place of a calendar that could have taken six months or more.

Lesson for your teamFix the time and let the scope move, and you will find the smaller version that does the job.

Sources: Shape Up, chapter 2: Principles of Shaping Shape Up, chapter 3: Set Boundaries

How to run it

  1. Set the appetite

    Before any design, decide how much time the problem is worth, such as two weeks or six, for a small team of a designer and one or two programmers.

  2. Find the baseline

    Ask a customer when they wanted the feature and what they were doing at that moment. Narrow the request to the specific need behind it.

  3. Shape it roughly

    Sketch the main elements and flow, name the rabbit holes, and list what is out of bounds. Leave the details to the team that builds it.

  4. Bet in the cool-down

    Choose which shaped projects to bet on in the break between cycles, then give each team the whole cycle without interruptions.

  5. Use the circuit breaker

    Work that is not done when the cycle ends gets no extension by default. Reshape it and bet again, or drop it.

Done when

  • every project in the cycle has a written appetite, a shaped pitch and a list of what is out of bounds.
  • the pitch names the specific customer need the request was narrowed to.
  • the cycle end date is fixed and unfinished work goes back to shaping instead of rolling over.
Template Markdown, ready to paste into a doc
# Pitch: ____

Appetite: ____ weeks, for ____ designer and ____ programmers
Cycle: ____ to ____ (no extension by default)

## Problem
The customer story behind the request (when did they want it, and what were they doing): ____
What exists today (baseline): ____
Narrowed need, in one sentence: ____

## Solution sketch
Main elements: ____
Main flow: ____
Sketch link: ____

## Rabbit holes
- ____ and how we avoid it: ____
- ____ and how we avoid it: ____

## Out of bounds
- ____
- ____
- ____

## Betting table (cool-down)
Bet on it: yes / no
Decided on: ____ by ____

## At cycle end
Shipped: yes / no
If not: reshape and bet again / drop
Notes: ____
Case study · Dropbox

Years of hard engineering ahead, for a product people struggled to grasp from a description.

The problem
Dropbox had to sync files across Windows, Mac, iPhone and Android, backed by an online service that needed high reliability. People also had a hard time understanding the product when it was explained, and Drew Houston faced the risk of waking up after years of development with a product nobody wanted.
What they did
Houston made a three-minute screencast, narrated by him, showing the product working as it was meant to, to test whether people would try it if the experience was better. He posted a version to Digg with about a dozen Easter eggs aimed at that audience.
What happened
The video drove hundreds of thousands of people to the site, and the beta waiting list grew overnight to 75,000 people, up from 5,000.

Lesson for your teamWhen building is expensive, give people something to react to before you build it.

Sources: TechCrunch: Dropbox case excerpted from The Lean Startup TechCrunch: how Dropbox got its first 10 million users

How to run it

  1. Write the leap-of-faith assumption

    State the belief that would sink the product if it were wrong, as a sentence you can test. For example: if we offer this experience, people will sign up to try it.

  2. Pick the smallest test

    Choose the cheapest thing that makes people act, such as a demo video, a landing page with a waitlist, a request form, or a manual service behind a simple front end.

  3. Set the bar before launch

    Write down the number that means yes, the number that means no, and the date you will read them.

  4. Show it where early users gather

    Put the test in front of the people most likely to want it first, in the places they already spend time and in the way they talk.

  5. Read the result and decide

    Compare the result with the bar. Keep going, change direction or stop, and log the decision with the evidence.

Done when

  • the assumption, the test, the bar and the read date are written down before the test goes live.
  • the test measures an action, such as a signup, a request or a payment, rather than an opinion.
  • the decision to keep going, change direction or stop is logged with the numbers behind it.
Template Markdown, ready to paste into a doc
# MVP test card: ____

Leap-of-faith assumption: If we ____, then ____ will ____.
Why this is the riskiest assumption: ____

## Test
Type: demo video / landing page with waitlist / request form / manual service / ____
Audience, and where we reach them: ____
Build budget: ____ days
Launch date: ____
Read date: ____

## Bar (set before launch)
Action we count: ____ (signup, request, payment)
Yes if: ____ or more by the read date
No if: ____ or fewer by the read date

## Result
| Measure | Bar | Actual |
|---|---|---|
| ____ | ____ | ____ |
| ____ | ____ | ____ |

What surprised us: ____

## Decision
Keep going / change direction / stop: ____
Evidence: ____
Next assumption to test: ____
Decided by: ____ on ____
Case study · Google (Gmail)

Seven-day active users counted anyone who opened Gmail once, and missed who relied on it.

The problem
Google teams commonly tracked seven-day active users, which counts everyone who visited at least once in the last week. The Gmail team wanted to understand how engaged its users were, and that count could not tell a daily user from one who visited once.
What they did
Working through the HEART categories and the goals, signals and metrics process, the team reasoned that engaged users check email as part of their daily routine. It chose a new engagement metric: the share of active users who visited on five or more days in the last week.
What happened
The researchers found that metric strongly predictive of longer-term retention, so it could serve as an early read on it. By the time of the paper, the authors had applied the framework to more than 20 Google products and projects.

Lesson for your teamChoose the behavior that shows the product is part of someone's routine, and measure its share instead of its count.

Sources: Google Research: Measuring the user experience on a large scale (CHI 2010) The paper as a PDF

How to run it

  1. Write the goals first

    For the product or feature, write what the user experience should achieve, using Happiness, Engagement, Adoption, Retention and Task success as prompts. Settle disagreements about goals at this step.

  2. Leave categories out on purpose

    Mark each category in or out, with a reason. Engagement may mean little for a tool people must use at work, for example.

  3. Find a signal for each goal

    Name the user action or attitude that would show the goal was met, and where it is logged or surveyed. Prefer signals that move only when the experience gets better or worse.

  4. Turn signals into rates

    Convert raw counts into ratios, percentages or averages per user, so a growing user base does not pass for a better experience.

  5. Check what the metric predicts

    Before a new metric goes on the main dashboard, test whether it tracks a longer-term outcome such as retention.

Done when

  • each metric on the dashboard links to a written goal and a named signal.
  • every HEART category is marked in or out with a reason.
  • no metric on the dashboard is a raw count that rises only because the user base grows.
  • at least one engagement or happiness metric has been checked against later retention.
Template Markdown, ready to paste into a doc
# HEART metrics: ____ (product or feature)

Owner: ____
Dashboard: ____
Review cadence: ____

## Goals, signals, metrics
| Category | In or out (why) | Goal | Signal | Metric (a rate or average) | Source |
|---|---|---|---|---|---|
| Happiness | ____ | ____ | ____ | ____ | survey |
| Engagement | ____ | ____ | ____ | ____ | logs |
| Adoption | ____ | ____ | ____ | ____ | logs |
| Retention | ____ | ____ | ____ | ____ | logs |
| Task success | ____ | ____ | ____ | ____ | logs or study |

## Rules
- Every metric is a ratio, a percentage or an average per user.
- Every signal is logged today, or has a ticket to log it: ____
- Traffic from bots and crawlers is filtered out: yes / no

## Validation
Metric checked against later retention: ____
Window: ____ weeks
Finding: ____

## Change log
| Date | Metric added, changed or removed | Why |
|---|---|---|
| ____ | ____ | ____ |
Case study · Duolingo

By mid-2018, daily users were growing in single digits after years of explosive growth.

The problem
By mid-2018, Duolingo's daily active users were growing at a single-digit rate year over year, which was troubling after the explosive growth the company had seen. Jorge Mazal, Head of Product since late 2017, needed to find which lever would restart growth.
What they did
The team built a growth model that sorts users into states such as new, current, reactivated and dormant, linked by retention rates, and ran a sensitivity analysis on years of daily history. Current user retention rate, the chance that someone who came in each of the past two weeks comes back this week, had 5 times the impact on DAU of the next-best input. Duolingo created a Retention Team with it as its North Star metric, and that team shipped leaderboards, streak improvements and better push notifications.
What happened
Mazal reports that current user retention rose 21% and DAU grew 4.5 times over four years, and that the leaderboards launch raised overall learning time by 17%.

Lesson for your teamModel which input moves your top number most, then give that input a team of its own.

Sources: Lenny's Newsletter: How Duolingo reignited user growth, by Jorge Mazal Duolingo blog: how the growth model sharpened product teams Amplitude: the North Star Metric and its three to five inputs

How to run it

  1. Name the North Star

    Pick one number that rises when customers get the value they came for, such as daily learners or orders delivered on time. It should lead revenue rather than be revenue.

  2. Break it into three to five inputs

    Write the North Star as a result of inputs a team can move, such as new users, current user retention and reactivation of lapsed users.

  3. Model the system

    Use your history to estimate how users move between states, then change each input by the same amount in the model and see which one moves the North Star most.

  4. Give the top input its own team

    Staff a team with that input as its goal, and have it run experiments that must move the input before they ship.

  5. Check that the link holds

    Confirm with A/B tests that moving the input moves the North Star. Rerun the model each quarter, since the strongest input can change.

Done when

  • the North Star and each input have an exact definition and a single owner.
  • a model ranks the inputs by their effect on the North Star, using your own history.
  • at least one experiment shows a change in an input followed by a change in the North Star.
Template Markdown, ready to paste into a doc
# North Star tree: ____ (product)

North Star: ____
Exact definition: ____
Why it reflects the value customers get: ____
Owner: ____

## Inputs (three to five)
| Input | Exact definition | Current value | North Star change from a 2% lift (model) | Team |
|---|---|---|---|---|
| ____ | ____ | ____ | ____ | ____ |
| ____ | ____ | ____ | ____ | ____ |
| ____ | ____ | ____ | ____ | ____ |

## Model
User states: new, current, reactivated, dormant, ____
History used: ____ months of daily data
Top input this quarter: ____
Its impact compared with the next-best input: ____ times

## Experiments on the top input
| Experiment | Input change | North Star change | Ship (yes / no) |
|---|---|---|---|
| ____ | ____ | ____ | ____ |
| ____ | ____ | ____ | ____ |

## Next model run
Date: ____
Run by: ____

Bring me the problem worth building.

A product idea, a team that needs an extra builder, or a business that wants its busywork automated. Tell me the problem and I will tell you honestly whether I can help.

LinkedIn Resume
↑ ↓ to moveEnter to openEsc to close