← Writing
Essay

A Working Guide to Product Discovery in the AI Era

Building got cheap. Judgment didn't. AI removed the governor, not the constraint.

July 16, 2026

Building got cheap. Judgment didn’t. AI removed the governor, not the constraint.

By: Michael Albers and Felix von Kunhardt.

Between us we have spent a couple of decades on opposite ends of the same problem. Michael spent years building and running consumer products at Yahoo at a scale most teams never touch, Yahoo Mail alone north of 250 million people. Building at that scale teaches a second discipline nobody enjoys: knowing when a product that still works has run its course, the way AIM, Answers, and Groups eventually did. Felix spent years on the other side of that glass, building the behavioral-analytics products (Decibel, then Medallia) that read what users did over what they said — and when the two disagreed, the behavior was right. One of us built at scale and learned what to end. The other learned to read the signal that tells you when to. AI just turned both of those into the whole job.

Building got cheap. Judgment didn’t. The constraint was never the building — it was always real customer value, and finding it is still the hard part. That makes judgment the scarce good in both directions: the bet on what’s worth building, and the call to kill what isn’t.

That muscle matters now, and almost nobody is training it, because for twenty years the system trained it for us by accident.

Here is what changed. The old product-design-engineering process had friction baked in. A PM gets excited about an idea. The designer iterates. The engineer pushes back, “we can’t do this, it takes too long, I have ten other things.” That pushback was the governor on what got built. Nobody owned it. It was just a property of working with humans on hard problems, and it quietly filtered out most of the bad ideas before they ever shipped. AI removed that friction. It did not remove the constraint underneath, which is still the same scarce thing it always was: real customer value. So the system can now produce far more than the constraint justifies. The kill function, the thing the friction used to do for free, has to become explicit team practice. Otherwise the team builds more, faster, worse.

When anyone can build anything, the team that wins is the one that knows what’s worth building — which makes discovery, and the product judgment driving it, more decisive than ever, not less. The role of the PM doesn’t shrink in the AI era. It concentrates.

That is the spine, and it has two sides: the discipline to kill what doesn’t earn its place, and the discovery to find what does. Everything below is one attempt to make both operational on a real team, weekly — without pretending a tool can do it for you.

What discovery is, and the trap

Discovery is everything between “we have an outcome to chase” and “we know, with reasonable evidence, what to build that will move it.” Interviews, synthesis, opportunity mapping, ideation, prototyping, assumption tests, the kill-or-iterate-or-proceed call. It is not the build itself, even when AI makes the build look like a finished product.

Marty Cagan, borrowing Jeff Patton, draws the same line in cleaner words: discovery is build to learn, delivery is build to earn¹. The entire loop here is build-to-learn work. Prototypes exist to answer questions, not to ship. The most damaging confusion of the AI era is treating a build-to-learn prototype as if it were build-to-earn, because it now looks like one. Keeping those two verbs apart is the cleanest one-line version of the scope this draws.

We lean on Teresa Torres for cadence and structure² (continuous, Opportunity-Solution-Tree “OST”-centric, a weekly trio rhythm) and on Cagan’s four risks for classifying assumptions. Both anchor it. Neither is the whole picture alone. What we built on top of them is the framework this piece lays out, and it is ours, Felix’s and Michael’s both: the governor framing, the kill-list-as-eval-set, the judgment ritual, shared context as the layer that compounds, and where the work actually sits when the org gets bigger or smaller.

The Foundation: three things shared, not assumed

Before the loop runs at all, three things have to be genuinely shared. Not assumed. Most teams skip this because the work feels like talking about talking, and teams want to get to artifacts.

  • The outcome. The metric at the top of the opportunity solution tree, grounded in company strategy. Strip that grounding and a team can run the whole loop with discipline and still optimize the wrong thing, faster.

  • The live OST. Not a workshop output gathering dust. The running state of the team’s shared understanding, updated every cycle.

  • Shared context. Style sheets and component assets, technical constraints, captured judgments from prior cycles (why we killed X, why Y worked), and the definition of good for every artifact the loop produces. This is the layer that compounds. Every cycle deposits learnings into it. Every future cycle benefits.

The third pillar is where almost every team falls down. If PM-A’s snapshot looks different from designer-B’s looks different from engineer-C’s, the team is running three workflows in parallel, not one. Same story when an AI prototype looks nothing like the product because nobody fed it the style guide. Same story when engineers complain about spaghetti code from AI, because nobody fed it the constraint context.

Now the load-bearing idea. The definition of good is the kill criteria is the eval set. One artifact, not three. Written upstream, before any prototype exists. It does three jobs: it defines good, it runs the kill later in the loop, and it teaches judgment, because a newer PM absorbs the team’s taste by helping write it. Define what good looks like first, and killing stops being a debate you relitigate every time someone gets attached to a demo. It becomes a test you run.

The Loop: four phases, weekly

Four phases that cycle on a weekly rhythm. The loop closes when the kill decision updates the OST, and the next pass starts from a smarter tree.

Phase 1, signal gathering. Opportunity interviews, story-based, focused on past behavior, no prototype. At least one a week per trio member. Each interview produces a snapshot that meets the shared definition of good. AI takes over the coding, tagging, and verbatim extraction completely. The judgment stays human. Phase 1 stays prototype-free on purpose, because the AI-era reflex to “always build something first” biases the conversation toward a solution you already imagined. When Felix was at eBay the instinct was to slim down the long listing form everyone complained about — but opportunity interviews showed many sellers liked it, while a whole other group would only list at all if it shrank to the few fields that matter. That insight — surfaced in interviews, not from testing a prototype — is what produced the Easylister at three pages instead of eight, resulting in a significant uplift in new listings.

Phase 2, pattern plus the judgment ritual. The team reads the week’s snapshots and converges on real versus noise versus ambiguous. AI is the bookends here, not the middle. It preps the synthesis brief before the room and captures the notes after. It does not run the discussion. The moment PMs hand the cognitive work to a tool, “AI did the synthesis, here’s the output, let’s move on,” they stop building the judgment the loop exists to build in them. The ritual stays slow on purpose. The slowness is the work. As an Anthropic leader put it, the one thing AI can’t do is get six people in a room to align.

Phase 3, choice and test framing. Pick the target opportunity. Bring solution ideas. Enumerate the assumptions and classify each by Cagan’s four risks: value, usability, feasibility, business viability. Knowing which risk an assumption falls under is half the work of deciding what to test next. Then set kill criteria upfront, as a team, documented, before any prototype exists. Otherwise they become political theatre the moment something working appears.

This is also where the discipline scales. Step one is pure discipline, no tooling: agree what good looks like and what would kill, written upstream. That pays off on its own. Step two arrives only when the human gate becomes the bottleneck. Michael watched an ad team go from a manager approving a handful of tests a week to drowning under 500 a week. At that volume you layer an eval agent on the same criteria to scale the gate. The criteria do not change. What scales is the application. You earn step two by doing step one.

Phase 4, build and decide. Build the test for a specific assumption. Run solution-validation interviews with the prototype as stimulus, targeting the assumption, not the customer’s whole world. Then the decision ritual: we agreed if X happened we would kill, X happened, we kill. This is the politically densest phase and the one where AI helps least, because something working now exists and the cost of killing it just went up.

The stopping signal is behavior, not opinion. Whether the prototype changed what the user did, not what they said they liked. Most teams instrument for opinions because opinions are easy to collect, and opinions are exactly what lets a working prototype survive on enthusiasm. The mature form of this is instrumented behavior read at volume, session and clickstream data interpreted by AI, not a survey. At Yahoo the click logs across the surface were the reaction data Michael trusted. The stated preference rarely was. He watched the product get better week over week off what people actually did inside it, a kind of mirrored-glass loop where the real reaction was always in the behavior, never in the comment box.

At Decibel, and then inside Medallia, the product Felix led was generating insights through session replays and experience analytics at enterprise volume. He watched teams argue for a redesign off a handful of vocal complaints while the behavioral data put the real friction three screens away. The heatmap doesn’t care who shouted loudest. Spend enough time on it and the focus group stops being an answer and becomes, at best, a hypothesis generator.

One move sits just off the loop: synthetic users. Before you spend a real interview slot, an AI stand-in can pressure-test a concept cheaply. The discipline is the whole game. A synthetic “no” is signal, a synthetic “yes” is an echo, because the models people-please toward yes. So they screen out, they never sign off. The green light always belongs to real users. It’s the highest-novelty, highest-risk move here, and it earns its own piece, which we’re writing.

Who runs it: the middle is a seat, not a headcount

Someone has to carry the downstream context, engineering, systems, constraints, compliance, and run validated prototypes from choice through to production. That function never disappears. The only question is whether it gets its own headcount.

In a larger org it does. The Feature PM, reframed as prototype orchestrator and productionizer, owns that downstream context the upstream Empowered PM often lacks. Cagan says that role is gone. Michael disagrees, and reframes it around downstream-context ownership rather than ticket-shuffling. In a scale-up of 50 to 150, the org is not big enough for a separate seat, so the function gets absorbed into the trio, usually by the tech lead, sometimes by the PM wearing the productionizer hat through the build phase. Same architecture. Two scales. The roles AI actually erases, the project PM running standups, the automation PM managing tickets, are the ones that never owned the orchestrator function in the first place. The headcount disappears. The function does not.

Where it collapses, and where it breaks down

Five places this breaks in practice. Run them as a field-test checklist.

  • The “what does good look like” conversation gets skipped because it feels like talking about talking. Then the kill function becomes pure authority, because there are no shared criteria to point to.

  • The OST stops being live and becomes one person’s property. The synthesis ritual goes performative.

  • Individual AI workflows fork. Prompts, context files, and agents drift apart, and the shared standard quietly stops being shared.

  • The judgment ritual gets optimized into an AI synthesis dump. The artifact looks fine. The judgment underneath is gone. This is the AI-era failure mode and the easiest one to slip into.

  • Engineers build upstream of discovery. The cost barrier that used to keep them downstream is gone, so they fill the gap with whatever is interesting. A solution looking for a problem.

In one B-stage turnaround Felix helped with, the product had been jerry-rigged for years, because whoever shouted loudest — usually sales, with a client commitment — set the build. There were no shared criteria for what was worth building, so every “no” turned into a fight with the founder instead of a test they ran. Installing a regular “what does good look like” conversation was the only thing that turned the kill decision from authority into evidence.

Now the part most frameworks skip. Neither of us has run this integrated loop end to end. We each have run pieces, the kills, the judgment rituals, the instrumented-behavior reads, the gate-scaling, but not the whole thing as one machine on one team. Polishing this document makes it better. It does not make it true. Another iteration is the comfortable move and the wrong one. The next move is a real team on a real slice, every claim marked practiced or designed as we go, and the designed ones earning their tag or getting cut.

So the question is not whether the framework reads well. It reads fine. The question is which team runs it first, and which claim falls down when they do.

Sources: 1. Marty Cagan, “Build to Learn vs Build to Earn,” Silicon Valley Product Group, April 16, 2026 — where Cagan credits Jeff Patton (User Story Mapping) with coining the phrase. https://www.svpg.com/build-to-learn-vs-build-to-earn/ ; 2. Teresa Torres, Continuous Discovery Habits: Discover Products that Create Customer Value and Business Value (Product Talk LLC, 2021) — the source for the continuous cadence, the opportunity solution tree, and the weekly product-trio rhythm. See also producttalk.org/opportunity-solution-trees/.

Originally published on Medium ↗

How I work → Bring me the actual situation →

If this hit close to home, bring me the actual situation.

Discuss your situation