When product principles become performance

In this post

Over more than fifteen years of building products and leading teams, I have seen how easily an organization can become fluent in the language of product management without examining whether the reasoning behind its decisions is sound. The most important lesson I have learned is that product leadership depends on keeping the connection between our assumptions, our decisions, and the customer’s experience open to examination, especially when doing so could overturn an approved plan or expose a mistake of our own.

This is what concerns me when “Day 1” becomes a description of an organization’s identity rather than a demand on its behavior. I have watched teams speak confidently about invention while treating an inherited roadmap as the boundary of what could be considered, and invoke high standards without being able to explain the judgment they expected someone to improve. Under those conditions, the vocabulary of curiosity can help preserve the very assumptions it was meant to challenge.

The risk is that we become better at executing plans than at understanding whether they will improve the customer’s experience. That can waste considerable effort, but it also limits what a team is capable of imagining. If we want people to build something meaningfully better, we have to give them both the intellectual tools to see the opportunity and an organization in which seeing it can change the work.

The reasoning beneath the roadmap

By a product mental model, I mean the explanation we are relying on when we expect a particular intervention to change a user’s behavior or circumstances, including the assumptions about what they need, what prevents them from getting it, and why the proposed change would produce an outcome worth pursuing. That explanation determines what we notice and what we build, even when no one has made it explicit enough to question.

Consider a hypothetical planning product whose recommendations customers rarely use. If the team assumes people have forgotten to return, reminders seem like a sensible investment; if it assumes they distrust the recommendations, improving the model seems more promising. But neither addresses a recommendation that arrives after the customer has committed to a decision, covers too little of the work to replace an existing process, or asks someone to take an action they lack permission to perform.

The first useful question is therefore what happens between receiving the recommendation and acting on it. An interview may reveal when plans become committed, but observing the next planning cycle can show whether the information is available before that point and whether someone can use it; another visit to the dashboard establishes neither. A customer who reads a recommendation, agrees with it, and still cannot act presents a different problem from someone who never sees it, and that distinction should change where we invest.

This is the substance of first-principles thinking. We separate what we know about the user’s task from what we have inherited about how the product should work, then examine the assumptions connecting the two. Repeatedly asking why helps when each answer leads to evidence or a question that could distinguish competing explanations; without that discipline, a root-cause exercise can produce an increasingly confident story about a cause we have never established.

Systems thinking extends the inquiry beyond the team’s component and beyond the conditions of its first release. If acting on a recommendation requires another team’s approval, a better model may increase the volume of requests without increasing the capacity to review them, so that the intervention which helped the first few customers becomes less useful as queues grow. The approval process may itself need to change. A product decision needs an account of what becomes scarce when it succeeds, rather than an assumption that benefits observed at small scale will simply multiply.

There is also a choice about what we are optimizing. A generation target directs attention toward producing more recommendations, whereas helping a customer complete a planning decision may require connecting several recommendations, resolving a missing input, or changing who can act. Product sense includes recognizing when the organization’s unit of work bears little resemblance to the customer’s unit of value.

A team can execute efficiently within the wrong model; the more concerning case is an organization whose measures and reviews prevent that model from being corrected.

How a team loses the ability to correct itself

Imagine a service in which 20 of 100 eligible customers complete a task, giving it a completion rate of 20%. Restrict access to the 50 customers most likely to succeed, with the same 20 completing, and the rate becomes 40% even though no additional customer has succeeded. Narrowing the audience might be a defensible commercial choice, but that is a different decision from improving the service for the original audience, and a goal that obscures the distinction invites the team to claim one while doing the other.

Every goal selects some consequences of the work for attention and leaves others in the background. A team rewarded for reducing support contacts could make help harder to reach, while one rewarded for adoption could make participation mandatory; in each case, the target can be achieved by behavior that defeats its purpose. Once that target is attached to recognition or career progression, leaders are responsible for examining what it encourages and what evidence of harm it might leave outside the review.

The same issue appears in how we evaluate people. When feedback such as “be more strategic” cannot be connected to a decision, an overlooked alternative, or an inadequately examined assumption, the recipient is left trying to predict what will satisfy the reviewer rather than learning how to improve the work. If expectations then shift after the outcome, it becomes difficult to establish whether the judgment was poor, the conditions changed, or the standard was never clear. We do not need to infer anyone’s motives to recognize that problem.

Sue Bolton’s account of product teams resisting learning describes one recognizable version of this problem.1 My concern extends to the arrangements around the team. A leader can ask for challenge while rewarding agreement, demand ownership while withholding decision rights, or praise ambition while funding only work that preserves existing commitments. Under those conditions, telling product managers to be more curious leaves the source of the contradiction intact.

In a field study of 51 manufacturing teams, Amy Edmondson found that psychological safety was associated with learning behavior, while confidence in the team’s capability was not associated with learning once psychological safety was taken into account.2 The study does not provide a causal estimate for product teams, but it gives us a reason to distinguish confidence that we can do the work from confidence that we can expose a gap without being punished for it. An impressive, self-assured team may still find the second much harder than the first.

We can also miss useful reasoning by confusing its presentation with its quality. Research on curiosity across cultures found variation in human-authored questions that language models often flattened into more familiar patterns.3 Although this concerns online questions and models rather than workplace performance, I take it as a caution against assuming that the form of inquiry most familiar to us is the only one worth hearing. A direct disagreement, a cautious question, and a carefully developed written objection may each contain an important challenge. We should examine the reasoning before judging the person, without assuming that nationality determines how an individual will think or speak.

None of this requires accepting every objection or lowering the bar. It requires a bar that distinguishes weak reasoning from unfamiliar expression, and a response to dissent that makes clear whether the concern was investigated, rejected for a reason, or deferred because another uncertainty mattered more. Otherwise, we risk filtering out the information that could correct us and treating the resulting agreement as evidence that we were right.

Some of my best experiences in a large organization were with two colleagues who became close friends. Those friendships are part of why I care about this problem; I want people to have room for demanding, generous work together without having to choose between intellectual honesty and belonging.

What understanding should make possible

Responding to these failures by requiring more questions in every review would leave us with another ritual unless the answers could change a material decision.

The Quriosity research distinguishes naturally occurring inquiries from questions constructed to test an answer already known, and examines the causal explanations people seek.4 I find that distinction useful when thinking about product reviews. An inquiry has value because we do not yet know what the answer should be. If every acceptable answer leads back to the proposed roadmap, we are testing the presenter’s ability to defend it rather than investigating whether it is right.

Return to the planning product. If the important obstacle is that customers cannot turn advice into action, then better understanding should affect the scope, the ownership, and perhaps the architecture of what we build. The product may need to connect data and permissions across a workflow, with recommendation quality becoming one capability inside a larger service. That is a different investment from another notification, and it may require a different team arrangement to succeed.

This is where product vision becomes useful. A vision should describe the improved experience clearly enough that we can identify what is missing, which capabilities already exist, and what would need to change for the whole system to work. Starting from today’s architecture can make necessary change appear impossible before it has been evaluated. Ignoring that architecture can produce a future that has no credible route from the present. Leadership has to hold both views long enough to make the comparison honestly.

A rewrite might remove a structural constraint that years of patching would preserve, or consume the team while delaying the very capability customers need; the comparison has to include migration risk, operating cost, dependencies, and the opportunities forgone while the work continues. The size of the initiative is a poor substitute for that analysis. Once the main constraint has shifted, further investment in the component we already know how to improve may produce very little additional value. The next important opportunity may sit elsewhere in the workflow.

Nor can we demand empirical proof of a future product before allowing anyone to explore it. A serious vision will contain assumptions that existing data cannot settle. The response should be to find a credible first test, choose how much to commit before the next uncertainty is resolved, and recognize what the test would leave unknown. Curiosity should improve the quality and timing of action. It should not become an indefinite exemption from making a decision.

What leaders should do

The practical changes begin with the things leaders control, because asking for better product thinking while leaving the same goals, funding decisions, review habits, and development standards untouched would preserve many of the reasons it was difficult in the first place.

First, we should make the reasoning behind an important investment inspectable before asking the team to defend its delivery plan. The explanation should connect a particular user’s problem to the proposed change, identify the assumption most likely to invalidate the investment, and describe what evidence would justify continuing, altering, or stopping it. For a small reversible test, the team may need only a brief explanation and a few days, while a difficult-to-reverse commitment warrants more scrutiny of the assumptions on which it depends. Neither should accumulate research simply because no one is prepared to decide. We should ask whether the next investigation is likely to improve the choice enough to justify waiting for it.

This requires senior leaders to make consequential choices, too. When evidence undermines a plan we sponsored, we should be prepared to change the commitment, move resources, and explain the revision ourselves. If every discovery still has to fit the original deadline and scope, the team has been given very little room to learn, however often we encourage it. Preserving what was reasonably believed at the time also lets us distinguish a justified change of course from an attempt to rewrite the history of a decision.

Second, we should design goals around a shared outcome and a defensible account of contribution. A team target needs a visible connection to the result the business wants, especially where improving one team’s measure could worsen another’s. If acquisition and retention goals conflict, leaders should establish how the trade-off will be resolved before asking teams to negotiate it through competing dashboards. Some trade-offs cannot be reduced to one number, but they can still have explicit limits and a person with authority to decide.

Within that shared objective, we should distinguish the product’s overall health from the incremental effect of an intervention. Total sign-ins may tell us that something important has changed; they do not, by themselves, establish what a particular team accomplished. Experiments, comparable cohorts, or a carefully qualified analysis can help estimate the difference its work made. When attribution is uncertain, we should say so, while continuing to hold the team responsible for investigating the result and making a sound next decision. Where success depends on another team, leadership needs to secure a joint commitment, narrow the promise, or choose a different plan. Telling the product manager to show more ownership does not create capacity or authority in someone else’s organization.

Third, we should develop judgment through the work itself. Instead of asking someone to sound more strategic, examine a real decision with them, show which assumption was unsupported or which alternative deserved consideration, and ask them to revise the analysis. A fair development plan gives the person a specific standard, the support needed to reach it, and a chance to respond to feedback, while protecting the team if important decisions continue to be made without adequate reasoning or follow-through. A useful test of development is whether the person can handle the next unfamiliar decision with less intervention from us, and explain what changed their view. Curiosity does not excuse repeated failure to act on what has been learned.

The same standard should apply to us. Inviting questions in writing, leaving room for people who need time to formulate an objection, and explaining why a challenge did or did not change our view make the reasoning available beyond a single meeting. Giving someone authority over a decision also requires access to the customer evidence and specialist knowledge needed to make it. We cannot delegate judgment while retaining all the conditions that make judgment possible.

Finally, we should make useful learning count when we allocate recognition and future responsibility. That includes an investigation that prevents a poor investment, a correction that makes another team’s work succeed, and a difficult decision to stop an idea that has not earned further funding. We should look at the quality of the reasoning and the response to evidence alongside the outcome, because a favorable result does not validate every decision that preceded it. Equally, a well-presented account of learning should not excuse repeatedly avoidable mistakes.

This asks more of us than promoting the right principles, because we have to expose our own reasoning, resolve conflicts we would otherwise pass down, and sometimes surrender a plan that has become part of our reputation. The benefit is a team that can take on larger problems with a credible way to learn as it goes.

We should make it possible for a product manager to bring us evidence that changes our understanding and leave with the authority, resources, and revised commitments needed to act on it. When that happens consistently, curiosity has become part of how the organization works.


Notes

Footnotes

  1. Sue Bolton, “When a Product Manager Thinks They Know It All, They Stop Learning” (9 February 2026). This is a practitioner’s account. The discussion here of organizational incentives and leaders’ responsibilities is my argument, rather than a finding demonstrated by that article. Back to reference 1

  2. Amy Edmondson, “Psychological Safety and Learning Behavior in Work Teams”, Administrative Science Quarterly 44(2), 350–383 (1999). This additional organizational-research reference concerns a multimethod field study of 51 teams in one manufacturing company. Its associations should not be presented as a causal estimate for software or product organizations. Back to reference 2

  3. Angana Borah, Zhijing Jin, and Rada Mihalcea, The Curious Case of Curiosity across Human Cultures and LLMs, version 2 (20 October 2025). The study compares country-associated patterns in online questions and model-generated questions, covering 18 countries and 16 topics. Platform selection, translation, and cultural coverage limit interpretation. It does not establish an individual’s curiosity, competence, or workplace behavior from their nationality; the leadership application is my inference. Back to reference 3

  4. Roberto Ceraolo and colleagues, Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries, version 4 (9 November 2025). The research analyzes 13,500 naturally occurring questions across search, human question-and-answer, and human–LLM interactions. Its online, English-language data and classification analyses do not establish the effect of curiosity on product-team performance. The distinction is used here to frame a question about reviews, not as proof of an organizational intervention. Back to reference 4

Back to top

THE INDEX

Find a thread.

Loading the index…