The short version. Many product and business decisions begin with a sense of certainty. Once an idea feels right, it is natural to start looking for evidence that supports it. This is not a sign of a poorly managed team. It is simply how the human mind tends to work. Decades of research in behavioral psychology, decision science, and large-scale online experimentation point to the same challenging reality: even smart, experienced people regularly misjudge which ideas will succeed. The answer is not to eliminate confidence, conviction, or intuition. It is to balance them with a small set of better questions, asked before significant time and money have already been committed.
What the numbers reveal
It helps to begin with the data, because it shows just how difficult it can be to predict which ideas will succeed.
Ron Kohavi, Thomas Crook, and Roger Longbotham studied the outcomes of ideas tested through Microsoft’s internal experimentation platform. These included features, redesigns, and other changes that teams had already developed because they believed those changes would help. Only around one-third of the tested ideas improved the metric they were designed to influence. The remaining ideas either produced no meaningful change or had an unintended negative effect.
Similar patterns have appeared elsewhere. Jim Manzi reported that only around 10 percent of controlled experiments at Google resulted in a business change. Avinash Kaushik suggested that teams misunderstand what customers want roughly 80 percent of the time, while Netflix, according to Mike Moran’s account, considered around 90 percent of the ideas it tested to be unsuccessful.
These figures do not suggest that the people behind the ideas lacked expertise. In fact, they come from some of the most experienced and data-informed organizations in the world. That is what makes the findings so valuable. They remind us that even strong instincts are not guarantees. Confidence can help a team move forward, but it becomes far more useful when paired with a process that allows assumptions to be tested and refined.
Experience shapes good ideas. Evidence helps us understand which ones work.
Why confidence does not always predict accuracy
If experienced teams can misjudge ideas so frequently, the next question is why confidence does not always correspond with accuracy. Research in behavioral psychology offers two useful explanations: how we evaluate our own certainty and how we estimate what a project will require.
Baruch Fischhoff, Paul Slovic, and Sarah Lichtenstein’s classic studies on calibration asked people to rate how confident they were in their answers to factual questions, then compared that confidence with the accuracy of their responses. Their findings, which have been replicated many times, showed that even when people described themselves as completely certain, they were still wrong a meaningful proportion of the time. The gap between perceived certainty and actual accuracy often became larger as confidence increased.
This does not mean that confidence has no value. It means that feeling certain and being correct are two different things. When a team feels strongly about an idea, that confidence can create momentum and alignment. It can also make testing feel less necessary, precisely at the moment when testing may be most useful.
Forecasting introduces another challenge. Roger Buehler, Dale Griffin, and Michael Ross’s research on the “planning fallacy” found that people regularly underestimate how long their own tasks will take, even when they have direct experience of similar tasks taking longer than expected.
Dan Lovallo and Daniel Kahneman extended this idea into business decision-making. Writing in Harvard Business Review, they argued that overoptimism, reliance on a single favorable scenario, and limited consideration of competitors’ possible responses can lead managers to evaluate major investments too positively. These investments might include mergers, market entries, or large capital projects. Their work suggests that disappointing outcomes are not always the result of unpredictable bad luck. They can also emerge from patterns of optimism built into the forecasting process itself.
When these findings are considered together, the pattern becomes clearer. Teams make predictions through a process that naturally leans toward optimism, then risk treating the confidence created by that process as though it were evidence. Recognizing this tendency allows teams to keep the positive energy of an idea while creating space to examine it more carefully.
The space between how certain we feel and what the evidence can support.
Why it can be difficult to step away from an idea
The initial decision is only one part of the cost. Another challenge appears after a project produces disappointing results and the team must decide whether to continue, adapt, or stop. At that point, walking away can feel much harder than it did before the work began.
Barry Staw’s now-classic experiment, “Knee-Deep in the Big Muddy,” presented participants with a business scenario in which an initial investment produced poor results. Participants were then asked how much additional money they would commit to it. Those who had personally made the original decision tended to allocate more resources to the struggling investment than those who inherited the same situation without making the initial choice.
This research describes what is known as escalation of commitment. Once people have invested time, money, effort, or professional credibility in a decision, the original investment can begin to influence what they choose to do next. Instead of considering only the project’s current prospects, they may also feel a need to protect or justify the earlier decision.
This is how the cost of confident guessing can grow over time. An initial assumption leads to an investment. When the expected results do not appear, the team may respond by investing even more, partly because so much has already been committed. This does not happen because people are careless or unwilling to learn. It happens because separating a project’s future potential from the emotional and practical weight of its past can be genuinely difficult.
Why warning signs are easy to overlook
There is another mechanism that helps explain why assumptions sometimes remain unchallenged for longer than expected. People do not always interpret new information neutrally, particularly when they already feel invested in a particular outcome.
Raymond Nickerson’s widely cited review describes confirmation bias as the tendency to seek or interpret evidence in ways that favor existing beliefs, expectations, or hypotheses. In a product or business context, a team that believes strongly in an idea may naturally pay more attention to encouraging metrics, find reasonable explanations for disappointing ones, or ask customers questions that make supportive answers more likely.
This is not necessarily a sign of dishonesty or deliberate manipulation. Confirmation bias is a normal feature of human thinking. Our existing beliefs help us make sense of new information, but they can also shape what we notice and how we interpret it. As a result, a strong belief in an idea can make the signals that challenge it more difficult to recognize.
Together, these mechanisms can create a reinforcing pattern. Overconfidence and the planning fallacy can produce an optimistic prediction. Confirmation bias can make warning signs less visible. Escalation of commitment can then make further investment feel more reasonable than changing direction. A team can be capable, sincere, and hardworking at every stage and still find itself caught in this pattern.
How optimism, selective attention, and continued investment can reinforce one another.
The answer is not less confidence. It is better questions.
If confidence can sometimes take the place of evidence, the goal is not to remove confidence from the process. Teams need belief, creativity, and momentum to explore new possibilities. The goal is to support that confidence with specific questions that can reveal uncertainty before substantial resources are committed.
Ask what could cause the idea to fail before focusing only on what would make it succeed. Gary Klein’s premortem technique, described in Harvard Business Review, asks a team to imagine that a project has already failed and then work backward to identify the possible reasons. Beginning with failure as the premise gives people permission to voice concerns that might otherwise feel overly negative or unsupportive.
A premortem does not necessarily require new data or an extensive research process. Its value comes from changing the framing of the conversation. Instead of asking team members to challenge an idea everyone is expected to support, it makes the search for risks a shared and constructive task.
Ask what the smallest meaningful test of the idea would look like before building the complete solution. This principle sits at the heart of the build-measure-learn loop popularized by Eric Ries in The Lean Startup. It also appears in Alberto Savoia’s concept of “pretotyping,” which focuses on learning whether people want an idea before investing heavily in proving that it can be built.
The order matters. When a team builds the complete solution before testing its core assumption, much of the investment the test was intended to protect has already been made. A smaller and less polished version can help the team learn whether the underlying idea is promising while the cost of changing direction is still manageable.
Ask questions that create room for an idea to be challenged. Rob Fitzpatrick’s framework in The Mom Test explains why questions such as “Would you use this?” often produce encouraging but unreliable answers. People can find it difficult to predict their future behavior, and they may also want to be supportive of the person asking.
Questions about specific past behavior tend to provide more useful information. What did the person do the last time they experienced this problem? What solution are they using now? Have they paid for an alternative? What have they already tried and stopped using? These questions reduce the pressure to make a prediction and give the team evidence grounded in real experience.
Run a test that can provide an honest answer, not only a reassuring one. In Trustworthy Online Controlled Experiments, Kohavi, Tang, and Xu argue that controlled experiments help teams distinguish between ideas that genuinely improve outcomes and those that only appear promising during internal discussions or demonstrations.
Where controlled testing is possible, the idea can be evaluated with a real audience, a comparison group, and a meaningful outcome metric. The purpose is not to prove that the original plan was wrong. It is to give the idea a fair opportunity to demonstrate its value on a limited scale before it reaches every customer.
A practical checklist for learning before making a larger commitment.
What this approach looks like in practice
The difference between confident guessing and a tested question is not merely theoretical. It can lead to a visible and measurable difference in how quickly a team learns and how much that learning costs.
Consider the base rate from Kohavi and his colleagues. On Microsoft’s experimentation platform, roughly two-thirds of the proposed ideas did not produce the outcome their teams expected. The organizations that benefited from this process were not those that somehow avoided incorrect predictions. Unsuccessful ideas are a natural part of experimentation. The advantage came from having a process that could identify them before they were introduced broadly.
A controlled experiment using a small portion of traffic can turn a costly and slow-to-detect problem into a smaller and faster learning opportunity. Without testing, a redesign might be released to every user and quietly reduce revenue for months before the effect becomes clear. With testing, the same issue may become visible within days, among a limited part of the audience, through a measurable result rather than an impression.
The same logic can be applied beyond software. Lovallo and Kahneman’s discussion of mergers, capital projects, and market entries describes a similar pattern on a much larger scale. An optimistic forecast may be created sincerely and in good faith, yet still move forward without a question capable of challenging its assumptions.
A premortem or reference-class forecast can introduce that challenge. Reference-class forecasting compares a new project with the actual outcomes of similar past projects instead of relying only on how promising the current plan feels. Neither technique requires more talented or less optimistic people. They simply introduce another perspective before resources are committed.
A practical framework for turning confidence into learning
- Before approval, run a premortem. Imagine that the idea has already failed and invite the team to identify every plausible reason. This creates room for concerns that may not emerge from a general discussion about whether the idea will work.
- Before building, define the smallest meaningful test. Ask what the simplest version of the idea would be that could provide an honest signal about whether its central assumption is valid.
- Before speaking with customers, prepare questions that can challenge your assumptions. Focus on specific past behavior rather than hypothetical future intentions. Ask what people actually did the last time they experienced the problem, not only what they believe they might do next time.
- Before a broad release, run a controlled test where possible. Allow a meaningful outcome metric from a real but limited audience to inform the final decision alongside the team’s experience and judgment.
- Before investing more resources in a struggling project, separate its history from its future. Ask what you would recommend to someone inheriting the decision today, without the personal weight of having made the original choice.
- Track what happens to your own team’s ideas. Understanding how often ideas succeed, fail, or produce no meaningful change can help the team calibrate its confidence and improve the way future decisions are made.
The takeaway
Confidence can feel like evidence because, from the inside, certainty is persuasive. Research suggests, however, that the relationship between confidence and accuracy is not always as strong as it seems.
This does not mean teams should wait until they feel completely certain before acting. Complete certainty is rarely available, especially when exploring something new. Instead, teams can move forward with a small set of thoughtful questions that make their assumptions visible and give the evidence a genuine opportunity to change the plan.
Unsuccessful ideas are not the real cost of experimentation. Every organization discussed here has tested ideas that did not work as expected. The greater cost is the time between committing to an idea and discovering what could have been learned earlier. Better questions shorten that distance. They help teams protect their resources, adapt sooner, and make decisions with confidence that has been informed by evidence.
References
The sources below provide the research and frameworks discussed throughout the article. Where a study or book is behind a paywall, the DOI or standard citation links to its official record.
- Kohavi, R., Crook, T., & Longbotham, R. (2009). Online experimentation at Microsoft. Third Workshop on Data Mining Case Studies and Practice Prize.
- Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press.
- Kohavi, R., Longbotham, R. (2010). Online controlled experiments and A/B testing. Encyclopedia of Machine Learning and Data Mining.
- Fischhoff, B., Slovic, P., & Lichtenstein, S. (1977). Knowing with certainty: The appropriateness of extreme confidence. Journal of Experimental Psychology: Human Perception and Performance, 3(4), 552-564. doi.org/10.1037/0096-1523.3.4.552
- Buehler, R., Griffin, D., & Ross, M. (1994). Exploring the “planning fallacy.” Journal of Personality and Social Psychology, 67(3), 366-381. doi.org/10.1037/0022-3514.67.3.366
- Lovallo, D., & Kahneman, D. (2003). Delusions of success: How optimism undermines executives’ decisions. Harvard Business Review, 81(7), 56-63.
- Staw, B. M. (1976). Knee-deep in the big muddy: A study of escalating commitment to a chosen course of action. Organizational Behavior and Human Performance, 16(1), 27-44. doi.org/10.1016/0030-5073(76)90005-2
- Nickerson, R. S. (1998). Confirmation bias: A ubiquitous phenomenon in many guises. Review of General Psychology, 2(2), 175-220. doi.org/10.1037/1089-2680.2.2.175
- Klein, G. (2007). Performing a project premortem. Harvard Business Review, 85(9), 18-19.
- Ries, E. (2011). The Lean Startup. Crown Business.
- Savoia, A. (2019). The Right It: Why So Many Ideas Fail and How to Make Sure Yours Succeed. HarperOne.
- Fitzpatrick, R. (2013). The Mom Test. CreateSpace Independent Publishing.