I have been looking for a model to think through the role of research and its possible replacement / transformation by AI in product development. Do we still need research? Done by humans? Or can synthetic users and / or researchers answer all the questions the PM may have?

On a recent Plain English episode, Derek Thompson and economist Alex Imas were discussing why the AI job apocalypse has not happened yet. Imas brought up a theory that resonated with me: Michael Kremer's O-ring model, published in 1993 and named after the gasket that doomed the Challenger, which happened 40 years ago in January.

The idea of the O-ring model is simple. In complex processes, tasks multiply rather than add. If all the tasks are executed perfectly then all is well. But ten excellent solutions and one careless one don't make a 90% excellent product: an oversalted omelet on beautiful china in a fancy restaurant overlooking the ocean ends up inedible.

I kept thinking about product development —what happens when one link in that chain gets quietly swapped for an AI approximation? Will it doom the product?

Well, no, at least not necessarily.

The O-ring model doesn't say a product without research fails. The O-ring model suggests a multiplication problem, where every factor has the potential to discount the whole. Skip research, or replace it with a bad approximation, and the product's expected value gets multiplied by the probability that your assumptions about users happen to be right.

The product is the product of task success probability values, so if research is only 0.5 trustworthy, then the other components (development, engineering, design) is multiplied by 0.5

The problem is that we perceive that good case scenario probability a lot closer to 1 (100%) than what it truly is. When founders win that bet, when those wins get retold at every product conference, we fall victim to the survivorship bias. We remember PayPal but not Fast for easy payment.  Not enough humble product managers write the post-mortem titled "we built on assumptions and the assumptions were wrong."  Basically, this fallacy ends up justifying saving money on research.

Kremer's model also predicts when the bet on lean product development gets expensive. Failure intolerance rises with complexity and stakes. A landing page may not need in-depth testing, but a clinical workflow tool, a banking app, or an insurance enrollment system cannot afford to skip usability testing with real humans. Those are the chains where a damaged O-ring is more likely to lead to product failure.

So what happens to the research factor when AI does the researching? We have some data.

Nielsen Norman Group tested synthetic users against their own prior studies with real participants. They found that the AI responses were not only too shallow to be useful, but worse, misleading as they were often overly uncritical and idealized. Synthetic users praised concepts that real users went on to question or reject. A tree test study by MeasuringU compared ChatGPT performance to that of 33 humans and found that ChatGPT aced a test designed to measure where humans fail. So, AI navigated like a 99th percentile superuser, not like actual users (assuming we are still designing for humans). Any upside for UX research? Well, if ChatGPT can't find something on a website, real users almost certainly can't either, so AI can be used as a cheap early smoke detector that beeps only for the biggest fire. But its silence cannot reassure us that everything is fine.

Bisbee et al. (2024) ‘s study on so-called synthetic survey data on public opinion revealed similar averages to that of their baseline data, but once the researchers dug deeper, they found that the silicon data was unusable for statistical analyses (with low variability and wrong regression coefficients), making it impossible to detect subgroups. They also found that minor changes in the prompt wording led to very different results. In other words, low quality data and no reliability

I taught statistics for more than a decade, so let me translate. Synthetic data looks right at the level of the summary slide and wrong at the level where teams make product decisions:  at the level of segments, variance, relationships.

The purpose of research is to learn about people you don't yet understand (that’s why I enjoy UX research).  A model generating plausible answers from past internet text is, by definition, incapable of giving you insights. So, it’s boring. But actually,  it is dangerous  because it will amplify the confirmation bias of decision makers, and lead them down a wrong path.

The Challenger failed in a way that is remembered for decades. The thorough research by the Rogers commission identified what went wrong. But product development failures don't work like that. The wrong product gets shipped, and the damage arrives eighteen months later as churn, low adoption, a quiet sunset. Diffuse, delayed, unattributable, with no headlines (and no smoking gun case study).

Back to Kremer’s O-Ring model. Transcription got automated, rightly, because transcription errors are cheap and self-evident — you catch them on the next read. Interpretation errors are expensive and go undetected. The tasks safest to automate are the ones where failure announces itself. The tasks companies are now rushing to automate are the ones where it doesn't.

For website references, please visit the linked webpages

  • Bisbee, J., Clinton, J. D., Dorff, C., Kenkel, B., & Larson, J. M. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis32(4), 401–416. doi:10.1017/pan.2024.5
  • Kremer, Michael (1993). "The O-Ring Theory of Economic Development". The Quarterly Journal of Economics108 (3): 551–575. doi:10.2307/2118400JSTOR 2118400