Written by: Anish Rao, Head of Growth, Listen Labs
Key Takeaways
- WTP testing tools combine survey methods like Van Westendorp, Gabor-Granger, and conjoint analysis with behavioral tests to estimate maximum customer prices.
- The right method depends on whether you are pricing a new feature or a new offering and whether you need stated preference or behavioral evidence.
- Stated WTP often exceeds actual WTP by 1.35 to three times, so pairing surveys with behavioral validation narrows the say-do gap.
- AI-moderated interviews now deliver adaptive, conversational price ladders at scale, with both quantitative curves and qualitative reasoning.
- Listen Labs pairs the Gabor-Granger price ladder with real-time AI probing and automated deliverables to compress weeks of research into less than 24 hours.
Talk With Listen Labs About WTP Testing
Method-Selection Framework: Answer Two Clarifying Questions
Method choice starts with two questions that shape every downstream decision.
Question 1: Are You Pricing A New Feature Or A New Offering?
Pricing a new feature within an existing product means a reference price and competitive context already exist. Conjoint or MaxDiff can show how much that feature is worth relative to others. Pricing an entirely new offering with no reference price calls for Van Westendorp first. That method establishes the acceptable range before you run any ladder.
Question 2: Do You Need Survey-Based Stated Preference Or Live Behavioral Evidence?
Survey-based methods (Van Westendorp, Gabor-Granger, conjoint) are fast, relatively inexpensive, and available before launch, but they carry hypothetical bias. Hypothetical stated WTP commonly exceeds actual WTP by roughly 1.35 to three times. Behavioral tests (fake door, pre-order, refundable deposit) measure what people actually do. These tests require a live product or landing page and often trade sample size for signal strength.
Map your actual question to the right method:
- “What price feels acceptable?” → Van Westendorp Price Sensitivity Meter
- “How does demand change across price points?” → Gabor-Granger price ladder
- “How much is each feature worth?” → Conjoint analysis or MaxDiff
- “Will people actually pay?” → Behavioral tests (fake door, pre-order, deposit)
This framework-first structure matters because the right tool follows from the right method.
See How Listen Labs Runs A Price Ladder
How To Measure Willingness To Pay: Core Methods And Formulas
Van Westendorp Price Sensitivity Meter
Developed in the 1970s by Dutch economist Peter van Westendorp, the Price Sensitivity Meter (PSM) asks four open-ended price questions:
- At what price would this product be so inexpensive that you would question the quality and not consider buying it? (Too cheap)
- At what price would this product feel like a great deal, a bargain for the money? (Cheap)
- At what price would this product start to feel expensive, but you would still consider buying it? (Expensive)
- At what price would this product be so expensive that you would not consider buying it, regardless of quality? (Too expensive)
The acceptable price range comes from the intersections of cumulative curves. The Point of Marginal Cheapness (PMC) is where the “too cheap” curve crosses the “expensive but acceptable” curve. This point marks the price floor below which more buyers question quality than feel excited about a deal. The Point of Marginal Expensiveness (PME) is where the “too expensive” curve crosses the “bargain” curve. This point marks the price ceiling above which more buyers are priced out than see value. The Optimal Price Point (OPP) is where the “too cheap” and “too expensive” curves intersect.
Formula: Price sensitivity = % change in quantity demanded / % change in price
Van Westendorp can produce directional results with as few as 50 respondents, though most sources recommend 100–300 respondents for reliable results. It works well when sample is limited or no reference price exists. The method does not account for competitive context and functions as a floor-level estimate.
Gabor-Granger
Developed in the 1960s by economists Andre Gabor and Clive Granger, this method presents an adaptive price ladder. Each buyer sees a price and answers yes or no. The next price moves up after a yes and down after a no until the ladder converges on the most that person would pay.
Worked example: “Would you pay $50? No. $40? Maybe. $30? Yes.” Each yes raises the floor and each no lowers the ceiling, narrowing the range until it converges on the most that person would pay.
Formula: Revenue index = Price × Cumulative % willing to buy at that price
Gabor-Granger typically needs 200–400 respondents, or 50+ per price point for monadic designs. It fits settings where a shortlist of candidate prices already exists. Many teams run Van Westendorp first to establish the acceptable range, then Gabor-Granger within that range to find the revenue-maximizing point.
Conjoint Analysis
Conjoint analysis decomposes a product into features and price levels, then derives part-worth utilities for each attribute. WTP for a feature equals its part-worth utility divided by the negative of the price coefficient.
Formula: WTP_feature = β_feature / −β_price
Traditional conjoint research typically costs $40K–$120K and requires roughly 300–500 respondents for stable part-worth utilities, though some sources cite lower minimums of 150–300. Van Westendorp plus Gabor-Granger often produces actionable pricing guidance at a fraction of the cost. Use conjoint when you need to understand how price interacts with specific features across competitive alternatives and when budget and sample support that depth.
Stated Vs. Revealed Preference: When A WTP Survey Misleads You
Hypothetical bias is the gap between what people say they would pay and what they actually pay. Stated WTP commonly exceeds actual WTP by roughly 1.35 to three times, according to Loomis in the Journal of Economic Surveys (2011). The say-do gap is largest in three situations:
- Novel categories with no reference price because respondents have nothing to anchor to, so they guess.
- Socially desirable purchases because people overstate WTP for organic, ethical, or premium products to signal good taste.
- Low-stakes hypothetical choices because when no money is at stake, answers drift upward.
Behavioral and transactional tests narrow this gap. Fake door tests measure interest in a concept, not a specific price point, and they prove that demand exists. Pre-order tests capture real commitment at a specific price and trade sample size for signal strength. Refundable deposit tests filter for genuine intent.
The moderated vs. unmoderated trade-off is real. Moderated sessions catch contradictions but do not scale past a handful of users. Unmoderated testing scales, but it records behavior nobody has time to watch. Listen Labs’ Visual Insights feature addresses this by having the AI interviewer observe on-screen behavior and probe contradictions in real time. This approach narrows the say-do gap at scale while preserving depth.
For more on AI-moderated pricing interviews, see How To Run AI-Moderated Pricing Interviews At Scale.
What Is WTA Vs. WTP?
Willingness to accept (WTA) is the minimum a person would accept to give up a good. Willingness to pay (WTP) is the maximum a person would pay to acquire it. Conventional economic theory expects these to be nearly equal, yet experiments routinely find a WTA-to-WTP ratio of around two to one, sometimes more.
The endowment effect explains this pattern. Once people feel ownership, parting with a good feels like a loss, and losses weigh more heavily than equivalent gains. Free trials and generous return policies work commercially because they move a person from buyer to owner before the price is settled, which raises their valuation.
For pricing research, WTP surveys measure acquisition value, not retention value. A customer who already uses your product will demand more to give it up than a prospect will offer to acquire it. These numbers differ, and treating them as interchangeable leads to pricing decisions that underestimate churn risk.
Is Willingness To Pay The Same As Consumer Surplus?
Consumer surplus is the gap between what a buyer would have paid and what they actually paid: CS = WTP − Price. WTP testing estimates the ceiling, the maximum a buyer will accept. Consumer surplus describes the value captured below that ceiling.
If a buyer values a tool at $120 but pays $90, their WTP is $120 and the $30 difference is consumer surplus the seller left on the table. Pricing is the exercise of deciding how to split the distance between WTP and cost. Willingness to pay testing tools help locate the ceiling, and pricing strategy determines how much of the distance between cost and ceiling to capture. With that distinction in mind, the tools below are organized by the question you are trying to answer.
Willingness To Pay Testing Tools: A Guide By Use Case
Organized by the question you are trying to answer. Each entry describes what the tool is built for and where it falls short.
Sawtooth Software focuses on conjoint and market simulator depth. Sawtooth’s Lighthouse Studio supports CBC, adaptive CBC, MaxDiff, and hierarchical Bayes estimation, with full licenses running to five figures annually. Users must source sample separately, so it suits teams with a dedicated analyst and longer timelines.
Conjointly focuses on accessible conjoint and pricing research. Conjointly bundles CBC, MaxDiff, Gabor-Granger, and Van Westendorp with templated setup and an integrated panel, but its reporting explains what won, not why. It fits teams that need structured pricing outputs more than deep qualitative reasoning.
PriceBeam focuses on pricing-specific research with global cloud-based WTP and revenue management capabilities. It suits teams that want pricing dashboards more than broad concept testing. It fits less well when you need to test features or messaging alongside price.
SurveyMonkey focuses on fast, inexpensive stated-preference surveys. SurveyMonkey is a standard platform for administering Van Westendorp and Gabor-Granger surveys. It does not provide adaptive probing, behavioral capture, or qualitative depth, so it fits teams that only need structured survey data. See also Best Survey Platforms For Pricing Research In 2026.
Jotform focuses on lightweight forms for simple price questions. It does not function as a full research platform, so it fits only very simple studies without methodology demands.
Listen Labs focuses on conversational, AI-moderated WTP testing that pairs the Gabor-Granger price ladder with qualitative “why” at scale. It combines the price ladder, adaptive qualitative probing, and automated deliverables in one platform, so it fits teams that want both curves and context.
Book A Demo To See The Conversational WTP Test
Listen Labs For Conversational WTP Testing
Listen Labs provides an end-to-end solution for willingness to pay testing by pairing the Gabor-Granger price ladder with the qualitative “why” behind every price decision at scale.

The Gabor-Granger Pricing Test runs as a conversational price ladder inside studies teams already field. Each buyer sees a product description and decides whether they would buy at a given price. Their answer determines the next price they see, higher after a yes and lower after a no, until the ladder converges on the most that person would pay.
Contextual follow-up questions separate “I cannot afford it” from “it is not worth that much.” These represent very different problems with very different fixes. The AI interviewer keeps probing, then codes those answers into themes so individual anecdotes become structured data.
Outputs include:

- Demand curve that shows how many buyers you keep at each price
- Revenue curve that highlights the price that earns the most
- Segment filters that let you see how key audiences behave without fielding a new study
Messy answers, handled: Buyers who would not purchase at any price stay in the count so demand is not artificially inflated. Buyers who say yes to a high price but no to a lower one are counted separately. Half-finished ladders drop from the analysis.
Why this improves on a static survey for WTP: Surveys deliver structured, quantitative data through pre-set questions with no ability to follow up or probe deeper. Listen Labs conducts conversational interviews where the AI adapts in real time, so you see both numbers and narrative.
MaxDiff capability supports prioritization without ties. Instead of rating scales, which produce ties, or full rankings, which exhaust respondents, MaxDiff shows small sets and asks for best and worst, then aggregates with a Hierarchical Bayes model into a clean ranking. The AI moderator follows up on each choice with prompts like “You rated Strawberry highest. What makes it a buy?”
Every point on the curve connects to a real interview, so when someone challenges the results, you can answer with clips of buyers explaining their choices in their own words.
Listen Labs compresses a 4–6 week research cycle to less than 24 hours, costs about a third of traditional research, reaches 50M+ verified respondents across 45+ countries and 120+ languages, and is trusted by enterprises including Microsoft, Google, Anthropic, P&G, Skims, and Sweetgreen.

For complex projects, Listen’s insights team of career researchers provides white-glove service to help find the right price and understand the why behind the number.
Get Your Price Ladder Results In Under 24 Hours
AI-Moderated And Synthetic WTP Testing: Where Each Fits
AI-moderated WTP testing complements validated methods like Van Westendorp, Gabor-Granger, and conjoint rather than replacing them.
Where synthetic respondents are reliable: early-stage directional testing and hypothesis generation. A 2026 Strat7 study found synthetic respondents in pricing modules gave WTP prices generally 16% above those provided by real people. That pattern supports directional insight, not precision. In the same study, synthetic respondents broke logical price ordering 68% of the time in a four-step price ordering exercise.
Where they fall short: any price decision that carries real revenue risk. Strat7’s group AI lead Hasdeep Sethi stated, “If you want precision, then you still need to ask real people.”
How AI-moderated interviews differ from synthetic respondents: AI-moderated interviews talk to real, verified participants and adapt follow-ups in the moment, while synthetic respondents generate simulated answers. Only the first produces research-grade evidence.
Common Pitfalls And How To Avoid Them
- Asking about price before establishing value. Respondents anchor on price before they understand what they are buying, so start with value perception questions before any mention of price.
- Using a single method for every question. Van Westendorp will not tell you feature value, and conjoint will not tell you if people will actually buy, so match the method to the question.
- Ignoring the say-do gap. As noted earlier, stated WTP overestimates actual WTP by 1.35 to three times, so pair stated-preference methods with behavioral tests when stakes are high.
- Running conjoint without the sample size to support it. Conjoint studies need the larger sample sizes noted above, so use Van Westendorp or Gabor-Granger when sample is limited.
- Treating a hypothetical yes as a purchase commitment. “Probably would buy” differs from buying, so apply a discount factor and validate with transactional data.
Most WTP mistakes are method-selection mistakes. For a related look at monadic approaches, see Monadic Price Testing Tools: A Pricing Research Guide and Monadic Price Testing Software: How To Choose The Right Tool.
Frequently Asked Questions
What’s The Difference Between Van Westendorp And Gabor-Granger, And Which Should I Use?
Van Westendorp uses four open-ended questions to find an acceptable price range. It works best when no reference price exists and you need to establish where the market’s psychological floor and ceiling sit before committing to specific price points. Gabor-Granger presents specific prices and asks yes or no at each to build a demand curve and identify the revenue-maximizing price. Use Van Westendorp when starting from scratch. Use Gabor-Granger when you have a shortlist of candidate prices and need to understand demand drop-off between them. Many teams run both in sequence, with Van Westendorp first to establish the range and Gabor-Granger within that range to find the optimal point.
How Many Respondents Do I Need For A Conjoint Study Versus A Simple WTP Survey?
Conjoint studies commonly require roughly 300–500 respondents for stable part-worth utilities, though some sources cite lower minimums of 150–300, and more if you plan to analyze subgroups separately, since each segment needs its own viable sample of 200–300. Van Westendorp can produce directional results with as few as 50 respondents, though most sources recommend 100–300 respondents for reliable results. Gabor-Granger typically needs 200–400 respondents, or 50+ per price point for monadic designs. Budget and timeline usually determine which method is practical. Van Westendorp plus Gabor-Granger in a single survey often suits teams that need actionable guidance without a full conjoint budget.
When Should I Run A Fake-Door Or Pre-Order Test Instead Of A Survey?
Run behavioral tests when the stakes justify behavioral evidence and you have a live product or landing page. Fake-door tests measure interest in a concept at a specific price and prove that demand exists, but they do not reveal the full demand curve. Pre-order tests capture real commitment at a specific price and represent the strongest stated-preference alternative to a full market launch. Use behavioral tests when stated preference alone will not convince stakeholders, when the category is novel enough that hypothetical bias is likely to be severe, or when you need to validate a survey-based finding before acting on it.
What Is The Difference Between WTP And WTA, And Between WTP And Consumer Surplus?
WTA is the minimum a person would accept to give up a good, and WTP is the maximum they would pay to acquire it. WTA typically exceeds WTP by a ratio of around two to one due to the endowment effect, where ownership makes losses feel heavier than equivalent gains. Consumer surplus is the gap between what a buyer would have paid and what they actually paid, expressed as CS = WTP − Price. WTP testing estimates the ceiling, and consumer surplus describes the value captured below it. Pricing strategy decides how much of the distance between cost and WTP ceiling to capture.
Can AI-Moderated Interviews Replace A Traditional Pricing Survey?
AI-moderated interviews with real participants can replace or augment traditional surveys when you need qualitative depth and adaptive probing. They surface the reasoning behind price decisions that static surveys cannot reach. They do not replace the statistical rigor of a well-designed conjoint study for feature-level WTP, and they differ from synthetic respondents, which generate simulated answers rather than real human testimony. A practical framework uses AI-moderated interviews for the “why” and qualitative context, Van Westendorp or Gabor-Granger for the quantitative price range and demand curve, and conjoint when you need feature-level trade-off data and have the sample and budget to support it.
Conclusion: Applying The Method-To-Tool Framework
Effective WTP work starts by identifying the question you are trying to answer, choosing the method that fits, and then selecting the tool that runs that method well.
Start by defining the pricing decision, since that choice determines which method fits. Run a small pilot before committing, and pair stated-preference results with behavioral evidence whenever the stakes justify it. Once the product is live, validate the stated WTP estimate against transactional data rather than treating it as final.
Teams that need the price ladder and the “why” behind it in one platform, with results in less than 24 hours, can use Listen Labs. Each point on the curve connects to a real interview, so you can respond to challenges with clips of buyers explaining their choices in their own words.


