MaxDiff (best/worst) scaling

Anchored, adaptive, and standard MaxDiff methodology and validation.
Shapley Values for MaxDiff Results: Overview and Recommendations for Practitioners (2026)
May 2026
21
Read
Bryan Orme, Sawtooth & David W. Lyon, Aurora Market Modeling LLC

Shapley values estimate each item’s contribution to TURF reach by averaging its marginal impact across all portfolio combinations. Unlike average utility scores, they account for redundancy and reward items that uniquely reach underserved segments. They are more stable than traditional TURF results and easier to interpret, providing a single importance score per item. Efficient methods like COA designs make them practical even for large studies. Their value depends on the reach definition used—threshold-based approaches tend to reveal the most insight. Overall, Shapley values extend TURF by delivering clearer, more reliable measures of item importance.

Real-Time Detection of Random Respondents in MaxDiff (2022)
February 2022
13
Read
Keith Chrzan and Bryan Orme, Sawtooth

In this paper we demonstrate a method that enables you to identify random respondents during a MaxDiff survey. This means that you may be able to tag and disqualify poor quality respondents before they complete your survey and before they qualify for compensation. After describing the method and illustrating step-by-step how to program it in our software, we describe the empirical tests we’ve conducted of the method, we apply it to some commercial studies, and we make recommendations about how best to apply it.

Statistical Testing - MaxDiff (2020)
November 2020
42
Read
Sawtooth

Chapter 6 exerpt from the book Applied MaxDiff.

Sparse, Express, Bandit, Relevant Items, Tournament, Augmented, and Anchored MaxDiff—Making Sense of All Those MaxDiffs! (2019)
May 2019
10
Read
Bryan Orme, Sawtooth

Over the last 15 years, new flavors of MaxDiff have been devised to extend its capabilities to measure more items or to address drawbacks. This short white paper describes seven different flavors of MaxDiff that have been discussed at Sawtooth Software conferences, followed by references so you can read more details about them.

How Good Is Best-Worst Scaling? (2018)
July 2018
0
Read
Bryan Orme, Sawtooth

Best-worst scaling (BWS) gives you better information with fewer respondents—it works better than traditional rating scales or constant sum questions, achieving better discrimination among the items and finding a larger number of statistically significant differences between groups of respondents. BWS yields more than just rank-order scaling and is able to recover relative metric differences among items. BWS provides more accurate information than 10-point ratings and constant sums for classifying respondents and making individual-level predictions. If the BWS scores need to be scaled relative to a buy/no buy or important/not important threshold, anchored BWS is a straightforward solution.

Bandit MaxDiff: When to Use It and Why It Can Be a Better Choice than Standard MaxDiff (2018)
February 2018
30
Read
Bryan Orme, Sawtooth

Bandit MaxDiff (best-worst scaling) achieves greater measurement precision than standard MaxDiff for items that have the highest utility scores. When that is your focus, you can vastly improve the efficiency for large MaxDiff projects, reducing data collection costs by 50% to 80%. Bandit MaxDiff is a new type of adaptive MaxDiff available in Lighthouse Studio v9.6 and later. Whether studying relatively few items or an enormous number of items, Bandit MaxDiff can significantly improve your results.

A Parameter Recovery Experiment for Two Methods of MaxDiff with Many Items (2015)
January 2015
41
Read
Keith Chrzan, Sawtooth

Clients don’t seem to be able to get enough of a good thing and this seems to apply more to MaxDiff than to some of the other methods we use: clients frequently ask for MaxDiff experiments that include more items than would allow us to expose each item to each respondent the recommended three or four times. In 2012, Wirth and Wolfrath introduced “Express” MaxDiff and “Sparse” MaxDiff as ways to handle more than the usual number of items. Express MaxDiff draws a subset of items for each respondent and shows each item multiple times. Sparse MaxDiff typically shows each item perhaps just once to each respondent, trying to obtain as much coverage as possible for each respondent over the items. The author, Keith Chrzan, investigates whether Express or Sparse MaxDiff does a better job of recovering “true” utilities using simulated robotic respondents. For two data sets that were based on real patterns of live humans’ preferences, he finds a modest edge in performance for Sparse MaxDiff.Express MaxDiff has the benefit of repeated measures at the individual level. Sparse has the benefit of avoiding so many missing items at the individual level. Both methods may be analyzed via counts, aggregate logit, latent class, or HB. Even though it is possible to estimate individual-level scores via HB, it is relatively imprecise under Express or Sparse MaxDiff compared to the recommended MaxDiff study that shows each item to each respondent about 3 or 4 times.

Common Scale Hybrid Discrete Choice Analysis: Fusing Best-Worst Case 2 and 3 (2013)
November 2013
51
Read
Bryan Orme, Sawtooth

The author (Orme) describes a hybrid discrete choice method that results in conjoint utilities on a common utility scale (where comparisons across levels of different attributes are supported). For certain research situations and when working with certain types of clients, there are valuable benefits to having commonly-scaled conjoint utilities. The key to obtaining the common utility scale is to employ both MaxDiff and CBC-type questions within the same questionnaire.

Although MaxDiff and CBC utilities are shown to not be equivalent, they tend to be very similar (correlation of about 0.90 or higher). Evidence is presented that MaxDiff scores can perform almost at the same level as CBC utilities in terms of predicting CBC-formatted holdout tasks. Orme suggests that if the goal is to predict CBC-looking holdouts (or real-world purchase decisions), the CBC tasks alone could be used to develop a market simulator. Not surprisingly, fusing MaxDiff and CBC slightly degrades predictions of CBC-formatted holdout tasks. If the goal is to present and interpret part-worth utilities on a common scale, while leveraging the strength of the conjoint approach to preference elicitation, then the fusion of MaxDiff and CBC tasks certainly has benefits.

MaxDiff Technical Paper (2020)
February 2013
2
Read
Sawtooth

This paper describes the technical procedures used in the MaxDiff System. MaxDiff (best-worst) scaling is a trade-off method for measuring the importance or preference for multiple items, such as brands, product features, political platforms, advertising claims, etc. Any time you are considering using a rating scale, ranking scale, or constant sum scale for multiple items, you can consider using MaxDiff.

The MaxDiff methodology, originally invented by researcher and academic Jordan Louviere, has gained in popularity over the last five years. Papers on MaxDiff have won "best presentation" awards at recent ESOMAR and Sawtooth Software research conferences. It has many similarities to, but is distinctively different, from conjoint methodology and is appropriate for a wider range of research opportunities.

Sawtooth Software’s MaxDiff System may be used for conducting web-based, CAPI, or paper-based MaxDiff studies. The software also supports asking the "best" half of the question only (not requiring respondents to identify the "worst" item in each set). The software may also be used for Method of Paired Comparisons research. Individual-level estimation of item scores employs Sawtooth Software’s popular hierarchical Bayes (HB) engine. Results may also be analyzed with the integrated Latent Class procedure for segmentation analysis.

Anchoring MaxDiff Scaling Against a Threshold - Dual Response and Direct Binary Responses (2010)
July 2010
32
Read
Kevin Lattery, Maritz Research

Maximum Difference Scaling is widely used to measure the relative values of items/attributes. Despite the strengths of MaxDiff, some analysts would prefer data that represented more than just relative scores; they would prefer absolute scores scaled with respect to each respondent’s importance threshold. In this article by Kevin Lattery (Maritz Research), Kevin tested two methods for anchoring MaxDiff scores to a threshold: Dual-Response MaxDiff suggested by Louviere and a more direct method asking respondents to choose which attributes are above a threshold (using 2-point scale grid questions).

He determined that theoretically (using synthetic respondent data) the direct method would be superior, especially as the number of attributes shown in a MaxDiff task increases. With six or more attributes shown per screen the indirect dual-response method should not be used, and even five attributes per screen may not capture individual anchoring that well. In comparing the two methods with human respondents, showing only four attributes per screen, results were very similar. The rank order of utilities at the respondent level was nearly identical. However, the anchoring in the direct method was more biased by the context of the total set of attributes. So if it is important for one to have a more neutral anchor for utilities then the indirect dual-response method may be slightly better, assuming four (and certainly no more than five) attributes are shown per screen.

(Originally published in the 2010 Sawtooth Software Proceedings).

Anchored Scaling in MaxDiff Using Dual-Response (2009)
September 2009
31
Read
Bryan Orme, Sawtooth

Traditional MaxDiff analysis leads to relative importance/preference scores. But, there is no possible way for respondents to express that (for example) all the items are important or none of the items are important. Some researchers have worried that the relative nature of the MaxDiff judgments and resulting scale means that meaningful differences between respondents or segments of respondents are lost.This article describes a practical way, proposed by Jordan Louviere (inventor of MaxDiff), to anchor the scale for each respondent based on an important/not important threshold. A dual-response questioning device is straightforward to include in MaxDiff questionnaires, and in analysis. The author (Orme) provides empirical evidence that the dual-response approach leads to meaningful discrimination among respondents and items, beyond the information provided by the standard MaxDiff tradeoffs. The pros and cons of the approach are discussed.

MaxDiff Analysis: Simple Counting, Individual-Level Logit, and HB (2009)
June 2009
50
Read
Bryan Orme, Sawtooth

This paper compares different methods of obtaining individual-level scores for MaxDiff surveys at the individual level: Simple counting, individual-level logit, and HB. Key to the success for all these methods was having enough information available for each respondent to estimate stable scores.

The author (Orme) finds that counting analysis provides reasonable population estimates of scores, but that the individual-level scores can lack precision. Precision is better under the logit model estimation methods: either individual-level logit or HB, which "borrows" information across the sample to improve the individual-level logit scores for individuals.

Despite the simplicity of the counting approach and its weaknesses, it tends to do quite well in predicting responses to holdout choices. But, across-respondent variance (heterogeneity) tends to be weaker than the other methods studied.

Using Calibration Questions to Obtain Absolute Scaling in MaxDiff (2009)
March 2009
22
Read
Bryan Orme, Sawtooth

In this paper, we create an artificial situation that demonstrates the relative scaling issue for MaxDiff at its worst. We collect a first wave of MaxDiff data on 30 items, and based on the items’ average scores we separate them into the best 15 and the worst 15 items. Then, we give two new sets of respondents MaxDiff questionnaires that include either the best 15 or the worst 15 items as determined from Wave 1 (plus a few calibration questions).

Political Landscape 2008: Segmentation Using MaxDiff and Cluster Ensemble Analysis (2008)
September 2008
40
Read
Bryan Orme & Chris King, Sawtooth

This article provides a case study regarding how MaxDiff and Cluster Ensemble analysis can be used to segment a population. Sawtooth Software conducted an online study among US respondents just prior to the 2008 presidential election between John McCain and Barack Obama. The 2008 Political Landscape study measured what policy positions would make people most and least want to vote for a US presidential candidate. We studied 25 different policy positions, including such items as:

  • Ensure the long-term health of Social Security
  • Bring the troops home from Iraq
  • Restrict carbon emissions to reduce global warming
  • Reduce the federal deficit
  • We used MaxDiff (best-worst) scaling to measure the influence of these items on respondents’ preference for a presidential candidate. And, of course, Democrats and Republicans have different priorities when it comes to policies and positions.

Following estimation of individual-level scores from MaxDiff, we used our CCEA (Convergent Cluster & Ensemble Analysis) software to find segments of respondents that had similar preferences.