MaxDiff (best/worst) scaling
Shapley values estimate each item’s contribution to TURF reach by averaging its marginal impact across all portfolio combinations. Unlike average utility scores, they account for redundancy and reward items that uniquely reach underserved segments. They are more stable than traditional TURF results and easier to interpret, providing a single importance score per item. Efficient methods like COA designs make them practical even for large studies. Their value depends on the reach definition used—threshold-based approaches tend to reveal the most insight. Overall, Shapley values extend TURF by delivering clearer, more reliable measures of item importance.
In this paper we demonstrate a method that enables you to identify random respondents during a MaxDiff survey. This means that you may be able to tag and disqualify poor quality respondents before they complete your survey and before they qualify for compensation. After describing the method and illustrating step-by-step how to program it in our software, we describe the empirical tests we’ve conducted of the method, we apply it to some commercial studies, and we make recommendations about how best to apply it.
Chapter 6 exerpt from the book Applied MaxDiff.
Over the last 15 years, new flavors of MaxDiff have been devised to extend its capabilities to measure more items or to address drawbacks. This short white paper describes seven different flavors of MaxDiff that have been discussed at Sawtooth Software conferences, followed by references so you can read more details about them.
Best-worst scaling (BWS) gives you better information with fewer respondents—it works better than traditional rating scales or constant sum questions, achieving better discrimination among the items and finding a larger number of statistically significant differences between groups of respondents. BWS yields more than just rank-order scaling and is able to recover relative metric differences among items. BWS provides more accurate information than 10-point ratings and constant sums for classifying respondents and making individual-level predictions. If the BWS scores need to be scaled relative to a buy/no buy or important/not important threshold, anchored BWS is a straightforward solution.
Bandit MaxDiff (best-worst scaling) achieves greater measurement precision than standard MaxDiff for items that have the highest utility scores. When that is your focus, you can vastly improve the efficiency for large MaxDiff projects, reducing data collection costs by 50% to 80%. Bandit MaxDiff is a new type of adaptive MaxDiff available in Lighthouse Studio v9.6 and later. Whether studying relatively few items or an enormous number of items, Bandit MaxDiff can significantly improve your results.
Clients don’t seem to be able to get enough of a good thing and this seems to apply more to MaxDiff than to some of the other methods we use: clients frequently ask for MaxDiff experiments that include more items than would allow us to expose each item to each respondent the recommended three or four times. In 2012, Wirth and Wolfrath introduced “Express” MaxDiff and “Sparse” MaxDiff as ways to handle more than the usual number of items. Express MaxDiff draws a subset of items for each respondent and shows each item multiple times. Sparse MaxDiff typically shows each item perhaps just once to each respondent, trying to obtain as much coverage as possible for each respondent over the items. The author, Keith Chrzan, investigates whether Express or Sparse MaxDiff does a better job of recovering “true” utilities using simulated robotic respondents. For two data sets that were based on real patterns of live humans’ preferences, he finds a modest edge in performance for Sparse MaxDiff.Express MaxDiff has the benefit of repeated measures at the individual level. Sparse has the benefit of avoiding so many missing items at the individual level. Both methods may be analyzed via counts, aggregate logit, latent class, or HB. Even though it is possible to estimate individual-level scores via HB, it is relatively imprecise under Express or Sparse MaxDiff compared to the recommended MaxDiff study that shows each item to each respondent about 3 or 4 times.
The author (Orme) describes a hybrid discrete choice method that results in conjoint utilities on a common utility scale (where comparisons across levels of different attributes are supported). For certain research situations and when working with certain types of clients, there are valuable benefits to having commonly-scaled conjoint utilities. The key to obtaining the common utility scale is to employ both MaxDiff and CBC-type questions within the same questionnaire.
Although MaxDiff and CBC utilities are shown to not be equivalent, they tend to be very similar (correlation of about 0.90 or higher). Evidence is presented that MaxDiff scores can perform almost at the same level as CBC utilities in terms of predicting CBC-formatted holdout tasks. Orme suggests that if the goal is to predict CBC-looking holdouts (or real-world purchase decisions), the CBC tasks alone could be used to develop a market simulator. Not surprisingly, fusing MaxDiff and CBC slightly degrades predictions of CBC-formatted holdout tasks. If the goal is to present and interpret part-worth utilities on a common scale, while leveraging the strength of the conjoint approach to preference elicitation, then the fusion of MaxDiff and CBC tasks certainly has benefits.
This paper describes the technical procedures used in the MaxDiff System. MaxDiff (best-worst) scaling is a trade-off method for measuring the importance or preference for multiple items, such as brands, product features, political platforms, advertising claims, etc. Any time you are considering using a rating scale, ranking scale, or constant sum scale for multiple items, you can consider using MaxDiff.
The MaxDiff methodology, originally invented by researcher and academic Jordan Louviere, has gained in popularity over the last five years. Papers on MaxDiff have won "best presentation" awards at recent ESOMAR and Sawtooth Software research conferences. It has many similarities to, but is distinctively different, from conjoint methodology and is appropriate for a wider range of research opportunities.
Sawtooth Software’s MaxDiff System may be used for conducting web-based, CAPI, or paper-based MaxDiff studies. The software also supports asking the "best" half of the question only (not requiring respondents to identify the "worst" item in each set). The software may also be used for Method of Paired Comparisons research. Individual-level estimation of item scores employs Sawtooth Software’s popular hierarchical Bayes (HB) engine. Results may also be analyzed with the integrated Latent Class procedure for segmentation analysis.
Maximum Difference Scaling is widely used to measure the relative values of items/attributes. Despite the strengths of MaxDiff, some analysts would prefer data that represented more than just relative scores; they would prefer absolute scores scaled with respect to each respondent’s importance threshold. In this article by Kevin Lattery (Maritz Research), Kevin tested two methods for anchoring MaxDiff scores to a threshold: Dual-Response MaxDiff suggested by Louviere and a more direct method asking respondents to choose which attributes are above a threshold (using 2-point scale grid questions).
He determined that theoretically (using synthetic respondent data) the direct method would be superior, especially as the number of attributes shown in a MaxDiff task increases. With six or more attributes shown per screen the indirect dual-response method should not be used, and even five attributes per screen may not capture individual anchoring that well. In comparing the two methods with human respondents, showing only four attributes per screen, results were very similar. The rank order of utilities at the respondent level was nearly identical. However, the anchoring in the direct method was more biased by the context of the total set of attributes. So if it is important for one to have a more neutral anchor for utilities then the indirect dual-response method may be slightly better, assuming four (and certainly no more than five) attributes are shown per screen.
(Originally published in the 2010 Sawtooth Software Proceedings).
Traditional MaxDiff analysis leads to relative importance/preference scores. But, there is no possible way for respondents to express that (for example) all the items are important or none of the items are important. Some researchers have worried that the relative nature of the MaxDiff judgments and resulting scale means that meaningful differences between respondents or segments of respondents are lost.This article describes a practical way, proposed by Jordan Louviere (inventor of MaxDiff), to anchor the scale for each respondent based on an important/not important threshold. A dual-response questioning device is straightforward to include in MaxDiff questionnaires, and in analysis. The author (Orme) provides empirical evidence that the dual-response approach leads to meaningful discrimination among respondents and items, beyond the information provided by the standard MaxDiff tradeoffs. The pros and cons of the approach are discussed.
This paper compares different methods of obtaining individual-level scores for MaxDiff surveys at the individual level: Simple counting, individual-level logit, and HB. Key to the success for all these methods was having enough information available for each respondent to estimate stable scores.
The author (Orme) finds that counting analysis provides reasonable population estimates of scores, but that the individual-level scores can lack precision. Precision is better under the logit model estimation methods: either individual-level logit or HB, which "borrows" information across the sample to improve the individual-level logit scores for individuals.
Despite the simplicity of the counting approach and its weaknesses, it tends to do quite well in predicting responses to holdout choices. But, across-respondent variance (heterogeneity) tends to be weaker than the other methods studied.
In this paper, we create an artificial situation that demonstrates the relative scaling issue for MaxDiff at its worst. We collect a first wave of MaxDiff data on 30 items, and based on the items’ average scores we separate them into the best 15 and the worst 15 items. Then, we give two new sets of respondents MaxDiff questionnaires that include either the best 15 or the worst 15 items as determined from Wave 1 (plus a few calibration questions).
This article provides a case study regarding how MaxDiff and Cluster Ensemble analysis can be used to segment a population. Sawtooth Software conducted an online study among US respondents just prior to the 2008 presidential election between John McCain and Barack Obama. The 2008 Political Landscape study measured what policy positions would make people most and least want to vote for a US presidential candidate. We studied 25 different policy positions, including such items as:
- Ensure the long-term health of Social Security
- Bring the troops home from Iraq
- Restrict carbon emissions to reduce global warming
- Reduce the federal deficit
- We used MaxDiff (best-worst) scaling to measure the influence of these items on respondents’ preference for a presidential candidate. And, of course, Democrats and Republicans have different priorities when it comes to policies and positions.
Following estimation of individual-level scores from MaxDiff, we used our CCEA (Convergent Cluster & Ensemble Analysis) software to find segments of respondents that had similar preferences.
The author (Orme) presents results from two studies testing a new procedure called Adaptive MaxDiff Scaling. Rather than focus equal attention on estimating respondents' preferences (or importances) for best AND worst items, A-MaxDiff focuses attention on estimating best/most important items with greater precision. The interview adapts to each respondent, learning from prior responses. Items marked "worst" are discarded from further consideration. The questionnaire proceeds in stages. In the first stage, K items are shown per set. In each subsequent stage, K-1 items are shown per set, until the respondent is doing paired comparisons among the surviving (most preferred) items. Later tasks reflect increased utility balance.
The results show better hit rates for "best" items in holdouts relative to standard MaxDiff. Average population parameters are essentially identical between standard and adaptive forms of MaxDiff. Respondents take slightly less time to complete the adaptive survey, and they perceive it to be more enjoyable and less monotonous than standard MaxDiff. Orme argues that A-MaxDiff should be especially preferred when simulation methods such as TURF are used with MaxDiff data. The main drawback is decreased precision of estimates for "worst" items.
The authors investigate how the number of items per MaxDiff set affects dropout rates, survey length, positional bias, parameter equivalence, and predictive validity. Three commercial studies are analyzed, where the number of items per set varied from 3 items/set to 8 items/set. The number of items/set has the most influence on task length, with respondents taking significantly longer to complete 8 items/set rather than 3 items/set. Statistically significant differences among the parameters were found, but the authors note that the overall results would lead to similar managerial decisions. The predictive validity tests "hint that 3 items per question may produce slightly worse predictions than questions with more items." They conclude: "Given the slight evidence of poorer hit rates and poorer out-of-sample for 3 items per question we recommend using 4 or 5 items per question in maxdiff experiments."
This paper communicates results of a Monte Carlo simulation study on how the precision of estimates for MaxDiff (best/worst) experiments is affected by:
- Number of items presented per set
- Number of sets presented to each respondent
- Number of items in the overall study
Results show that it may not be useful to ask more than about 5 items per set. The data also suggest that displaying each item 3 or more times per respondent works well for obtaining reasonably precise individual-level estimates with HB. Asking more tasks, such that the number of exposures per item is increased well beyond 3, seems to offer significant benefit, provided respondents don't become fatigued and provide data of reduced quality.
This article offers a case study demonstrating how best/worst scaling may be used for estimating the price sensitivity of automobile buyers to different car options, such as warranty, anti-lock brakes, and keyless entry. Using best/worst scaling, the author (Keith Chrzan), shows how price sensitivity curves may be developed for each car option, presented on a common scale. Chrzan contrasts the best/worst approach with conjoint analysis, and explains the benefits of using best/worst for this particular application rather than conjoint analysis. The results validate closely to self-reported past purchase behavior for options on the most recently purchased car.
Chrzan provides some background on best/worst, and describes how Sawtooth Software's latent class and HB software may be used during analysis. This paper was voted "best presentation" at the 2004 Sawtooth Software Conference.
Maximum Difference (MaxDiff, or best/worst) scaling is a relatively new technique for measuring the importance or preference of multiple items. In MaxDiff tasks, respondents see sets of items (typically 4 to 6). In each set, respondents indicate which item is most important (preferred) and least important (preferred). Steve Cohen describes the methodology and presents results for a methodological study comparing MaxDiff measurement with monadic ratings and paired comparisons, and also a case study focusing on using MaxDiff for segmentation work. MaxDiff is shown to provide results that have greater between-item and between-respondent discrimination, and greater predictive accuracy than either monadic ratings or paired comparisons. Steve won the "best presentation" award with this paper at the 2003 Sawtooth Software Conference.