Synthetic data is a polarizing topic, simultaneously touted as a solution for hard-to-reach populations and small-sample studies, and condemned as confidently misrepresenting the voice of the customer. In this session, we present a validation framework along with findings from our ongoing evaluation of synthetic data across common tasks including cross tabs, key driver analysis, and conjoint. Attendees will leave with practical guidance for evaluating synthetic data outcomes along with validated use cases.
We outline a practical framework for improving online survey data quality by integrating behavioral, device, and historical feedback signals. Traditional checks like RLH risk discarding genuine respondents or retaining fraudulent ones, undermining insights. The research demonstrates how combining in-survey behavioral indicators, device fraud detection, and cross-platform feedback loops creates a dynamic, adaptive quality management system. By leveraging Data Quality Co-op's API within Lighthouse Studio, the study highlights holistic, iterative methods for identifying fraud and preserving authentic consumer voices.
Conjoint analysis often begins with static feature glossaries, but are respondents really reading them? This research-on-research study compares glossary-based warm-ups with a Kano-inspired warm-up, with and without the Kurz-Binner Priming questions. We test impacts on comprehension, ANA, engagement, LOI, and utilities. Attendees will walk away with evidence-based, practical guidance for designing conjoint warm-ups that improve both the respondent experience and the quality of the data.
Clients often request feature-level willingness to pay (WTP) insights, but additive methods face scaling and feasibility issues. Sawtooth's Sampling of Scenarios (SOS) improves realism through scenario averaging yet still overstates values. SKIM is testing alternative extrapolation functions (quadratic, exponential, and piecewise) to better capture price responses, with preliminary results showing a 16% WTP reduction using quadratic extrapolation versus linear. Future work also includes broader scaling approaches comparing "all-features-on" versus "all-features-off" simulations.
Automated Machine Learning (AutoML) identifies and then executes the best algorithm fitting to each specific data set automatically with little or no human intervention. Expanding AutoML to clustering is gaining attention but facing additional hurdles due to the lack of ground-truth for model training. Making it work for market research is even more challenging because of the art part. But the value can be big. This research explores an automated system for market segmentation, shares learnings and novel solutions for future expansion.