Pages

Wednesday, June 4, 2014

The Unavoidable Instability of Brand Image

"It may be that most consumers forget the attribute-based reasons why they chose or rejected the many brands they have considered and instead retain just a summary attitude sufficient to guide choice the next time."
This is how Dolnicar and Rossiter conclude their paper on the low stability of brand-attribute associations. Evidently, we need to be very careful how we ask the brand image question in order to get test-retest agreement over 50%. "Is the Fiat 500 a practical car?" Across all consumers, those that checked "Yes" at time one will have only a 50-50 chance of checking "Yes" again at time two, even when the time interval is only a few weeks. Perhaps, brand-attribute association is not something worth remembering since consumers do not seem to remember all that well.

In the marketplace a brand attitude, such as an overall positive or negative affective response, would be all that a consumer would need in order to know whether to approach or avoid any particular brand when making a purchase decision. If, in addition, a consumer had some way of anticipating how well the brand would perform, then the brand image question could be answered without retrieving any specific factual memories of the brand-attribute association. By returning the consumer to the purchase context, the focus is placed back on the task at hand and what needs to be accomplished. The consumer retrieves from memory what is required to make a purchase. Affect determines orientation, and brand recognition provides performance expectations. Buying does not demand a memory dump. Recall is selective. More importantly, recall is constructive.

For instance, unless I have tried to sell or buy a pre-owned car, I might not know whether a particular automobile has a high resale value. In fact, if you asked me for a dollar value, that number would depend on whether I was buying or selling. The buyer is surprised (as in sticker shock) by how expensive used cars can be, and the seller is disappointed by how little they can get for their prized possession. In such circumstances, when asked if I associate "high resale value" with some car, I cannot answer the factual question because I have no personal knowledge. So I answer a different, but easier, question instead. "Do I believe that the car has high resale value?" Respondents look inward and ask themselves, introspectively, "When I say 'The car has high resale value,' do I believe it to be true?" The box is checked if the answer is "Yes" or a rating is given indicating the strength of my conviction (feelings-as-information theory). Thus, perception is reality because factual knowledge is limited and unavailable.

How might this look in R?

A concrete example might be helpful. The R package plfm includes a data set with 78 respondents who were asked whether or not they associated each of 27 attributes with each of 14 European car models. That is, each respondent filled in the cells of a 14 x 27 table with the rows as cars and the columns as attributes. All the entries are zero or one identifying whether the respondent did (1) or did not (0) believe that the car model could be described with the attribute. By simply summing across the 78 different tables, we produce the aggregate cross-tabulation showing the number of respondents from 0 to 78 associating each attribute with each car model. A correspondence analysis provides a graphic display of such a matrix (see the appendix for all the R code).



Well, this ought to look familiar to anyone working in the automotive industry. Let's work our way around the four quadrants: Quadrant I Sporty, Quadrant II Economical, Quadrant III Family, and Quadrant IV Luxury. Another perspective is to see an economy-luxury dimension running from the upper left to the lower right and a family-sporty dimension moving from the lower left to the upper right (i.e., drawing a large X through the graph).

I have named these quadrants based only on the relative positions of the attributes by interpreting only the distances between the attributes. Now, I will examine the locations of the car models and rely only the distances between the cars. It appears that the economy cars, including the partially hidden Fiat 500, fall into Quadrant II where the Economical attributes also appear. The family cars are in Quadrant III, which is where the Family attributes are located. Where would you be if you were the BMW X5? Respondents would be likely to associate with you the same attributes as the Audi A4 and the Mercedes C-class, so you would find yourself in the cluster formed by these three car models.

Why am I talking in this way? Why don't I just say that the BMW X5 is seen as Powerful and therefore placed near its descriptor? I have presented the joint plot from correspondence analysis, which means that we interpret the inter-attribute distances and the inter-car distances but not the car-attribute distances. It is a long story with many details concerning how distances are scaled (chi-square distances), how the data matrix is decomposed (singular value decomposition), and how the coordinates are calculated. None of this is the focus of this post, but it is so easy to misinterpret a perceptual map that some warning must be issued. A reference providing more detail might be helpful (see Figure 5c).

Using the R code at the end of this post, you will be able to print out the crosstab. Given the space limitation, the attribute profiles for only a few representative car models have been listed below. To make it easier, I have ordered the columns so that the ordering follows the quadrants: the Mazda MX5 is sporty, the Fiat 500 is city focus, the Renault Espace is family oriented, and the BMW X5 is luxurious. When interpreting these frequencies, one needs to remember that it is the relative profile that is being plotted on the correspondence map. That is, two cars with the same pattern of high and low attribute associations would appear near each other even if one received consistently higher mentions. You should check for yourself, but the map seems to capture the relationships between the attributes and the cars in the data table (with the exception of Prius to be discussed next).

Mazda MX5 Fiat 500 Renault Espace BMW X5 VW Golf Toyota Prius
Sporty 65 8 1 47 29 8
Nice design 40 35 17 31 20 9
Attractive 39 40 12 36 33 10
City focus 9 58 5 1 30 26
Agile 22 53 9 15 40 10
Economical 3 49 17 1 29 42
Original 22 37 7 8 5 19
Family Oriented 1 3 74 41 12 39
Practical 6 39 52 23 44 16
Comfortable 12 6 47 46 27 23
Versatile 5 5 39 30 25 21
Luxurious 28 6 10 58 12 11
Powerful 37 1 9 57 20 9
Status symbol 39 12 6 51 23 16
Outdoor 13 1 20 46 6 4
Safe 4 5 23 40 40 19
Workmanship 13 3 4 28 14 19
Exclusive 17 14 3 19 0 8
Reliable 17 11 17 38 58 27
Popular 5 24 27 13 55 10
Sustainable 8 7 18 19 43 29
High trade-in value 4 3 0 36 41 4
Good price-quality ratio 11 20 15 7 30 21
Value for the money 9 7 12 8 24 10
Environmentally friendly 6 32 7 2 20 51
Technically advanced 17 2 6 32 10 46
Green 0 10 2 2 6 36

Now, what about Prius? I have included in the appendix the R code to extract a third dimension and generate a plot showing how this third dimension separates the attributes and the cars. If you run this code, you will discover that the third dimension separates Prius from the other cars. In addition, Green and Environmentally Friendly can be found nearby, along with "Technically Advanced." You can visualize this third dimension by seeing Prius as coming out of the two-dimensional map along with the two attributes. This allows us to maintain the two-dimensional map with Prius "tagged" as not as close to VW Golf as shown (e.g., shadowing the Prius label might add the desired 3D effect).

The Perceptual Map Varies with Objects and Features

What would have happened had Prius not be included in association task? Would the Fiat 500 been seen as more environmentally friendly? The logical response is to be careful about what cars to include in the competitive set. However, the competitive set is seldom the same for all car buyers. For example, two consumers are considering the same minivan, but one is undecided between the minivan and a family sedan and the other is debating between the minivan and a SUV. Does anyone believe that the comparison vehicle, the family sedan or the SUV, will not impact the minivan perceptions? The brand image that I create in order complete a survey is not the brand image that I construct in order to make a purchase. The correspondence map is a spatial representation of this one particular data matrix obtained by recruiting and surveying consumers. It is not the brand image.

As I have outlined in previous work, brand image is not simply a network of association evoked by a name, a package, or a logo. Branding is a way of seeing, or as Douglas Holt describes it, "a perceptual frame structuring product experience." I used the term "affordance" in my earlier post to communicate that brand benefits are perceived directly and immediately as an experience. Thus, brand image is not a completed project, stored always in memory, and waiting to be retrieved to fill in our brand-attribute association matrix. Like preference, brand image is constructed anew to complete the task at hand. The perceptual frame provides the scaffolding, but the specific requirements of each task will have unique impacts and instability is unavoidable.

Even if we attempt to keep everything the same at two points in time, the brand image construction process will amplify minor fluctuations and make it difficult for an individual to reproduce the same set of responses each time. However, none of this may impact the correspondence map for we are mapping aggregate data, which can be relatively stable even with considerable random individual variation. Yet, such instability at the individual level must be disturbing for the marketer who believes that brand image is established and lasting rather than a construction adapting to the needs of the purchase context.

The initial impulse is to save brand image by adding constraints to the measurement task in order to increase stability. But this misses the point. There is no true brand image to be measured. We would be better served by trying to design measurement tasks that mimic how brand image is constructed under the conditions of the specific purchase task we wish to study. The brand image that is erected when forming a consideration set is not the brand image that is assembled when making the final purchase decision. Neither of these will help us understand the role of image in brand extensions. Adaptive behavior is unstable by design.

Appendix with R code:

library(plfm)
data(car)
str(car)
car$freq1
t(car$freq1[c(14,11,7,5,1,4),])
 
library(anacor)
ca<-anacor(car$freq1)
plot(ca, conf=NULL)
 
ca3<-anacor(car$freq1, ndim=3)
plot(ca3, plot.dim=c(1,3), conf=NULL)
Created by Pretty R at inside-R.org

Tuesday, May 27, 2014

Which came first, the preference or the choice?

Obviously, preference precedes choice because choices are made to maximize preference. That is certainly the way we conduct our marketing research. We generate factorial designs and write descriptions full of information about products and services. Our subjects have no alternative but to use the lists of features that we provide them in our choice sets.

All of this can be summarized in the information integration paradigm shown in the figure above. Features are the large S real-world stimuli that become corresponding small s perceptions with their associated utilities inside our heads. The utility function does the integrating and outputs an internal small r response, the total utility of the feature bundle. The capital R represents the answer on the questionnaire because there is some additional translation needed to provide a rating or to select an alternative from a choice set.

One feature (big S) elicits one preference (little s). If one likes the features, then one will like the product, which is nothing more than a bundle of features. So, if you like strawberry jam, but not too sweet and especially not too expensive, then you will prefer Brand X strawberry preserve if it possesses the optimal combination of your preferred features. And you know this from a blind taste test of all the strawberry preserves on our supermarket shelf? Or, do you buy Brand X strawberry preserve because it was what you ate as a child or was a gift from someone important to you or was recommended by a trusted companion? A good deal of your product knowledge comes from observational learning where we watch and copy the purchase choices made by others. Are your choices determined by the features that you prefer, or are you seduced into purchase and infer feature preference from your choices?

Will Choice Blindness Change Your Mindset?

Let's head to the supermarket for a taste test of different flavors of jams and an experimental assessment of which comes first preference or choice. Better yet, we have this YouTube video from a BBC program on decision making. The details of the study have been published, so that all I need to do is describe the following figure.

The woman with the spoon is the respondent tasting two different flavors of jam, one in the blue jar and the other in the red jar. Here is the videotape associated with this picture. What you should note is that the woman controlling the jars is the experimenter. She opens the red jar first, the taster tastes, and the experimenter turns the red jar over. Unknown to the taster, the red and blue jars are identical. Both are double jars that can be opened by unscrewing either the top or the bottom, and both contain the same two jams. What you need to notice is that the experimenter turns the jar over after the first tasting. Consequently, the top of the red jar now has the same flavor jam as the blue jar. The respondent tastes the blue jar, which given the look of disapproval on the taster's face, we will assume has been rejected. Now, the respondent is asked to taste again the flavor that they preferred, which has been switched and replaced with rejected flavor, and  to tell us their reasons for their preference. This video and especially the BBC video illustrate that our tasters have no problem giving reasons why they like the flavor that they just rejected.

In two-thirds of the trials our tasters did not notice that the flavors were switched. They liked the blue on their left or the red on their right, and they were sticking with their choice. When asked to taste again and tell us why, they gave reasons why they preferred the flavor that they originally rejected. If you are questioning if the two jams were similar enough to be mistaken for each other, you can be reassured that in a separate experiment different subjects had no problems discriminating between the two flavors. The first video from the BBC contains clips of respondents explaining their preferences for the flavor that they rejected but were fooled into believing that they preferred. It should be noted that the reasons seem entirely reasonable as if the tasters fell for the trick and took the second tasting as preferred. If this one study does not convince you, you can find many more available from the choice blindness lab.

What are the implications for statistical modeling?

Does the product as feature bundle seem to work in our research because we have removed so much of the actual purchase context and all that the respondents have left is the information we provided in our vignettes and scenarios? For marketing researchers this implies that choice modeling may be of limited usefulness because there are only a few purchase contexts where we can generalize from the laboratory to the marketplace. Otherwise, the preference construction process induced by our hypothetical experiment does not match the preference construction process unfolding in real purchase contexts.

In fact, this is why we have discrete choice modeling. As Daniel McFadden explains in his Nobel Prize acceptance speech, the decision to drive alone, carpool, take the bus or the metro required a new statistical model reflecting the specific preference construction process required to make such a choice. What is true for transportation choices is true for many occasions when consumers decide what to buy. It is not difficult to think of actual purchase situations where we have narrowed our choices down to two or three alternatives that we are actively considering and where we compare these offers by trading off features. Online purchases by sites that enable you to create side-by-side comparisons will satisfy the criteria underlying information integration. We can observe the same phenomena in the store when a shopper takes two containers offer the shelf to compare the ingredients.

The R package bayesm will handle such data and return individual-level estimates for all the utilities, although we still need to be concerned when presenting multiple choice sets to each respondent. Feature importance is determined by the range and frequency with which feature levels are systematically manipulated in an experimental design. Thus, what is not important when presented as a single choice task becomes important when varied over several choice sets. Moreover, there are good reasons to limit the number of features listed for each alternative. We must resist client demands to cram more information into the product description than they would be willing to include on the packaging or in advertising.

Now, what about all the other purchases that we make every day that do not involve feature comparisons? A possible answer is offered by John Hauser in his paper on the role of recognition-based heuristics in marketing science. Recognition is one of many heuristics from the Fast and Frugal Paradigm, which holds that simple processes can yield good results with little cognitive effort when the environment is structured appropriately. For example, larger brands with more customers tend to be more easily recognized by everyone. If the greater market share results from the brand's ability to satisfy more customers, then the market is structured so that my recognition of the brand is an indication that I am likely to be satisfied after buying the brand. This is an example of choice without feature preference.

Sell the Sizzle Not the Steak

What is the basis for competition in the product category? Is it feature comparison (steak type, grade, degree of marbling, maturity, fat content, origin), or is the focus on benefit (happy steak consumption experiences with loved ones)? Often, the market is engaged in a battle to frame the purchase in term of benefits or features with the market leader emphasizing benefits and competitors struggling to gain share by pushing feature or price advantages. Statistical modeling plays on this same battlefield when we collect and analyze choice data by varying feature levels. Conjoint analysis frames the purchase task and encourages the impact of features. Who is surprised to discover strong price effects when prices are varied repeatedly over a sizeable range of values?

Branding, on the other hand, draws our attention away from features and toward benefits. It invites different data collection procedures and alternative statistical models. Treating the brand as just another feature in an experimental design diminishes the role that brand plays in the market. In the actual purchase context brand dominates as recognizable and familiar, as an intentional agent promising benefits, or as I argued in an earlier post, as an affordance. Consequently, when we measure brand perceptions, we consistently find a highly correlated set of responses: a strong first principal component (linear) or a clear low-dimensional manifold (nonlinear). This is the brand schemata, a pattern of strengths and weaknesses, that I have analyzed using item response theory.

Focusing on the brand takes us down an analytic path toward pattern recognition and matrix decomposition. It moves us out of the econometric task view on the CRAN into machine learning. We begin looking at linear and nonlinear dimension reduction as statistical models of how the brand holds it all together. Recommender systems acquire a special appeal. Moreover, experimental designs begin to seem too obstructive, and we are drawn toward more naturalistic data collection. If consumers rely on brand information to assist them in their decision journey, then we must be careful not to remove those supports and create a self-fulfilling prophecy where feature preferences dominate because feature comparisons are all that is available.

Monday, May 19, 2014

The Purchase Funnel Survives the Consumer Decision Journey

The journey metaphor is almost irresistible. All one needs is a starting point and a finish line, plus some notion of progression. Thus, life is a journey, and so is love. Why not apply the metaphor to your next purchase? McKinsey & Company takes such a metaphoric leap in their very popular paper on the consumer decision journey. They start with a purchase trigger and move through four primary phases that they see as four potential battlegrounds:  initial consideration, active evaluation, closure (purchase), and post-purchase. If you are thinking that this seems to be a variation of awareness-interest-desire-action (AIDA) or the purchase funnel, as it is also known, you would not be wrong. In fact, McKinsey & Company begin with the purchase funnel but reject its claims that consumers move sequentially through invariant stages with a narrowing of the number of brands considered until only one victor remains.

According to the consumer decision journey, the pathway is no longer linear or invariant, nor is the momentum consistently forward. The notion of progression remains, but progress can be piecemeal with one step forward followed by two steps back. The internet makes a lot more information available and transfers control of marketing activities from the brand to the consumer. Consumers are in command and can seek information as they wish from any source and in any order.

Yet, the journey metaphor maintains the goal of narrowing options and selecting what is best for each consumer. The prize for the brand remains loyalty as indicated by continued purchase and advocacy through recommendation and positive word of mouth. Brand awareness still matters, as does customer satisfaction and support services. Competitive pricing and new offerings have not lost their ability to steal away customers. The outcome that we seek to maximize has not changed whether you call it brand equity, attachment, involvement, engagement or loyalty. Moreover, that outcome retains its status as a latent variable that can be observed only through its effects on consumer responses to the brand, such as awareness, interest, desire and action. In fact, as Joakim Nilsson shows in the diagram below, this latent dimension has an impact after the purchase with social media enlarging the funnel and amplifying the reach of each satisfied or dissatisfied customer. By beginning and ending the process with impressions, one gets the sense that the consumer decision journey evolves over time as more customers join and usage diversifies.



So why is the purchase funnel dead? Are consumers considering products with which they have no awareness? Are they purchasing products without considering them first? If a sizable percentage of your customers reported that they have recommended your brand to others, cannot we conclude that your brand has made it to the end of the decision-making process with a more or less stable customer base? On the other hand, would you not agree that your brand was stuck in the starting blocks if most consumers had a difficult time naming your brand when asked about the product category?

When one wants to intervene in the purchase process to improve its brand's standing, it is important to know the different paths that consumers are taking to learn about the brand. However, improving your brand's status means that you also need to know how well the brand is doing, which is what the purchase funnel provides. Brands that fail to achieve awareness among prospective customers will not succeed, so we measure brand achievement by asking customers about their brand awareness. Brand awareness is good, but consideration is better, and purchase is best. The percentage of prospects lost as we move from awareness to consideration to purchase would seem to be a good indicator of brand value. We assess the brand by measuring consumers. Therefore, the purchase funnel is not dead, at least not as an instrument for brand appraisal.

Thinking Like an Item Response Theorist

An item response theorist sees people lined up, single file, one after another with each person possessing more of some property than the person before but less of that property than the next person in line. "Alright, everyone in line with the shortest person in front and the tallest at the end." But what if the property were not visible, such as one's location along the consumer decision journey? How would the item response theorist know which consumer was farther along and which had just started the trip? Would they not write items for an achievement test?

We already have some idea of what to measure based on the battlegrounds identified in the McKinsey & Company article. Successful brands are those that come to mind when making a purchase and are able to maintain a loyal customer base. The item response theorist would place a "sensor" at each of these locations:  one item measuring awareness and another item measuring recommendation. If a brand activated the awareness sensor but not the recommendation sensor, we would know where the brand fell along the path. In fact, we would want to put a "sensor" at each gate leading from one stage to the next (e.g., attention, positive image, consideration, purchase, and recommendation). These stages are achievement milestones along the consumer decision journey.

Brand equity is found in the value that consumers see in the brand. Thus, the brand appraisal process proceeds by asking for consumer ratings. The items assess the final achievement for the brand and not the consumer learning process. We measure if the brand is considered, but not whether that consideration results from an advertisement or a showroom visit or an online review or a YouTube video. The first step is brand appraisal. The next step is tracing the journey or at least assessing the impact of touchpoints on brand standing.

Item response theory (IRT) will guide us in the brand appraisal process. We begin looking for an underlying continuum, a latent variable that will account for our observations. Brands attract consumers. The strength of this attraction is what we wish to measure. Small levels of attraction appear initially in awareness by getting the brand noticed. Does the consumer follow-up and get acquainted? Stronger attraction places the brand into the consideration set, and the strongest brands pull us even closer to them so that we continually purchase more and more frequently. At the highest levels of attraction, we become advocates and begin spreading the word. The black hole pushes the analogy too far, but the image ought to make the point memorable.



On the other hand, if you would like a more formal justification for this hierarchy, you can find it in the literature on consumer-based brand equity (the CBBE model). Simply replace the purchase funnel with the brand resonance pyramid since a funnel is nothing more than an inverted pyramid.

I have provided the code in previous posts showing how the r package ltm will perform the analysis for checklists or rating scales. To learn about item response theory, it is best to start with binary items (e.g., yes/no or present/absent). This is because ratings follow the same logic as checklists by treating the rating as if it were an ordered sequence of checklists. While a binary item divides brand attraction into two parts, a rating scale partitions the same continuum into the number of values on the scale. In both cases an item response model will give us some number of cutpoints for each observed indicator of the underlying latent variable, either one cutpoint for a checklist or the number of scale values minus one ordered cutpoints for rating scales. The same data generating process is responsible regardless of the type of scale.

The brand attracts consumers with milestones indicating the strength of that attraction. The item response model calculates the position of these transition points from inattention to interest to consideration to purchase to satisfaction to retention to advocacy. Then, it positions each respondent along that same continuum showing the strength of the brand's attraction for that respondent. The distribution of consumers for each brand is a measure of the brand's total attraction yielding not only a central tendency but also showing the nature and shape of the heterogeneity among respondents.

It's Alive!

The purchase funnel has survived as a reliable tool for measuring brand strength. Yet, I have offered no guidance for how to assess the impact of the consumer decision journey. The number of predictors explodes when the consumer is in charge. There are hundreds of possible touchpoints where prospective customers can learn about the brand. This figure is but a start.


Yet, each individual is likely to have only limited contact with only a subset of all possible touchpoints, yielding a high-dimensional predictor space that is quite sparse (not unlike what we see with recommendation systems like Netflix). Moreover, if we take the journey concept seriously, then touchpoint effects depend on where we are in purchase process. It's a continuum in the sense that first comes awareness and then comes consideration followed by purchase, but what gets a prospect from awareness to consideration is not the same as what gets them from consideration to purchase. All this will need to wait for a later post.

Thursday, May 15, 2014

The Mind Is Flat! So Stop Overfitting Choice Models


Conjoint analysis and choice modeling rely on repeated observations from the same individuals across many different scenarios where the features have been systematically manipulated in order to estimate the impact of varying each feature. We believe that what we are measuring has substance and existence independent of the measurement process. Nick Chater, my source for this eerie figure depicting the nature of self-perception, lays to rest this "illusion of depth" in a short video called "The Mind is Flat." We do not possess the cognitive machinery demanded by utility theory. When we "make up our mind," we are literally making it up. Features do not possess value independent of the decision context. Features acquire value as we attempt to choose one option from many alternatives. Consequently, whatever consistency we observe results from reoccurring situations that constrain preference construction and not because of some seemingly endless store of utilities buried deep in our memories.

Although it is convenient for the product manager to think of their offerings as bundles of features and services, the consumer finds such a representation to be overwhelming. As a result, the choice modeler is forced to limit how much each respondent is shown. The conflict in choice modeling is between the product manager who wants to add more and more features to the bundle and the analyst who needs to reduce task complexity so that respondents will participate in the research. At times, fractional factorial designs fail to remove enough choice sets, so we turn to optimal configurations with acceptable confounding (see design of experiments in R). Still, even our reduced number of choice scenarios may be too many for any one individual, so we show only a few scenarios to each respondent, make a few restrictive assumptions about homogeneity (e.g., introduce hyperparameters specifying the relationships between individual- and group-level parameters), and then proceed with hierarchical Bayes to compute separate estimates for every person in the study.

We justify such data collection by arguing that it is an "as-if" measurement model. Of course, people cannot retain in memory the utilities associated with every possible feature or service level. Clearly, no one is capable of the mental arithmetic necessary to do the computation in their heads. Yet, we rationalize the introduction of such unrealistic assumption claiming that they allow us to learn what drives choice and decision making. Thus, by asking a consumer to decide among feature bundles using only the information provided by the experimenter, one can fit a model and estimate parameters that will predict behavior in this specific setting. But our findings will not generalize to the marketplace because we are overfitting. The estimated utilities work only for this one particular task. What we have learned from behavioral economics over the last 30 years is that what is valued depends on the details of the decision context.

For those of you wishing a more complete discussion of these issues, I will refer you to my previous posts on Context Matters When Modeling Human Judgment and Choice, Got Data from People?, and Incorporating Preference Construction into the Choice Modeling Process.

Ecological Data Collection and Latent Variable Modeling

I am not suggesting that we abandon choice modeling or hierarchical Bayes estimation. A well-designed choice study that carefully mimics the actual purchase context can reveal a good deal about the impact of varying a small number of features and services. However, if our concern is learning what will happen in the marketplace when the product is sold, we ought to be cautious. Order and context effects will introduce noise and limit generalizability. Multinomial logistic models, such as those in the bayesm R package, teach us that feature importance depends on the range of feature levels and the configuration of all the other features varied across the choice scenarios. We are no longer in the linear world of rating-based conjoint via multiple regression with its pie charts indicating the proportional contribution of each feature.

A good rule of thumb might be to include no more features than the number that would be shown on the product package or in a print ad. Our client's desire to estimate every possible aspect will only introduce noise and result in overfitting. On the other hand, simply restricting the number of features will not eliminate order effects. Whenever we present more than one choice scenario, we need to question whether our experimental arrangements have induced selection strategies that would not be present in the marketplace. Does varying price focus attention on price? Does the inclusion of one preferred feature level create a contrast effect and lower the appeal of the other feature levels? These effects are what we mean when we say the preference is not retrieved from a stable store kept in long-term memory.

It is unnecessary for consumers to store utilities because they can generate them on the fly given the choice task. "What do you feel like eating?" becomes a much easier question when you have a menu in your hands. We use the choice structure to simplify our task. I read down the menu imaging how each item might taste and select the most appealing one. I switch products or providers by comparing what I am using with the new offer. The important features are the ones that differentiate the two alternatives. If the task is difficult or I am not sure, then I keep what I have and preserve the status quo. In both cases context comes to our rescue.

The flexibility that characterizes human judgment and decision making flows from our willingness to adapt to the situation. That willingness, however, is not a free choice. We are not capable of storing, retrieving and integrating feature level utilities. You might remember the telephone game where one person writes down a message and whispers it to a second person, who whispers the message they heard to a third, and so on. Everyone laughs at the end of the telephone line when the last person repeats what they think they had heard and it is compared to what was written. Such is the nature of human memory.

We can avoid overfitting by reducing error and simplifying our statistical models. These are the two goals of statistical learning theory. We keep the choice task realistic and avoid order effects. Occam's razor will trim our latent variables down to a single continuous dimension or a handful of latent classes. For example, the offerings within a product category are structured along a continuum from basic to premium. The consumer learns what is available and decides where they personally fall along this same continuum. Do they get everything they want and need from the lower end, or is it worth it to them to pay more for the extras? The consumer exploits the structure of the purchase context in order to simplify their purchase decision. If our choice modeling removes those supports, it no longer reflects the marketplace.

Choice remains complex, but now the complexity lies in the recognition phase. That is, choice begins with problem recognition (e.g., I need to retrieve email away from my desktop or I want to watch movies on the go or both at the same time). Framing of the choice problem determines the ad hoc or goal-derived category, which in turn shapes the consideration set (e.g., smartphones only, tablets only, laptops only, or some combination of the three product categories) and determines the evaluative criteria to be used in this particular situation. This is why I called this section ecological data collection. It is the approach that Donald Norman promotes when designing products for people. For the choice modeler, it mean a shift in our statistical modeling from estimating feature-level utilities to pattern recognition and unsupervised learning.


Friday, May 9, 2014

Customer Satisfaction and Loyalty: Structural Equation Model or One-Dimensional Dissonance

Causal thinking is seductive. Product experience comes first, then feelings of satisfaction, and finally intentions to continue as a customer. Although customer satisfaction and loyalty data tend to be collected all at one time within the same questionnaire, who does not see the work of the invisible hand of causation? We call product and service ratings "drivers" of satisfaction because it is so easy to imagine experience impacting affect and intention. Thus, no one will be surprised to see the R package semPLS (Structural Equation Modeling using Partial Least Squares) using the customer satisfaction model as its example.

However, what if we were to ignore the causal model and look only at the data? The mobi dataset from the semPLS package is a data matrix with 250 rows containing ratings on a scale from 1 to 10 across 24 items measuring mobile phone customer satisfaction. All you need to know in order to run the SEM and interpret the results can be found in the above link to the well-written Journal of Statistical Software article. I, on the contrary, will pretend that I have no causal model and treat the data set as any other battery of ratings. Let us start by exploring the intercorrelations among these variables.

A principal component analysis for this 250 x 24 data matrix yields a first principal component accounting for 39.5% of the total variation and a second principal component that is only one-sixth the size of the first with 6.6%. The size of the first principal component tells us a great deal about the amount of redundancy in the data matrix. The biplot will provide the visualization.

A biplot is a graphic display showing the 24 variables as vectors and the 250 observations as points projected onto the two-dimensional principal component space. That is, we can create a two-dimensional map with the first principal component as the x-axis and the second principal component as the y-axis. We calculate the two principal component scores for every respondent and use points to represent each row. We know the correlation between each variable and the principal components (i.e., the factor loadings), and we can use that knowledge to project the variables as lines or vectors onto the principal component space. As you might recall, the higher the correlation between two variables, the smaller the angle between the lines representing these variables.
This plot was created using the BiplotGUI R package. It shows the distribution of the 250 rows across the first two principal components as red boxes. The lines are the 24 ratings with tick marks indicating the scores from 1 to 10. Only abbreviations are provided, but you can find the actual questions in the semPLS documentation. With the exception of one variable, all the ratings point in the same direction toward the right. Given the small angles between all these lines, one would expect the correlation matrix to be filled with positive correlations. The "L" in the CUSL# indicates loyalty. The second loyalty measure (would you switch providers for price discount) seems to go its own way.

You may have noticed an arc of non-red circles. I picked one of the points toward the right side to show the arc of predicted values for one respondent with a high first principal component score. The same way that you would drop a perpendicular from a point to the x- and y-axes in order to determine the point's location on those dimensions, one can read out the predicted rating by the perpendicular projection of a point onto that variable's line. These are the dotted lines shown by BiplotGUI. One can clearly see that a high first principal component score results in uniformly high ratings across all 24 ratings. This predicted rating would had been the actual rating had the first two principal components accounted for 100% of the variation.

The "GUI" in BiplotGUI indicates that the function opens its own window that allows you to interact with the biplot. This interaction will enable you to experience the relationship between a point's position in the two-dimensional principal component space and its scores on the 24 variables. The "arc" of varying colored circle tells us that those respondents toward the positive end of the first dimensions tended to rate everything higher.
The above biplot is identical to the first biplot except for the location of the arc, which is moved toward the mean. It illustrates the effect of a decreasing first principal component score. The first two principal components account for less than half of the total variation (39.5% + 6.6%), so there will be some discrepancy between our predicted and the actual ratings. Although it is not shown here, the BiplotGUI window provides a frame where you can see the actual and predicted ratings as you move the cursor across the biplot.

Let me show one more arc from the other side of the first principal component. Now the arc is located toward the lower end of the x-axis and shows how these respondents give uniformly lower ratings.
I would encourage you to install BiplotGUI and semPLS, copy the few lines of R code at the end of this post, and interact with the biplot by moving your cursor across the space. As I note in the R code, you will need to active "Predict points closest to cursor position" by right clicking on the biplot display. By selecting the Prediction tab in the adjacent frame, you will be able to see the actual and predicted ratings for each respondent. Seeing the consistency with which all the ratings move together as different respondents are selected might just change your mindset.

Do I really need all these separate latent constructs with their causal connections? Does Occam's Razor shred the structural equation model? True, the network display of a causal model, as shown below, does provide an organization for the 24 ratings. Conceptually, we can distinguish between performance, satisfaction, and loyalty. Obviously, over time, experience comes first, followed by feelings of satisfaction and then loyalty intentions. The directed graph makes causation a compelling inference.
But the ratings are all that I have, and those ratings were all collected on a single occasion. A longitudinal study may be able to separate performance, satisfaction, and loyalty. By the time I make by measurements, the directed arrows have become feedback loops. Thus, my evaluation of my brand's performance depends on whether or not a better alternative is available or how much effort is needed to switch providers. Sometimes it is convenient to believe that all companies are the same. On the other hand, once I decide to switch, all my ratings fall, that is, until I investigate further and discover problems with my "new" provider. I am not arguing that all the ratings will be identical. Every provider will have strengths and weaknesses. But mean-level item differences are not separate latent variables. As long as the ratings move together as a cohesive unit, we have one latent dimension.

To be clear, I am not claiming that unresolved problems will not impact satisfaction and retention. However, the key word is "unresolved" and the frequency of such problems tend to be low among current customers. The unhappy churn unless there are barriers to exit. Our data matrix is a mixture of respondents: a few "hostages" waiting to be freed, some "shoppers" looking for a better deal, a plurality of "inerts" who prefer not to think about it, and a brand-specific percentage of "advocates" making recommendations and active on social media. I have ordered these four components of our mixture model as they might appear along the first principal component. They are not well-separated, which is why a dimensional representation works so well (see this previous post for a more complete discussion of this issue).

In the end, perhaps it is cognitive dissonance, and not cause-and-effect, that binds the ratings together? Attitudes serve behavior. Do I switch or continue using or simply not think about it at all? Do I become an advocate for the brand? Each of these alternative courses of action result from a complex interplay of product experience with each customer's usage situation and personal needs. We cannot simply assume that whether or not one recommends the brand to others depends solely on the brand and not because there are personal gains and losses associated with the recommendation process that have nothing to do with the brand. Complex systems of thought and action can be described but not by causal models.

R code to run analysis in this post:

#load semPLS and datasets
library(semPLS)
data(mobi)
data(ECSImobi)
 
#runs PLS SEM
ecsi <- sempls(model=ECSImobi, data=mobi, E="C")
ecsi
 
#calculate percent variation
(prcomp(scale(mobi))$sdev^2)/24
 
#load and open BiPlotGUI
library(BiplotGUI)
Biplots(mobi, PointLabels=NULL)
#right click on biplot
#select "Predict points closest to cursor positions"
#Select Prediction Tab in top-right frame
Created by Pretty R at inside-R.org

Sunday, March 23, 2014

Warning: Clusters May Appear More Separated in Textbooks than in Practice

Clustering is the search for discontinuity achieved by sorting all the similar entities into the same piles and thus maximizing the separation between different piles. The latent class assumption makes the process explicit. What is the source of variation among the objects? An unseen categorical variable is responsible. Heterogeneity arises because entities come in different types. We seem to prefer mutually exclusive types (either A or B), but will settle for probabilities of cluster membership when forced by the data (a little bit A but more B-like). Actually, we are more likely to acknowledge that our clusters overlap early on and then forget because it is so easy to see type as the root cause of all variation.

I am asking the reader to recognize that statistical analysis and its interpretation extend over time. If there is variability in our data, a cluster analysis will yield partitions. Given a partitioning, a data analyst will magnify those differences by focusing on contrastive comparisons and assigning evocative names. Once we have names, especially if those names have high imagery, can we be blamed for the reification of minor distinctions? How can one resist segments from Nielsen PRIZM with names like "Shotguns and Pickups" and "Upper Crust"? Yet, are "Big City Blues" and "Low-Rise Living" really separate clusters or simply variations on a common set of dwelling constraints?

Taking our lessons seriously, we expect to see the well-separated clusters displayed in textbooks and papers. However, our expectations may be better formed than our clusters. We find heterogeneity, but those differences are not clumping or distinct concentrations. Our data clouds can be parceled into regions, although those parcels run into one another and are not separated by gaps. So we name the regions and pretend that we have assigned names to types or kinds of different entities with properties that control behavior over space and time. That is, we have constructed an ontology specifying categories to which we have given real explanatory powers.

Consider the following scatterplot from the introductory vignette in the R package mclust. You can find all the R code needed to produce these figures at the end of this post.

This is the Old Faithful geyser data from the "datasets" R package showing the waiting time in minutes between successive eruptions on the y-axis and the duration of the eruptions along the x-axis. It is worth your time to get familiar with Old Faithful because it is one of those datasets that gets analyzed over and over again using many different programs. There seems to be two concentrations of points: shorter eruption that occurs more quickly and longer eruptions that have a longer waiting period. If we told the Mclust function from the mclust package that the scatterplot contains observations from G=2 groups, the function would produce a classification plot that looked something like this: 

The red and the blue with their respective ellipses are the two normal densities that are getting mixed. It is such a straightforward example of finite mixture or latent class models (as these models are also called by analysts in other fields of research). If we discovered that there were two pools feeding the geyser, we could write a compelling narrative tying all the data together.

The mclust vignette or manual is comprehensive but not overly difficult. If you prefer a lecture, there is no better introduction to finite mixture than the MathematicalMonk YouTube video. The key to understanding finite mixture models is recognizing that the underlying latent variable responsible for the observed data is categorical, a latent class which we do not observe, but which explains the location and shape of the data points. Do you have a cold or the flu? Without medical tests, all we can observe are your symptoms. If we filled a room with coughing, sneezing, achy and feverish people, we would find a mixture of cold and flu with differing numbers of each type.

This appears straightforward, but for the general case, how does decide how many categories, in what proportions, and with what means and covariance matrices? That is, those two ellipses in the above figure are drawn using a vector with means for the x- and y-axes plus a 2x2 covariance matrix. The means move the ellipse over the space, and the covariance matrix changes the shape and orientation of the ellipse. A good summary of the possible forms is given in Table 1 of the mlcust vignette. 

Unfortunately, the mathematics of the EM algorithm used to solve this problem gets complicated quickly. Fortunately, Chris Bishop provides an intuitive introduction in a 2004 video lecture. Starting at 44:14 of Part 4, you will find a step-by-step description of how the EM algorithm works with the Old Faithful data. Moreover, in Chapter 9 of his book, Bishop cleverly compares the workings of the EM and k-means algorithms, leaving us with a better understanding of both techniques.

If only my data showed such clear discontinuity and could be tied to such a convincing narrative.

Product Categories Are Structured around the Good, the Better, and the Best

Almost every product category offers a range of alternatives that can be described as good, better, and best. Such offerings reflect the trade-offs that customers are willing to make between the quality of the products they want and the amount that they are ready to spend. High-end customers demand the best and will pay for it. On the other end, one finds customers with fewer needs and smaller budgets accepting less and paying less. Clearly, we have heterogeneity, but are those differences continuous or discrete? Can we tell by looking at the data?

Unlike the Old Faithful data with well-separated grouping of data points, product quality-price trade-offs look more like the following plot of 300 consumers indicating how much they are willing to spend and what they expect to get for their money (i.e., product quality is a composite index combining desired features and services).

There is a strong positive relationship between demand for quality and willingness to pay, so a product manager might well decide that there was opportunity for at least a high-end and a low-end option. However, there is no natural breaks in the scatterplot. Thus, if this data cloud is a mixture of distinct distributions, then these distributions must be overlapping. 

Another example might help. As shown by John Cook, the distribution of heights among adults is a mixture two overlapping normal distribution, one for men and another for women. Yet, as you can observe from Cook's plots, the mixture of men's and women's height does not appear bimodal because the separation between the two distributions is not large enough. If you follow the links in Cook's post, eventually you will find the paper "Is Human Height Bimodal?", which clearly demonstrates that many mixtures of distributions appear to be homogeneous. We simply cannot tell that they are mixture by looking just at the shape of the distribution for the combined data. The Old Faithful data with its well-separated bimodal curve provides a nice contrast, especially when we focus only on waiting time as a single dimension (Fitting Mixture Models with the R Package mixtools).

Perhaps then, segmentation does not require gaps in the data cloud. As Wedel and Kamakura note, "Segments are not homogeneous groupings of customers naturally occurring in the marketplace, but are determined by the marketing manager's strategic view of the market." That is, one could look at the above scatterplot, see the presence of a strong first principal component from the closest of the data points to the principal axis of the ellipse and argue that customer heterogeneity is a single continuous dimension running from the low- to the high-end of the product spectrum. Or, one could look at the same scatterplot and see three overlapping segments seeking basic, value and premium products (good, better, best). Let's run mclust and learn if we can find our three segments.


When instructed to look for three clusters, the Mclust function returned the above result. The 300 observed points are represented as a mixture of three distributions falling along the principal axis of the larger ellipse formed by all the observations. However, if I had not specified three clusters and asked Mclust to use its default BIC criterion to select the number of segments, I would have been told that there was no compelling evidence for more than one homogeneous group. Without any prior specification Mclust would have returned a single homogeneous distribution, although as you can see from the R code below, my 300 observations were a mixture of three equal size distributions falling along the principal axis and separated by one standard deviation.

Number of Segments = Number Yielding Value to Product Management

Market segmentation lies somewhere between mass marketing and individual customization. When mass marketing fails because customers have different preferences or needs and customization is too costly or difficult, the compromise is segmentation. We do not need "natural" grouping, but just enough coherence for customers to be satisfied by the same offering. Feet come in many shapes and sizes. The shoe manufacturer can get along with three sizes of sandals but not three sizes of dress shoes. It is not the foot that is changing, but the demands of the customer. Thus, even if segments are no more than convenient fictions, they can be useful from the manager's perspective.

My warning still holds. Reification can be dangerous. These segments are meaningful only within the context of the marketing problem created by trying to satisfy everyone with products and services that yield maximum profit. Some segmentations may return clusters that are well-separated and represent groups with qualitatively different needs and purchase processes. Many of these are obvious and define different markets. If you don't have a dog, you don't buy dog food. However, when segmentation seeks to identify those who feel that their dog is a member of the family, we will find overlapping clusters that we treat differently not because we have revealed the true underlying typology, but because it is in our interest. Don't be fooled into believing that our creative segment names reveal the true workings of the consumer mind.

Finally, what is true for marketing segmentation is true for all of cluster analysis. "Clustering: Science or Art?" (a 2009 NIPS workshop) raises many of these same issues for cluster analysis in general. Videos of this workshop are available at Videolectures. Unlike supervised learning with its clear criterion for success and failure, clustering depends on users of the findings to tell us if the solution is good or bad, helpful or not. On the one hand, this seems to make everything more difficult. On the other hand, it frees us to be more open to alternative methods for describing heterogeneity as it is now and how it evolves over time.  

We seek to understand the dynamic structure of diversity, which only sometimes takes the form of cohesive clusters separated by gaps. Other times, a model with only continuous latent variables seems to be the best choice (e.g., brand perceptions). And, not unexpectedly, there are situations where heterogeneity cannot be explained without both categorical and continuous latent variables (e.g., two or more segments seeking alternative benefit profiles with varying intensities).

Yet, even these three combinations cannot adequately account for all the forms of diversity we find in consumer data.  Innovation might generate a structure appearing more like long arrays or streams of points seemingly pulled toward the periphery by an archetypal ideal or aspirational goal. And if the coordinate space of k-means and mixture models becomes too limiting, we can replace it with pairwise dissimilarity and graphical clustering techniques, such as affinity propagation or spectral clustering. Nor should we be wedded to the stability of our segment solution when those segments were created by dynamic forces that continue to act and alter its structure. Our models ought to be as diverse as the objects we are studying.

R code for all figures and analysis

#attach faithful data set
data(faithful)
plot(faithful, pch="+")
 
#run mclust on faithful data
require(mclust)
faithfulMclust<-Mclust(faithful, G=2)
summary(faithfulMclust, parameters=TRUE)
plot(faithfulMclust)
 
#create 3 segment data set
require(MASS)
sigma <- matrix(c(1.0,.6,.6,1.0),2,2)
mean1<-c(-1,-1)
mean2<-c(0,0)
mean3<-c(1,1)
set.seed(3202014)
mydata1<-mvrnorm(n=100, mean1, sigma)
mydata2<-mvrnorm(n=100, mean2, sigma)
mydata3<-mvrnorm(n=100, mean3, sigma)
mydata<-rbind(mydata1,mydata2,mydata3)
colnames(mydata)<-c("Desired Level of Quality",
                    "Willingness to Pay")
plot(mydata, pch="+")
 
#run Mclust with 3 segments
mydataClust<-Mclust(mydata, G=3)
summary(mydataClust, parameters=TRUE)
plot(mydataClust)
 
#let Mclust decide on number of segments
mydataClust<-Mclust(mydata)
summary(mydataClust, parameters=TRUE)

Created by Pretty R at inside-R.org