Search This Blog

24.1.11

Alternative Perceptual Mapping Techniques: Relative Accuracy and Usefulness

Perceptual mapping has been used extensively in marketing. This powerful technique is used in new product design, advertising, retail location, and many other marketing applications where the manager wants to know

1. The basic cognitive dimensions consumers use to evaluate "products" in the category being investigated; and

2. The relative "positions" of present and potential products with respect to those dimensions.

When used correctly perceptual mapping can identify opportunities, enhance creativity, and direct marketing strategy to the areas of investigation most likely to appeal to consumers.

Perceptual mapping has received much attention in the literature. Though varied in scope and application, this attention has been focused on refinements of the techniques, comparison of alternative ways to use the techniques, or application of the techniques to marketing problems.' Few direct comparisons have been made of the three major techniques—similarity scaling, factor analysis, and discriminant analysis. In fact, most of the interest has been in similarity scaling because of the assumption that similarity measures are more accurate measures of perception than direct attribute ratings despite the fact that similarity techniques are more difficult and more expensive to use than factor or discriminant analyses.

In practice, a market researcher has neither the time nor the money to simultaneously apply all three techniques. He/she usually selects one method and uses it to address a particular marketing problem. The market researcher must decide whether the added insight from similarity scaling is worth the added expense in data collection and analysis. Furthermore, if the market researcher selects an attribute-based method such as factor analysis or discriminant analysis he/she wants to know which method is better for perceptual mapping and how such maps compare with those from similarity scaling. To answer these questions, one must compare the alternative mapping techniques.

One way to compare these techniques is theoretically. Each technique has theoretical strengths and weaknesses and the choice of technique depends on how consumers actually react to the alternative measurement tasks.

Another comparison approach is Monte Carlo simulation. This useful investigative tool has been employed by researchers to explore variations in similarity scaling and other techniques.

In perceptual mapping Monte Carlo simulation can compare the ability of various techniques to reproduce a hypothesized perceptual map, but it requires that the researcher assume a basic cognitive structure of the individual. Monte Carlo simulation leaves unanswered the empirical question of whether the analytic technique can adequately describe and predict an actual consumer's cognitive structure.

The comparison procedure we use is practical and is based on theoretical arguments and empirical analyses to identify which procedures yield results most useful for marketing research decisions. If the theoretical arguments are supported, researchers can continue to subject a technique to empirical tests in alternative product categories. In this way, one gains insight about the techniques by learning their strengths and weaknesses. If and when the hypotheses are falsified new theories will emerge.

We chose the following guidelines for the comparison.

1. The marketing research environment should be representative of the way the techniques are used empirically.

2. The sample size and data collection should be large enough to avoid exploiting random occurrences and should have no relative bias in favor of the techniques identified as superior.

3. The use of the techniques should parallel as closely as possible the recommended and common usage.

4. The criteria of evaluation should have managerial and research relevance. To fulfill these criteria, we chose an estimation sample and a saved data sample of 500 consumers each, drawn from residents of Chicago's northern suburbs. The application area is perceptions of the attractiveness of shopping areas in the northern suburbs and the criteria are the ability to predict consumer preference and choice, interpretability of the solutions, and ease of use.

Comparison of Similarity Scaling and Attribute-Based Techniques

A major difference between similarity scaling and the attribute-based techniques is the consumer task from which the perceptual measures are derived. Attribute ratings are more direct measures of perceptions than similarity judgments, but may be incomplete if the set of ratings is not carefully developed. Similarity judgments introduce an intermediate construct (similarity) but the judgments are made with respect to the actual product rather than specific attribute scales. A priori, if the set of attributes is relatively complete there is no theoretical reason to favor one measure over the other.

Another difference is the treatment of variation among consumers. In the attribute-based techniques a common structure is assumed, but the values of individual measures are not restricted. In similarity scaling it is restricted to be at most a stretching of the common measure.

Finally, similarity scaling is limited by the number of products. At least seven or eight are needed for maps in two or three dimensions (Klahr 1969). There are no such restrictions for factor analysis. The restriction for discriminant analysis is the number of products minus one. This argument favors attribute-based techniques if the number of products in a consumer's evoked set is small; it favors neither technique if the number of products is large. In practice, the evoked set averages about three products (Silk and Urban 1978).

On the basis of these arguments, if the attribute set is reasonably complete, attribute-based techniques should provide better measures of consumer perception than similarity scaling.

Comparison of Factor Analysis and Discriminant Analysis

Factor analysis is based on the correlations across consumers and products. Discriminant analysis is limited to dimensions that, on average, distinguish among products. Thus factor analysis should use more attributes than discriminant analysis in the dimensions and therefore produce richer solutions. For example, consider Mercedes Benz and Rolls Royce. Suppose that the true perceptual dimensions are country of origin and reliability and that only reliability affects preference and choice. Suppose that perceptions of country of origin differ among products. Suppose that average perceptions of reliability are the same for both cars but individual perceptions differ among consumers. Discriminant analysis will identify only country of origin. Factor analysis will identify both dimensions.

On the basis of this type of argument, one expects factor analysis to provide a richer perceptual structure than discriminant analysis. It should be able to use more of the attribute ratings and should identify perceptual dimensions that predict preference and choice better than discriminant analysis dimensions.

IMPLICATIONS AND FUTURE RESEARCH

Perceptual mapping is an important marketing research tool used in new product planning, advertising development, product positioning, and many other areas of marketing. Strategies based on perceptual maps have led to increased profits, better market control, and more stable growth. Furthermore, much research is based on implications of market structure as identified by perceptual maps. Because of this interest and use, it is crucial that the best mapping technique available be employed in these applications.

Factor analysis is likely to be superior in categories where:

1. The number of products in the average consumer's evoked set is relatively small (seven stimuli or less).

2. There is variation in the way consumers perceive products in the category.

3. Qualitative research has identified a set of attributes likely to represent the product category.

The presence or absence of these characteristics does not ensure the superiority of one technique, but without evidence to the contrary they can serve as guidelines.

Source: “Alternative Perceptual Mapping Techniques: Relative Accuracy and Usefulness“ By JOHN R. HAUSER and FRANK S. KOPPELMAN *

Submitted By:

Yoshita Malkotia

MBA-Finance

12115

Perceptual Mapping- Know where you stand and where you wanna be (Sandipan Sarkar-12045)


Perceptual mapping is a graphics technique used by asset marketers that attempts to visually display the perceptions of customers or potential customers. Typically the position of a product, product, brand, or company is displayed relative to their competition.
Perceptual maps can have any number of dimensions but the most common is two dimensions. Any more is a challenge to draw and confusing to interpret. In case its multi dimensional then its represented in two dimensions with distance between the objects being adjusted. Displaying consumers’ perceptions of related products is only half the story. Many perceptual maps also display consumers’ ideal points. These points reflect ideal combinations of the two dimensions as seen by a consumer.
A company considering introducing a new product will look for areas with a high density of ideal points. They will also look for areas without competitive rivals. This is best done by placing both the ideal points and the competing products on the same map.
Some maps plot ideal vectors instead of ideal points. The map below, displays various aspirin products as seen on the dimensions of effectiveness and gentleness. It also shows two ideal vectors. The slope of the ideal vector indicates the preferred ratio of the two dimensions by those consumers within that segment. This study indicates there is one segment that is more concerned with effectiveness than harshness, and another segment that is more interested in gentleness than strength.

Perceptual Map of Competing Products with Ideal Vectors
Perceptual maps need not come from a detailed study. There are also intuitive maps (also called judgmental maps or consensus maps) that are created by marketers based on their understanding of their industry. Management uses its best judgment. It is questionable how valuable this type of map is. Often they just give the appearance of credibility to management’s preconceptions.

When detailed marketing research studies are done methodological problems can arise, but at least the information is coming directly from the consumer. There is an assortment of statistical procedures that can be used to convert the raw data collected in a survey into a perceptual map. Preference regression will produce ideal vectors. Multi dimensional scaling will produce either ideal points or competitor positions. Factor analysis, discriminant analysis, cluster analysis, and logit analysis can also be used. 
The data required for perceptual mapping thus comes from rating scales where the subjects of the map, from products to populations, are described on the basis of selected attributes. The validity of the map depends on both the overall set of attributes and the subjects of the study as well as the subset of attributes and subjects evaluated by each respondent.

Most studies suffer from too many attributes. Manufacturers and service providers see hundreds of ways in which their products and services differ—or might differ—from those of their competitors. Often the research analyst is unable to impose the discipline necessary to develop a reasonably short list of attributes. In most studies, it is usually desirable (or necessary) to select a subset of attributes for respondents to rate. This can be accomplished by using one of two approaches:
1. Select a subset of most important attributes. Each respondent rates all attributes on importance. The questionnaire is programmed to select a subset of the important attributes for rating. This may assure more meaningful questionnaires for respondents.
2. Randomly select a subset of attributes. The questionnaire randomly selects a subset of attributes for each respondent. This has the advantage that there will be roughly equal sample sizes for each of the evaluative criteria. The obvious disadvantage is that the respondent task may be less interesting.

Data Analysis and Presentation
Multiple discriminate analysis uses the “F ratio” to determine attribute and product or subject location in the perceptual space. The F ratio is a ratio of the variance between ratings of different products/subjects to the variance of ratings within products/subjects. In an attribute study, these variations among ratings are generally of two types:
1. The differences between products/subjects, revealed in the difference between average ratings for different products.
2. The differences within products, revealed in the differences among respondents’ ratings of the same product.

An attribute would have a higher F ratio either if its product averages were more different from one another, or if there were more agreement among respondents rating the same product. Multiple discriminate analysis finds the optimal weighted combination of all the
Attributes which would produce the highest F ratio of between-product to within product variation. That weighted combination of attributes becomes the first
dimension of the perceptual map.

An example of this is given below, employing attributes describing corporate activities,
such as “Manages operations safely,” Prepared for the future,” “Contributes to balance
of trade,” and “Old fashioned” among others.


This first dimension contains the most information possible for any single perceptual map, in terms of accounting for perceived differences among products/subjects, measured in terms of the variation among ratings of the same product.

Reference: http://en.wikipedia.org/wiki/Perceptual_mapping
                http://www.populus.com/files/Perceptual%20Mapping_f_1.pdf

Interpretation of Cluster Analysis - by Arunava Guha (12016)


On the second day of our course of Business Analytics, we primarily learnt Cluster Analysis, which is a popular classification technique frequently used to analyze market research data which divides the data into groups. Data appears in rows and columns. Rows can then be clustered with respect to columns or columns with respect to rows. For example, clustering techniques can be used to identify demographic or psychographic characteristics of consumers with similar purchasing histories, or to isolate differences between groups of products. Market researchers can then study the individual clusters of consumers or products in more detail in order to maximize results from future marketing strategies.

Cluster analysis is most often used in cases in which it is unknown, prior to the analysis, the number of groups in the data or which observations belong to which groups. Objects associated with a specific cluster should be quite similar and generally clusters should be distinct, i.e. not overlapping. Hierarchical methods, in which clusters are defined according to similarity or dissimilarity measures, remain the most popular method of analysis.

The validity of the conclusions drawn from cluster analysis techniques is sometimes questioned since very different clusters can be formed from the same data depending on how the analysis is performed. This apparent ambiguity results from the decisions made by researchers prior to the actual analysis. Careful attention to these pre-analysis decisions will increase your ability to obtain meaningful results. The first of these issues is which variables should be included in the analysis. If important variables are ignored, future results will be suspect. Secondly, a distance measure must be determined. These distances will be used to form the actual clusters - observations with small distances will be grouped together. Euclidean distance is commonly used, and is the default in many software packages. This option, however, may be a poor choice if variables are measured on different scales. Variables on larger scales will have stronger effects on the clusters formed. In fact, in this situation the clusters may simply represent the variables which are measured on similar scales. Appropriate transformations of selected variables may be needed to eliminate this feature of the data. The city-block distance is another type of distance measurement that can be chosen. It is considered to be more robust, or resistant, against the effect of outliers in the data. A final decision on the proper distance choice will depend, not only on the data, but the objectives of the analysis.

Lastly, the algorithm used to form the clusters must be defined. While there are literally hundreds of such algorithms, selection is effectively limited to the half dozen or so available through whatever software package is being used. Still, the choice among these half dozen algorithms is not always clear, and different choices can again lead to different cluster solutions. The single linkage method is the simplest method of joining clusters and the most commonly used. It is robust to small perturbations in the distances, but a small addition of data can significantly alter the clustering. The distances, as defined above, are computed between each object not in a cluster to the nearest cluster member. The new object is then assigned to the nearest cluster. Other methods commonly used are the complete linkage and the average linkage. Complete linkage distances are computed between non-clustered objects and the farthest cluster member; average linkage distances are computed between non-clustered objects and the mean of the clustered objects.

If changes among the distances and algorithms used can result in changes to cluster membership, how are meaningful results to be obtained? How much confidence can researchers place in future marketing strategies based on these clustering results? The first step to determining the validity of the clusters begins well before the analysis stage. Researchers should fully understand both the data and the research objectives. Exploratory data analysis is easily done in all software packages. Box-plots and histograms can be examined to detect outliers or differences in scale among the variables. Cross-tabs and scatter plots can be examined to look for strong relationships between variables.

Secondly, again prior to analysis, thought should be given to what the conclusions of the study are expected to be. It is rare that those involved with the project will not have a general idea of what the outcome might look like. Previous analysis results, or simply the intuition of a researcher who has been absorbed in similar data for several months, can be used to get a feeling for what the final outcome might look like.

Lastly, it is important that different methods of cluster analysis work better with certain shaped clusters. For example, methods which work well when groupings within the data tend to be spherical may perform poorly if groupings within the data tend to be elliptical. Thus, when performing the analysis, one should examine several types of distance measures and algorithms. Then, the resulting clusters are to be compared if they are showing similar pictures. Differences tend to occur when there are no natural clusters in the data. In the case of different outcomes, it is to be observed if some of the clusters contradicting or if at all the described relationships make sense and are actionable. Answers to the pre-analysis issues may indicate distances and algorithms which are more appropriate. Results from several of the distance or algorithm options may be non-informative.

In cases where several different conclusions are obtained and there is no reason to view any as preferable to the others, the study objectives are to be re-examined. If multiple results answer the questions of interest, the solution is t be kept as simple as possible. In cluster analysis, simplicity usually corresponds to the number of clusters in the solution. Above all, it is important to note that cluster analysis is an exploratory tool and different algorithms may very well detect different patterns in the data, none of which may be "wrong" - simply a different method of "slicing" the data.

Reference:

Predictive Analytics: the Future of Business Intelligence

Introduction

In 1918, four years after he was hired by the DuPont Corporation, electrical engineer F. Donaldson Brown was given the task of untangling the finances of a company in which Du Pont had just invested. (This company was General Motors and DuPont had purchased 23 percent of its stock.)
Brown's work led to the development of a system of planning and control for all operating decisions within a firm - the analytical system which became the dominant form of financial analysis in corporations throughout the world. (The Dupont Analysis)

Fast-forward to 2010, when Ted Plush, a 32-year veteran of DuPont, was given the task of making DuPont "a predictive enterprise" -- using analytics not just to report and analyze the past, but provide the insights that will guide the global organization
strategy going forward.

The market is witnessing an unprecedented shift in business intelligence (BI), largely because of technological innovation and increasing business needs. The latest shift in the BI market is the move from traditional analytics to predictive analytics. Although predictive analytics belongs to the BI family, it is emerging as a distinct new software sector.

Analytical tools enable greater transparency, and can find and analyze past and present trends, as well as the hidden nature of data. However, past and present insight and trend information are not enough to be competitive in business. Business organizations need to know more about the future, and in particular, about future trends, patterns, and customer behavior in order to understand the market better. To meet this demand, many BI vendors developed predictive analytics to forecast future trends in customer behavior, buying patterns, and who is coming into and leaving the market and why.

Traditional analytical tools claim to have a real 360° view of the enterprise or business, but they analyze only historical data—data about what has already happened. Traditional analytics help gain insight for what was right and what went wrong in decision-making. Today’s tools merely provide rear view analysis. However, one cannot change the past, but one can prepare better for the future and decision makers want to see the predictable future, control it, and take actions today to attain tomorrow’s goals.

What is Predictive Analytics?

Predictive analytics are used to determine the probable future outcome of an event or the likelihood of a situation occurring. It is the branch of data mining concerned with the prediction of future probabilities and trends. Predictive analytics is used to automatically analyze large amounts of data with different variables; it includes clustering, decision trees, market basket analysis, regression modeling, neural nets, genetic algorithms, text mining, hypothesis testing, decision analytics, and more.
The core element of predictive analytics is the predictor, a variable that can be measured for an individual or entity to predict future behavior. For example, a credit card company could consider age, income, credit history, other demographics as predictors when issuing a credit card to determine an applicant’s risk factor.
Multiple predictors are combined into a predictive model, which, when subjected to analysis, can be used to forecast future probabilities with an acceptable level of reliability. In predictive modeling, data is collected, a statistical model is formulated, predictions are made, and the model is validated (or revised) as additional data become available.

Predictive analytics combine business knowledge and statistical analytical techniques to apply with business data to achieve insights. These insights help organizations understand how people behave as customers, buyers, sellers, distributors, etc.
Multiple related predictive models can produce good insights to make strategic company decisions, like where to explore new markets, acquisitions, and retentions; find up-selling and cross-selling opportunities; and discovering areas that can improve security and fraud detection. Predictive analytics indicates not only what to do, but also how and when to do it, and to explain what-if scenarios.
A Microscopic and Telescopic View of Your Data

Predictive analytics employs both a microscopic and telescopic view of data allowing organizations to see and analyze the minute details of a business, and to peer into the future. Traditional BI tools cannot accomplish this functionality. Traditional BI tools work with the assumptions one creates, and then will find if the statistical patterns match those assumptions. Predictive analytics go beyond those assumptions to discover previously unknown data; it then looks for patterns and associations anywhere and everywhere between seemingly disparate information.

Predictive Analytics and Data Mining

The future of data mining lies in predictive analytics. However, the terms data mining and data extraction are often confused with each other in the market. Data mining is more than data extraction It is the extraction of hidden predictive information from large databases or data warehouses. Data mining, also known as knowledge-discovery in databases, is the practice of automatically searching large stores of data for patterns. To do this, data mining uses computational techniques from statistics and pattern recognition. On the other hand, data extraction is the process of pulling data from one data source and loading them into a targeted database; for example, it pulls data from source or legacy system and loading data into standard database or data warehouse. Thus the critical difference between the two is data mining looks for patterns in data.

A predictive analytical model is built by data mining tools and techniques. Data mining tools extract data by accessing massive databases and then they process the data with advance algorithms to find hidden patterns and predictive information. Though there is an obvious connection between statistics and data mining, because methodologies used in data mining have originated in fields other than statistics.
Data mining sits at the common borders of several domains, including data base management, artificial intelligence, machine learning, pattern recognition, and data visualization. Common data mining techniques include artificial neural networks, decision trees, genetic algorithms, nearest neighbor method, and rule induction.

Major Predictive Analytics Vendors

Traditional powers such as SAS and SPSS(IBM) still sit atop the predictive analytic market. The predictive analytic and data mining leaders have supported large enterprise customers' advanced analytics needs for many years. All offer mature, high-performance, scalable, flexible, and robust [predictive analytic and data mining] solutions that combine a wide range of statistical algorithms with integrated support for in-database analytics and a broad range of information types
Other players such as KXEN Inc., Oracle Corp., and Portrait Software Fair Isaac Corp. (FICO), Angoss Software Corp., and TIBCO Software Inc. are creditable competitors and field very functional or respected offerings

User Recommendations

Depending on an organization’s needs, some predictive analytics tools will be more relevant than others. Each has its strengths and weakness and can be highly industry-and model-specific—the algorithms and models built for one industry are not applicable to other industries. Financial industries, for example, have different models than what are used in manufacturing and research industries.
Selecting the appropriate predictive analytics tools is not a simple task. The following capabilities must be taken into consideration: algorithm richness, degree of automation, scalability, model portability, web enablement, ease of use, and the capability to access large data sets. The more diversified the business, the more functions and unique models are required. Model portability is important even within different business units in the same company. The scalability of the solution and its ability to handle expanded functionality should also be verified and based on a business’ growth.

Users require extensive training and expertise to use the core functionalities of the predictive analytics solutions, such as identifying data, building the predictive model with right predictors, data mining knowledge to align with business strategy etc. Furthermore, predictive analytics automates model building, but does not automate the integration of business processes and knowledge. Thus expertise and training are required to evaluate the best software relevant to an organization’s unique business model.

If a company has or is willing to attain the expertise required to use predictive analytics it can definitely benefit from the tool. Although most large enterprises use some sort of traditional BI tool or platform, their tools do not provide predictive analytics functionality. Incorporating predictive analytics into an existing BI infrastructure can provide organizations’ a competitive advantage in their industry. Consequently, the integration of BI tools is a key consideration when selecting a predictive analytical tool, as is its integration with key applications such as enterprise resource planning, (ERP), customer resource management (CRM), and supply chain management (SCM) etc. Ultimately, since predictive analytics is currently the only way to analyze and monitor the business trends of the past, present, and future, selecting the right tool can be a key success factor in a company’s BI strategy.

Submitted By
Amiteshwar Singh
Roll No 12067

Compiled from various souces:
http://blogs.hbr.org/events/2010/05/moving-analytics-from-what-hap.html
http://www.slideshare.net/robertdpalmer/hbr-competing-on-analytics
http://searchbusinessanalytics.techtarget.com/news/1506983/Predictive-analytics-software-next-battleground-for-BI-vendors?ShortReg=1&mboxConv=searchBusinessAnalytics_RegActivate_Submit&

Prediction Analytics

Robert Nozick once said, “There is no justifiable prediction about how the hypothesis will hold up in the future; its degree of corroboration simply is a historical statement describing how severely the hypothesis has been tested in the past.”

A finance manager has to predict the future EPS just like a marketing manager has to predict the future revenues. An operations manager has to predict the future inventory requirements just like a hr manager has to predict the future attrition rate and finally a GM has to predict the overall growth of the firm and all of them are responsible to someone or the other. But what if the predictions go wrong?

Prediction modeling is the process by which data is modeled and diagnosed to try to best predict the probability of an outcome. But what is the data that is being referred here? The companies realized the importance of this ‘data’ sometime in 1960s. Since then, the trend of data mining and warehousing has picked up in a big way. Almost all the companies across industries started storing and accumulating the data in the form of past performances, demographics of the society, research papers etc. This data is today being used to predict the future trends. Take the example of the banking industry. They use past information related to credit history, job information, loan history and other parameters about a person to determine his credit worthiness and therefore predict whether he or she will default in the future. This is a classic example of prediction analytics, a topic that has been the talk of the day since the past two days at college during the SPSS workshop.

Apart from predicting future behavior, predictive models also anticipate the consequences of change. According to Victor Holman (Performance Management Expert at Lifecycle Performance Professionals), there are three main types of models associated with Predictive analytics. First being the prediction model, second being the descriptive model and third being the decision model.

Predictive analytics' central building block is the predictor, a single value measured for each customer. For example, 'most recent', which is based on the number of weeks since the customer's last purchase, has higher values for more recent customers. This predictor is usually a reliable campaign response predictor: you will receive more responses from those customers more highly ranked by 'most recent'. That means that if you contact your customers in order of 'most recent' - first, call the most-recent customer; next, call the next-most-recent customer; and so on - you will improve your response rate. For each prediction goal, there are an abundance of predictors that will help rank your customer database. For example, consider a customer's online behavior: Customers who spend less time logged on may be less likely to renew their annual subscription. In this case, retention campaigns can be cost-effectively targeted to customers with a low monthly usage predictor value.

Descriptive models quantify the relationships between data in order to classify customers into groups. While predictive models focus on predicting one customer's behavior, descriptive models identify relationships between several customers or products. Descriptive models do not predict a target value, but focus more on the intrinsic structure, relations, interconnectedness, etc. The example that I mentioned earlier about the banks using past demographic data to predict the creditworthiness of the person, follows a descriptive model.

On the first day of the workshop, we were given a small training on Cluster analysis. It is a descriptive modeling technique that identifies clusters embedded in the data. A cluster is a collection of data objects that are similar in some sense to one another. Another descriptive modeling technique’s mention that I found on Victor Holmon’s blog is the k-means algorithm. K-means algorithm is a distance-based clustering algorithm that partitions the data into a predetermined number of clusters (provided there are enough distinct cases). The k-means algorithm works only with numerical attributes. Distance-based algorithms rely on a distance metric (function) to measure the similarity between data points.

Finally, Decision models describe the relationship between all decision elements and predict the results of decisions, allowing you to try different scenarios, and optimize results. Decision Support Systems and Management Information Systems at hospitals use predictive analysis in the health care industry to determine at risk patients and sometimes to determine which course of action would be best given a multiple array of variables. Rational decision models are based around a cognitive judgment of the pros and cons of various options. It is organized around selecting the most logical and sensible alternative that will have the desired effect. The decisions are normally organized through a detailed analysis of alternatives and a comparative assessment of the advantages of each. Weighted criteria scoring are an example of rational decision models.

To sum up, organizations these days have to spend considerable amount of resources on essential tools of survival like data mining and prediction analytics because the future is highly uncertain.

Submitted by:
Kunal Shah
Roll no. 12029
Finance Batch
SIBM Bangalore

(Certain text has been taken from Victor Holman’s knoll on the same topic)
http://knol.google.com/k/victor-holman/three-basic-predictive-analysis-models/2650srcv52e8t/62#

Exploring marketing ideas with Different Types of Perceptual Maps

Making marketing strategies is a complex process requiring research, judgment and creativity. Perceptual mapping is a powerful tool for exploring data and generating hypotheses. This article discusses three types of perceptual maps: preference, multidimensional scaling (MDS) and correspondence.
Marketing research helps marketers establish an objective (or, simulated/virtual) marketplace to understand their customers and their products, or answering questions like: How do my customers use my product? What are the strengths and weakness of my product relative to my competition? Where does my product fit in the overall market consuming such products? Who are the targeted customers for my product?
Once this structured framework is established and understood, it then becomes a guide and analytic platform for creative strategists to design innovative, targeted strategies (to fill the gaps, or to raise the existing product to a higher ground, etc.).
Perceptual mapping is one of the many techniques used in the analytic steps, and an extremely popular one. Its beauty is in its graphical display: Simpler to interpret than a listing of numerical results, it quickly points to potential relationships, connections, and patterns in the data. Its deficiency is that the graph is only an approximate representation of the real data, because of the amount of data condensation/transformation the procedure requires. Therefore, perceptual mapping should not be used alone to reach any conclusions, and must be accompanied by other mathematical means to verify its findings. In general, perceptual mapping is a powerful tool for exploring data, and for coming up with hypotheses.
There are three ways of producing perceptual maps, although most people are familiar with only one: the MDS map. The three types of maps are produced by three different techniques and have different usages:
1. Preference map
2. Multidimensional scaling (MDS) map
3. Correspondence map
Each map requires a different view of the input data, and the maps are used to study different aspects of the marketing problem.
1. Preference map (for study of consumer preferences) A basic preference map shows consumers’ preferences for a set of products. It is more useful than presenting a table of mean ratings. In a typical preference analysis, consumers are surveyed for their preferences for a set of products. For example, 15 consumers are asked to rate their preferences for 10 U.S.-made cars on a rating from 1 to 10 (1 is the least preferred, 10 is the most preferred). Preference analysis performs a principal component analysis on the rating data, and then plots the first two principal components from the analysis to create an approximate two-dimensional display of the consumer preferences for the 10 cars.
2. Multidimensional scaling map (for analysis of product competitiveness) Multidimensional scaling is a graphic technique for analysing the similarities (or dissimilarities) between products. It is not meant for studying consumer preference, but for analysing competitive positioning of the products in the minds of the consumers. The data: For a multidimensional scaling survey, it would be ideal, but highly impractical, to ask every consumer to rate the degree of similarity (or dissimilarity) between all possible pairs of products, because the number of pairs of products to rate would be too large if there are many products. Alternatively, each consumer is asked to place the products into groups of similar products. Consumers can decide as many or as few groups as they like. Multidimensional scaling performs an initial principal component analysis of the original data, and then improves on the solution iteratively. When the solution can no longer be improved, the procedure stops and produces an optimal two-dimensional map of product distances.
3. Correspondence map (to explore information in any frequency table) Correspondence analysis is an ingenious device to explore the associative relationships and clustering patterns in the frequency data. For example, you can use the correspondence map to examine the association between a categorical variable that identifies a group of customers and another categorical variable that distinguishes your product. It is even equipped to display multiple categorical variables simultaneously (such as in multi-way tables of frequency), each having a large number of levels, although with some sacrifice (i.e., the distances between all points in the plot become meaningless).

Submitted By:
Amritha Shrikumar
( Roll No. -12069)

Application of Clustering in selection of Mutual Funds.


Introduction

One of the reasons investors prefer mutual funds is that the funds allow the investors to diversify across asset classes with limited capital. Also since diversifying over uncorrelated assets is a good way to reduce risk, people watch asset return correlation closely so they can make informed decision in portfolio construction.

Correlation is important not only in identifying diversification opportunities but also in understanding investment strategies' characteristics so that we can make informed decision in fund selection.

In this article the author has used return correlation as a similarity measure to cluster some model portfolios. Expected result is that the return correlations would reveal these portfolios' characteristics not easily readable from stated strategies or stock holdings.

Method

Model portfolios with longer than two-year history were selected from a platform, and it resulted in a set of 30. For each portfolio, monthly returns since inception as a time-series were calculated. The correlation of monthly returns for each portfolio pairs was computed which led to construction of a correlation matrix.

Now hierarchical clustering was used to cluster the portfolios, where the distance measure between each pair of portfolios is given by (1 - return correlation). Note that correlation is the only metric used to measure and cluster the portfolios.

Results

The portfolio cluster tree constructed by hierarchical clustering is presented in the graph below. Each leaf node is a portfolio, with its ID associated. From bottom up, you can see that hierarchical clustering builds up the cluster tree iteratively by grouping together clusters of portfolios into bigger clusters. The very top level represents the biggest cluster, i.e. the cluster of all 30 portfolios. The height of the tree represents the distance between pairs of clusters, calculated using complete linkage.

By eye-balling the tree, one can see that the tightest cluster is the one in the left bottom corner. It's a cluster of nine portfolios, including portfolios 18, 3, 9, 17, 24, 4, 14, 2 and 22. It actually makes a lot of sense, because among them:

- Five portfolios are dividend and growth strategy (portfolio 18, 3, 17, 4 and 14)
- Three portfolios are large-cap strategy (portfolio 9, 24, 22)
- One portfolio is value strategy (portfolio 22)
If you delve into these portfolios' holdings, you will see that they mostly held/hold large-cap and dividend-paying stocks, and their holdings have a fair amount of overlap. This cluster could be categorized as a large-cap-dividend cluster.

If we draw a horizontal line anchored on the large-cap-dividend cluster sub-tree's root, you will see another smaller cluster in the middle with similar tightness. It's the cluster of four including portfolios 11, 12, 13 and 16. Their respective strategies are:
- Portfolios 11 and 12: small+mid-cap strategy
- Portfolio 16: mid-cap strategy
- Portfolio 13: quantitative strategy

Another fairly tight cluster is the one with two portfolios, portfolio 19 (strategic asset allocation) and 20 (tactical asset allocation). They both use ETFs to implement asset allocation strategies. No surprise that they ended up in the same cluster.

On the uncorrelated side, portfolio 30 (healthcare), portfolio 28 (value), portfolio 26 (growth) and portfolio 15 (absolute return) are not correlated with the other portfolios (i.e. they are far away from the others on the tree). It makes sense as well because portfolios 30, 28 and 26 are the only portfolios that ever held short positions and thus less correlated with the other long-only portfolios. The most uncorrelated portfolio, i.e. portfolio 15, is absolute return strategy. This portfolio/manager only traded in and out long positions in gold-related instruments, such as gold mutual funds, ETFs and gold miners. That explains its difference from the other stock-focused portfolios.

Implication

If we only used stated strategies or stock holdings, we wouldn't have been able to recognize the large-cap-dividend cluster, nor the small+mid-cap cluster identified in this analysis. Nor could we have been able to separate the uncorrelated portfolios into different clusters. Correlation, a statistical measure as basic as it is, through the lens of hierarchical clustering, does offer additional insights into portfolios' characteristics.

If one were to invest in these portfolios/managers, one would pick a good manager from the large-cap-dividend cluster, a good manager from the small+mid-cap cluster, and maybe portfolio 15 also (this portfolio is not only uncorrelated with others but also has delivered solid performance consistently). It's interesting and valuable that return correlation plus hierarchical clustering offer all this practical interpretation for manager selection.

References

http://en.wikipedia.org/wiki/Cluster_analysis

http://eng.wealthfront.com/2011/01/cluster-portfolios-using-return.html

Submitted By:
Sujoy Gupta
12167
SIBM-Bangalore

Weather forecast using SPSS Statistical Methods


With a lot of interesting quotes and quotations in the field of statistics a thought that has intrigued my mind always is of our very own Saul Barron

"The weather man is never wrong. Suppose he says that there's an 80% chance of rain. If it rains, the 80% chance came up; if it doesn't, the 20% chance came up!"

From that time I have been wondering whether at all statistics helps to predict weather forecasts/rains. Exploring this I stuck up with an interesting article regarding whether at all SPSS helps us to predict rains.

I went through a paper which presents a case study of using SPSS 13.0 in weather prediction. The data were collected from 2001 till 2005 and it was made a prediction of future temperatures encountered in that region. For this, he used two methods provided by SPSS 13.0, such as factor analysis and linear regression.

Introduction

Weather prediction has always being a matter of debate. Every time the MET department advices us to carry umbrellas we end up having no rain and vice-versa. But nevertheless, weather warnings are an important forecast because they are used to protect life and environment.

Starting the MET dept. data, we can use SPSS software for weather forecast. SPSS uses its advanced mathematical and statistical expertise to extract predictive knowledge. Thus, together with data and this software; we can predict the outcomes before they occur.

Case Study

For this application, the public available data provided by Hong Kong Observatory site was used.

The structure of the database was:

date – the year when the data were registered;

month – the month when the data were registered;

Avg_press –average pressure;

Max_Temp – temperature daily maximum;

Avg_Temp – temperatures average;

Min_temp – temperature daily minimum;

Relative_humidity – average relative humidity;

Avg_cloud – average amount of cloud;

rainfall – total rainfall.

The temperatures from the database were registered in Celsius degrees. Then, the database is completed with the data about weather corresponding from 2001 to 2005

Outside of the primary variables, they transformed average temperatures from Celsius degrees in Fahrenheit degrees. The transformation formula is (where C is Celsius and F is Fahrenheit)

F = 9/5*C + 32

Linear Regression

Regression analysis is a statistical tool that can produce data predictions. The basic principle behind regression is to use one or more variables to predict another variable of interest. The predictor variables are known as independent variables and the variable that is being predicted is known as the dependent variable.

In this case, they considered the regression between the atmospheric pressure and the average temperature. There is a predictor variable (Avg_pressure) and a dependent variable avg_temperature

Using SPSS regression linear method, we obtain regression coefficients, as shown below.

Avg_temp = B+ Beta * Avg_pressure

Now in case the atmospheric pressure is registered being 1017.5, we can estimate the average temperature using the relation = 736.274:0.704*1017.5=19.95 .

The values predicted in this way are estimations, while the correlation between variables is not perfect. The error from estimations is not directly proportional with the correlation between variables (the correlation graph points being more far away from the regression right line).

Factor Analysis

Now this is a reductive method in which new factors are built based on existing relationships between variables. A period of time, factor analysis was used only in psychology. After noticing the good results of this method, it was applied in the economic analysis and it has become an established statistical method.

The main challenge of factor analysis is to find variables that have as much variation (information) "common" as possible, so that after concentration little useful information is lost.

Using SPSS, we can apply this method by accessing the menu:

Statistics – Data Reduction – Factor

and the “Factor Analysis window will appear”.

We can select the variables needed, as well as various characteristics of the variables that may be calculated like mean, standard deviation, correlation matrix, and number of factors to be determined, factor score coefficient matrix. The exclusion of the case list wise and the suppression of the absolute values less than 0.10 can be done as well.

After the proper selection of the needed characteristics, we can press the OK button. Beside the results described above, the method output presents the Total Variance Explained, the Component Matrix and the new variables found.

On Extraction Sums of Squared Loadings columns, there are explained variance and Cumulative variance for two factors in the context of initial factorial solution.

Variant explained by each factor is distributed between different factors.

Ex:-

If factor 1 : 56,606% and factor 2 : 25.359%.

Together, they explain 81.964% of the variation of analyzed values.

The rest up to 100% remain unexplained by this model factorial.

On Rotation Sums of Squared Loadings columns, we have the same values, but after the procedure of rotation. A redistribution of variation explained by each of the items:

Ex:-

If factor 1 – 43.324% and factor 2 – 38.640% has been made.

Together, they explain 81.964% of the variation of analyzed values, but with redistributed weights.So, using a rotation method of the redistribution between the two factors has been made

Conclusions

Thus we have seen an interesting case as to how the MET DEPT. gives us the prediction. The present application can forecast an unknown value, on the basis of some real, known values, using techniques that don’t use too many scientific details. Also the factor analysis is important because it reduces a large number of variables into fewer factors. But the question to ask again is, Are this Statistics a substitute to our judgment??

Submitted By

Naresh Adwani

Roll No: - 12033

SIBM Bangalore

Ref:- http://bulletin-mif.unde.ro/docs/20091/13SCHIOPU_DANIELA.pdf

Hierarchical Agglomerative Clustering

Hierarchical clustering algorithms are either top-down or bottom-up. Bottom-up algorithms treat each document as a singleton cluster at the outset and then successively merge (or agglomerate) pairs of clusters until all clusters have been merged into a single cluster that contains all documents. Bottom-up hierarchical clustering is therefore called hierarchical agglomerative clustering or HAC . Top-down clustering requires a method for splitting a cluster. It proceeds by splitting clusters recursively until individual documents are reached.

An HAC clustering is typically visualized as a dendrogram. Each merge is represented by a horizontal line. The y-coordinate of the horizontal line is the similarity of the two clusters that were merged, where documents are viewed as singleton clusters. We call this similarity the combination similarity of the merged cluster. The combination similarity of a singleton cluster is by moving up from the bottom layer to the top node, a dendrogram allows us to reconstruct the history of merges that resulted in the depicted clustering.

Process

· Assign each object to a separate cluster.

· Evaluate all pair-wise distances between clusters

· Construct a distance matrix using the distance values.

· Look for the pair of clusters with the shortest distance.

· Remove the pair from the matrix and merge them.

· Evaluate all distances from this new cluster to all other clusters, and update the matrix.

· Repeat until the distance matrix is reduced to a single element.

Advantages

· It can produce an ordering of the objects, which may be informative for data display.

· Smaller clusters are generated, which may be helpful for discovery.

Disadvantages

· No provision can be made for a relocation of objects that may have been 'incorrectly' grouped at an early stage. The result should be examined closely to ensure it makes sense.

· Use of different distance metrics for measuring distances between clusters may generate different results. Performing multiple experiments and comparing the results is recommended to support the veracity of the original results.

SUBMITTED BY:

SONIKA PRADHAN

ROLL NO. 1 2164

MBA (HR)

SIBM- Bangalore

Satisfaction survey analysis using statistics

Satisfaction surveys are an important tool for assessing the satisfaction of your customers, employees, patients and readers. Without a reliable way to analyze the responses, however, you risk making important decisions based on incomplete or superficial information.

Spreadsheets and databases give you simple summaries and basic row-and-column math. To best interpret and understand survey responses, you need in-depth analysis unavailable in spreadsheets and databases. When you use statistical software for satisfaction survey analyses, you get the most value from your data.

With statistical analysis, you can translate your survey responses into meaningful information and gain more insight into the responses. More insight, in turn, leads to better decisions. Using statistics to analyze your survey data helps ensure you’ll be delighted with the results.

7 ways statistics are better than spreadsheets for satisfaction surveys

Whether you are a beginner or a savvy, experienced survey researcher, these 12 ways show you how to better analyze your survey responses and present your results using statistics. They demonstrate why statistical software is a necessity for your analytical solution.

1. Use all your data efficiently combining and manipulating data holds the key to important results.

For a thorough analysis of your satisfaction survey, you need flexible data management. Combining responses from separate studies can be the key to spotting trends or patterns in your data. For example, merge the responses from your 1995 survey with the 1996 survey, and you can compare satisfaction scores between quarters or years to monitor improvements over time. To identify key segment differences you need to manipulate data within one survey. Aggregating your responses, for example, may help you determine whether department A responds differently than department B, or if different customer groups perceive your service differently. It’s important that your survey analysis tool offers flexible data management. SPSS lets you combine and manipulate data. Plus, you can aggregate, merge, split, subset and recode data.. SPSS also reads spreadsheet files and data in databases, such as Oracle or Microsoft Access.

Since all your data are valuable, it’s important to use a tool that can handle them. Don’t compromise your analysis because of software limitations. SPSS works with extremely large data sets, so you can combine several files without a problem. Whether you have 200 or 200,000 responses, SPSS handles your data easily. Typical spreadsheets, on the other hand, can handle only 16,000 responses. Once you reach that limit, the program does not accept any more data.

Match response samples to the true population.

Sometimes your response sample is different from your true population. With SPSS weighting, you can match the proportion of your responses to that of your population, so you don’t over- or under-represent groups in your analysis.

For example, you survey four regions and the response from the Northeast is underrepresented. SPSS’ weighting compensates for the low response, so all regions are equally represented. SPSS ensures your results are accurate, takes the worry out of who responds and reflects the reality of your population.

2. Know when there’s a problem with your data

Unusual responses affect the results of your satisfaction analysis and influence the decisions you make. It is important to know whether an unusual response is the result of a data entry error, and should be corrected, or whether it reflects a true relationship that exists in the data and should be considered in your decision. SPSS helps you easily spot data entry errors, respondent errors or unusual responses that you may want to leave out of your analysis, or look at more closely. A scatterplot gives you an overview of your data, helping you draw preliminary conclusions about possible relationships.

It also helps identify “outliers” which bear closer examination.

You can check data that do not follow general patterns or groupings to ensure they were not caused by a data entry error. Click on a point in the graph and see the response highlighted in the Data Editor. In spreadsheets, no link exists between your graphs and data. Once you click on the questionable point, it’s labeled on the graph and highlighted in the data. Further exploring this response is easy — you know its name (or label).

In this case, reviewing the document you used to enter the response showed the $185,000 income was a data entry error. Simply correct it and proceed. The $175,000 income was not a data entry error. You may choose to exclude it from further analyses.

3. Work easily with words, instead of numbers

Responses to satisfaction surveys use many questions with answer choices in categories, such as Male/Female, Yes/No, age ranges and scales of 1-5. When dealing with unfamiliar data, it can be difficult to remember what every coded answer represents. Often it’s more intuitive to work with words rather than numbers. The SPSS Data Editor, can show your data in words (labels) in place of numeric values (codes).

For example, it’s easy to see that 1 represents “extremely likely” so you don’t confuse it with 5 “not at all likely.” Switch between labels and codes in one mouse click, so you can better understand what is being analyzed. SPSS uses the underlying codes, so the calculations are still fast. And, your labels are automatically applied to your graphs and tables, so your results are easy to read and understand.

Satisfaction surveys often re-use questions and response options from previous surveys or within the same survey. For example, one survey may include several items asking about a customer’s rating of various products, all on a scale of 1 = poor to 5 = excellent, and coding “no answer” responses as 9.

SPSS stores all of your labels so you can reapply them to new surveys or additional questions whenever you need to. You save valuable time and reduce errors when preparing your data for analysis.

4. Get accurate results even when some data are missing

For many reasons, satisfaction survey respondents do not always answer every question. Missing responses occur when a question is not applicable, a respondent refuses to answer, or the respondent simply doesn’t know the answer. Gaps in your data influence your analysis and results. When responses are missing, SPSS accounts for them in the analysis, so you get accurate and meaningful results. SPSS lets you compare percentages with and without missing values to see the difference.

If you don’t consider the missing responses, you would overstate satisfaction. With SPSS, you can also specify multiple types of missing values so you can tell the differences between them. For example, respondents answer “don’t know” when they haven’t heard about the product or service; they respond with “N/A” when it does not apply to them.

You have the choice of eliminating different types of missing data during analysis, so you can find and understand patterns in respondents who answer “don’t know” versus “not applicable.” SPSS looks at responses on a question-by-question basis and can include a survey respondent’s answers only for those questions with a valid response

A typical spreadsheet package counts only data that have a blank entry as missing data. It does not allow you to separate other data you may wish to leave out of the analysis. This inflexibility may cause you to miss critical differences that exist. Some spreadsheets offer a work-around solution by suggesting a “hand-tailored” approach to the formulas in the individual cells. This work-around can be time consuming and error-prone.

5. Make better decisions by knowing what’s significant

It’s not enough to look at simple reports and try to draw conclusions from them. Often you notice differences or relationships that look interesting. For example, satisfaction scores may differ between groups. Perhaps new customers are more dissatisfied with your delivery times than long-time customers. Or, satisfaction with store A looks lower than store B. Perhaps the satisfaction score from this year is lower than last year. But are these findings really important? Are the differences enough to be “statistically significant?” You don’t have to be a statistician to understand significance. The comprehensive significance statistics in SPSS help you make better decisions by telling you immediately if your results are significant or if differences are random.

For example, in 1994, customers rated their satisfaction with a product at 6.2 on a 7-point scale with 1 = highly dissatisfied and 7 = highly satisfied. In 1995, the satisfaction rating increased to 6.4. Statistics help tell us if the increase is significant. In this case, a p-value of 0.001 tells us there is a difference between the two years’ scores.

SPSS significance statistics are easy to use and can usually be run along with another analysis with a simple click of your mouse button. To get explanations, definitions and rules of thumb for statistics or results, simply click “What’s This?”. SPSS has more significance tests than spreadsheets, so you can be confident in the interpretation of your results.

6. Save time and money using small samples

Sometimes, you survey less than 50 respondents or only a handful of people return the questionnaire.

Other times, you want to subset your data into small groups. For example, you analyze results by department, but many departments have only a small number of employees. With SPSS, you can work with smaller data sets and still get good results. Traditionally, you need large samples to get reliable significance testing, but a statistical technique called exact tests gives you correct statistics even with a few responses per question. Spreadsheets and databases don’t have exact tests. Don’t miss valuable opportunities because you don’t think you can rely on the results due to a small number of responses. With SPSS, simply click a single check box and you can act on your findings with confidence.

For example, a newspaper’s reader survey produced over 200 responses. In “slicing and dicing” the survey data, they discovered an interesting subset of 10 responses: women who read the “Daily News” and earn $50,000 enjoy the business section. To determine if this relationship was significant, the newspaper performed a significance test.

Without exact tests, the test result was insignificant (the traditional p-value of .056 was high).With exact tests for small samples, the result was significant (the p-value of .034 was low).With this evidence, the newspaper used the survey results to target the business section to women with an income of $50,000.

7. Use the right tool for the job to save time and increase productivity For satisfaction survey analysis, SPSS gives you all the tools you need for better, more informed decisions.

With SPSS, the answers to your questions are easy to find and understand.You’ll be productive quickly, with the online tutorial that teaches you the basics of data analysis and gives you step-by-step operations for common tasks. Examples guide you through the program and get you up and running quickly. The statistical glossary provides pop-up definitions to clarify unfamiliar statistical terms. And, “What’s This?” offers “rules of thumb” that help explain and define results. SPSS manuals are rich with examples that guide you through understanding the basics and obtaining results you can trust. And, as your analytical needs change, SPSS grows with you. You start with the basics, customized to meet your specific needs. Then, as your needs develop further, you can easily add more sophisticated survey research tools.

SPSS does the work for you.

Survey analysis tasks are often repetitive. You may reuse the same questionnaire periodically or run a set of standard reports for each survey project. With SPSS, you can work more efficiently by processing data unattended. Set up the reports and graphs you want, enter new data, and all you have to do is substitute the time period and file name in the dialog box. SPSS does the rest.

Submitted By:

SNEHI SHRESHTHA

ROLL NO. – 12044

SIBM- Bangalore