Search This Blog

3.2.11

Cluster analysis Usage in various fields....... By Ajay Amarnath A


It is used mainly for software project planning control and management, an accurate estimate of software development cost is important. Past research has focused on using parametric models to predict development cost. The integration a neural network method with cluster analysis to estimate development cost.
Clustering is an economic development model signifying growth of similar kinds of industries at one geographical location. Locating near other similar firms provides numerous competitive advantages, including sharing a common labor pool, enhancing close working relationships between firms, reducing transaction costs and travel times between customers and suppliers, and enhancing the spread of technology through firms in the region.
As a cluster in a region takes root and expands, synergies often develop between firms and institutions, spurring additional growth and innovation. The existence of demand centre and concentration of Service Providers around the cluster also contributes to the growth of the cluster in terms of number of units. Other stakeholders like consultants, equipment manufacturers, Government Agencies etc also get concentrated in the cluster.
A cluster analysis was then performed to identify aspects of low, medium, and high risk projects. An examination of risk dimensions across the levels revealed that even low risk projects have a high level of complexity risk. For high risk projects, the risks associated with requirements, planning and control and the organization become more obvious. The influence of project scope, sourcing practices, and strategic orientation on project risk dimensions was also examined. Results suggested that project scope affects all dimensions of risk, whereas sourcing practices and strategic orientation had a more limited impact.
In marketing, cluster analysis is used for segmenting the market and determining target markets. Product positioning and New Product Development Selecting test markets the basic procedure. Formulate the problem - select the variables that you wish to apply the clustering technique.

How to solve human machine interface problem...... By KATHIRESHAN R


I have doubts regarding the human machine interface of the tool. i.e.
while we are evaluating the null hypothesis we are checking the significance value in the chi-square table and come to a conclusion whether to accept or reject the hypothesis. But the problem is, the chi-square value remains same and the way we define our hypothesis plays an important role. It can clearly be explained by 2 scenarios.

Scenario 1:
Null Hypothesis: Suppose there is no relationship between males who are educated are early married.
We found the significance value to be 0.05 from the chi-square table and hence we come to a conclusion that there is a relationship between males who are educated and early marriage.

Scenario 2:
Null Hypothesis: Suppose if we define there is relationships between male who are educated are early married.
Since the significance value which we obtained from the chi-square table is not changed then it implies that there is no relationships between male who are educated are early married.

These two scenarios contradict each other.

How is this possible and how to solve the human machine interface problem?And how to overcome the assumption made by the software? Is there a better way of defining the hypothesis in the tool?

28.1.11

Discriminant Analysis: By Vaibhav Khaparde


Discriminant Analysis (DA) undertakes the same task as multiple linear regressions by predicting an outcome. However, multiple linear regressions is limited to cases where the dependent variable on the Y axis is an interval variable so that the combination of predictors will, through the regression equation, produce estimated mean population numerical Y values for given values of weighted combinations of X values. But many interesting variables are categorical, such as political party voting intention, migrant/non-migrant status, making a profit or not, holding a particular credit card, owning, renting or paying a mortgage for a house, employed/unemployed, satisfied versus dissatisfied employees, which customers are likely to buy a product or not buy, what distinguishes Stellar Bean clients from
Gloria Beans clients whether a person is a credit risk or not, etc.
DA is used when:
The dependent is categorical with the predictor at interval level such as age, income, attitudes, perceptions, and years of education, although dummy variables can be used as predictors as in multiple regression. Logistic regression can be of any level of measurement. There are more than two DV categories, unlike logistic regression, which is limited to a dichotomous dependent variable.

The major underlying assumptions of DA are:

The observations are a random sample;

Each predictor variable is normally distributed;
Each of the allocations for the dependent categories in the initial classification are correctly classified;
There must be at least two groups or categories, with each case belonging to only one group so that the groups are mutually exclusive and collectively exhaustive (all cases can be placed in a group);

Each group or category must be well defined, clearly differentiated from any other group(s) and natural. Putting a median split on an attitude scale is not a natural way to form groups. Partitioning quantitative variables is only justifiable if there are easily identifiable gaps at the points of division;

For instance, three groups taking three available levels of amounts of housing loan; the groups or categories should be defined before collecting the data;

The attribute(s) used to separate the groups should discriminate quite clearly between the groups so that group or category overlap is clearly non-existent or minimal;

Group sizes of the dependent should not be grossly different and should be at least five times the number of independent variables.




There are several purposes of DA:

To investigate differences between groups on the basis of the attributes of the cases, indicating which attributes contribute most to group separation. The descriptive technique successively identifies the linear combination of attributes known as canonical discriminant functions (equations) which contribute maximally to group separation.
Predictive DA addresses the question of how to assign new cases to groups. The DA function uses a person’s scores on the predictor variables to predict the category to which the individual belongs.

To determine the most parsimonious way to distinguish between groups.

To classify cases into groups. Statistical significance tests using chi square enable you to see how well the function separates the groups.

To test theory whether cases are classified as predicted.

Methods implemented in this area are Multiple Discriminant Analysis, Fisher's Linear Discriminant Analysis, and K-Nearest Neighbors’ Discriminant Analysis.
Multiple Discriminant Analysis
(MDA) is also termed Discriminant Factor Analysis and Canonical Discriminant Analysis. It adopts a similar perspective to PCA: the rows of the data matrix to be examined constitute points in a multidimensional space, as also do the group mean vectors. Discriminating axes are determined in this space, in such a way that optimal separation of the predefined groups is attained. As with PCA, the problem becomes mathematically the eigen reduction of a real, symmetric matrix. The eigen values represent the discriminating power of the associated eigenvectors. The nYgroups lie in a space of dimension at most nY - 1. This will be the number of discriminant axes or factors obtainable in the most common practical case when n > m > nY(where n is the number of rows, and m the number of columns of the input data matrix).
Linear Discriminant Analysis
is the 2-group case of MDA. It optimally separates two groups, using the Mahalanobis metric or generalized distanceIt also gives the same linear separating decision surface as Bayesian maximum likelihood discrimination in the case of equal class covariance matrices.
K-NNs Discriminant Analysis
: Non-parametric (distribution-free) methods dispense with the need for assumptions regarding the probability density function. They have become very popular especially in the image processing area. The K-NNs method assigns an object of unknown affiliation to the group to which the majority of its K nearest neighbors’ belongs.

Submitted By:

Vaibhav Khaparde
12054
Marketing





Factor Analysis : Introduction

Factor analysis is a collection of methods used to examine how underlying constructs influence the responses on a number of measured variables.

There are basically two types of factor analysis: exploratory and confirmatory..

1>Exploratory factor analysis (EFA) attempts to discover the nature of the constructs influencing a set of responses.

2>Confirmatory factor analysis (CFA) tests whether a specified set of constructs is influencing responses in a predicted way.

Both types of factor analyses are based on the Common Factor Model, illustrated in figure 1.1. This model proposes that each observed response (measure 1 through measure 5) is influenced partially by underlying common factors (factor 1 and factor 2) and partially by underlying unique factors (E1 through E5). The strength of the link between each factor and each measure varies, such that a given factor influences some measures more than others. This is the same basic model as is used for LISREL analyses.

Factor analyses are performed by examining the pattern of correlations between the observed measures. Measures that are highly correlated (either positively or negatively) are likely influenced by the same factors, while those that are relatively uncorrelated are likely influenced by different factors.

Exploratory Factor Analysis:

The primary objectives of an EFA are to determine

1. The number of common factors influencing a set of measures.

2. The strength of the relationship between each factor and each observed measure.

Some common uses of EFA are to

Ø Identify the nature of the constructs underlying responses in a speci¯c content area.

Ø Determine what sets of items \hang together" in a questionnaire.

Ø Demonstrate the dimensionality of a measurement scale. Researchers often wish to develop scales that respond to a single characteristic.

Ø Determine what features are most important when classifying a group of items.

Ø Generate \factor scores" representing values of the underlying constructs for use in other analyses.

Miscellaneous notes on EFA:

1>To have acceptable reliability in your parameter estimates it is best to have data from at least 10 subjects for every measured variable in the model. This number should be increased if you expect that the influence of the common factors is relatively weak. You should also have measurements from at least three variables for every factor that you want to include in your model.

2> You should endeavour to have a wide variety of measurements for your EFA. The more accurately that your selection of measurements properly represents the population" of measurements that could be taken, the more generality you will have in your findings.

3> EFA can be performed in SAS using proc factor. Principal component analysis can be performed in SAS using proc princomp, while it can be performed in SPSS using the Analyze/Data reduction/Factor analysis menu selection. EFA cannot actually be performed in SPSS (despite the name of menu item used to perform PCA).

confirmatory factor analysis

Confirmatory factor analysis (CFA) is a special form of factor analysis. It is used to test whether measures of a construct are consistent with a researcher's understanding of the nature of that construct (or factor).

CFA is commonly used in social research. CFA is frequently used when developing a test, such as a personality test, intelligence test, or survey. CFA is also frequently used as a first step to assess the proposed measurement model in a structural equation model. Many of the rules of interpretation regarding assessment of model fit and model modification in structural equation modeling apply equally to CFA. CFA is distinguished from structural equation modeling by the fact that in CFA, there are no directed arrows between latent factors.[clarification needed] In the context of SEM ,the CFA often is called 'the measurement model', while the relations between the latent variables (with directed arrows) are called 'the structural model'


Source: Wikipedia, Asparouhov, T.;Muthén, B. (2009). "Exploratory structural equation modeling". Structural Equation Modeling, 16, 397-438., http://www.stat-help.com/factor.pdf


Blog By: Akhilesh Agarwal (Operations_SIBM)

Factor analysis titus raju 12053

Factor analysis
Factor analysis attempts to discover the nature of the constructs influencing a set of responses.
Both types of factor analyses are based on the Common Factor Model where model proposes that each observed response (measure 1 through measure 5) is influenced partially by underlying common factors (factor 1 and factor 2) and partially by underlying unique factors (E1 through E5). The strength of the link between each factor and each measure varies, such that a given factor influences some measures more than others.
Factor analysis helps you find these undetermined variables by looking at the variables you have actually collected.

An important element of the factor analysis output is the standardized factor score coefficients (Output-Table), which gives location of each product on each factor.
The vectors are obtained based on the amount of correlation the original attitudes possess with the factor scores (represented as factors). The direction of the vectors indicates the factor with which each attribute is associated, and the length of the vector indicates the strength of association. Thus, on the left map the “filling” attribute has little association with any factor, whereas on the right map the “filling” attribute is strongly associated with “refreshing” factor.Although a factor is not observable like the other original variables, it is still a variable. One output of most factor analysis programs is the values for each factor for all respondents. These values are termed factor scores and are shown above for three factors that were found to underlie the five input variables. Thus, each beverage has a factor score on each factor, in addition to the beverage’s rating on the original eight attributes. In subsequent stages while doing perceptual mapping this factor scores are used to position beverages in the perceptual map.

My learning from class:
Factor analysis is basically what is done when you have a large number of variables to work with. Such a large number of variables makes it very difficult to organize data, and analyze it. Thus, we can use factor analysis, to reduce the number of variables so that it becomes much easier to work.
Factor analysis uses correlation between variables to see which ones are related, and can be eliminated, without causing significant impact on the output of analysis.
For this purpose, we take a default eigen value >1 and use verimax rotation so that the first few variables have the maximum effect.
Using Rotated Component Matrix, we can very well determine which variables are to be combined into factors, and keeping a limit on the eigen values to be accepted, we can determine what percentage of the data is contained in the factors we are choosing. According to this, we can increase or decrease the number of factors to simplify our calculations.
We can look at the correlation between the variables in scatter plot, in which visually clustered patterns can be made out and combined.
The remaining factors can be combined to give more meaning to the analysis.



Regards
Titus Raju

27.1.11

factor analysis

Understanding of factor analysis describes us a method for investigating whether a number of variables of interest are linearly related to a smaller number of unobservable factors. It determines the strength of relationship between each factor and each observed measure. By the exercises to comprehend factor analysis we derive that it is a statistical tool to account for variability in a set of measured items in terms of a smaller numbers factors. There are several ways to conduct factor analysis and the choice of method depends on many things. We have the options pertaining to the retention of factors. The choice of either selexting factors with eigen values greater than a user- specified value or retaining a fixed number of factors. By looking at the scree plot and the Eigen values over 1 will lead us to retain the same number of factors then continue with analysis.
The objective of the factor analysis is to objectively detect natural groupings of variables. It also aims to extract quantitative information from large matrices using objective statistical criteria. It reduces the redundant data in the list.
The software uses the sample size from which communalities after extraction should probably be above 0.5. The system then finds a factor solution to a set of variables. When the first factor solution does not reveal the hypothesized structure of the loadings, it is customary to apply rotation in an effort to find another set of loadings that fit the observations equally well but can be more easily interpreted. The most widely used of these is the varimax criterion. It seeks the rotated loadings that maximize the variance of the squared loadings for each factor; the goal is to make some of these loadings as large as possible, and the rest as small as absolute value. It encourages the detection of factors each of which is related to few variables. It discourages the detection of factors influencing all variables.
The interpretability of factors can be improved through rotation. Rotation maximizes the loading of each variable on one of the extracted factors whilst minimizing the loading on all other factors. Rotation works through changing the absolute values of the variables whilst keeping their differential values constant.

The scores factor allows one to save factor scores for each subject in the data editor. It reates new column for each factor extracted and then places the factor score for each subject within that column. These scores can then be used for further analysis, or simply to identify groups of subjects who score highly on particular factors. Another feature provided by SPSS is options which helps use to list variables by size.
To analysis the output in SPSS software, the R matrix shows the Pearson correlation coefficient between all pairs of questions whereas the bottom half contains the one tailed significance of these coefficients. We can use this matrix to check the pattern of relations. The significance values details provides the confidence level or the reliability of the data. If any data of significance value more than 0.9 is found, then one should be aware that a problem could arise because of singularity in the data which demands for the check of determinants of the correlation matrix and if necessary eliminate one of the two variables causing the problem. Generally, the data is considered based on the null hypothesis proven right or wrong. Based on the acceptance or rejection of null hypothesis, further analysis is done on the information obtained.

SPSS lists the Eigen values associated with each linear component before extraction, after extraction and after rotation. The eigen values associated with each factor represent the variance explained by that particular linear component and it also displays the eigen values in terms of the percentage of variance explained.

Another output table displays the communalities before and after extraction. The principal component analysis works on the initial assumption that all variance are common, therefore before extraction he communalities are 1. Another way to look at these is in terms of the proportion of variance explained by the underlying factors. The output also explains the component matrix before rotation. This matrix contains loadings of each variable onto each factor.
Another important output is the rotated component matrix which provides with loads of information. It is the matrix that contains same information as the component matrix except that it is calculated after rotation. Before rotation most variables loaded highly onto first factor and the remaining factors didn’t really get a look in.
Therefore, the detailed analysis report explains that SPSS as a tool is every exploratory and is it should be used to guide the researcher to make various decisions. The conclusions derived out of the SPSS analysis can help the leaders to take strategic decisions and build strategies to achieve them.

Regards,
Nisha

Live...let live!

Applications of Multivariate Analysis (CLUSTER Analysis) in International Tourism Research: The Marketing Strategy Perspective of NTOs


Source - Applications of Multivariate Analysis in International Tourism Research: The Marketing Strategy Perspective of NTOs by Satish Chandra & Dennis Menezes

In recent times International tourism has increased exponentially. With this growth the industry has become significantly more competitive, and the marketing role of National Tourism Organizations (NTOs) has taken on added significance. Correspondingly, research related to the marketing aspects of international tourism has increased.

The paper focuses on:

1. identifying and describing the key components of marketing strategy that must be addressed by NTOs, and

2. identifying and describing the multivariate statistical techniques most relevant to research that relates to enhancing the marketing strategies of NTOs along with citing some of the recent related research.

Multivariate Techniques used for achieving the Marketing Strategies.

Refer Above Diagram. Prior to addressing these tasks, a SWOT analysis should be completed.

Critical Element & Technique Used

Cluster Analysis (In Baseline/Post Hoc Segmentation): In Baseline/Post Hoc Segmentation tourists are classified into clusters on the basis of their appropriate attribute similarities.

Baseline segmentation involves analyzing a large cross sectional sample of tourists where data has been collected on a variety of variables, such as psychological, life style, demographic, and other variables of interest. The preferred mode of analyzing this large set of data is Cluster analysis. In the baseline segmentation approach using Cluster analysis, the segments are produced analytically.

Cluster analysis classifies the subjects into clusters, so that each subject is very similar to other subjects in that cluster with respect to selected criterion variables. The clusters formed exhibit high within cluster homogeneity and high between cluster heterogeneity. Thus, when good classification is achieved, subjects within clusters will be close together when plotted geometrically, but different clusters will be far apart. (Refer Figure 3 above).

In the context of segmenting tourism markets, Cluster analysis can be used to identify different clusters of tourists that exist within a larger group or market of tourists. As a result, Cluster analysis may be used to develop a taxonomy of different types of tourist segments and thereby gain a better understanding of the composition of the larger population of tourists. The within cluster similarity of the tourists is typically determined using an inter subject Euclidean distance measured on two variables.

Conclusion

International tourist arrivals increased from approximately 25 million in 1950 to 625 million in 1998, an increase of 2,500 percent. A WTO survey of NTOs and leading experts in tourism envision the following :

(1) international tourism arrivals by 2020 to be 1.6 billion, with spending in excess of 2 trillion U.S. dollars,

(2) the percent of the traveling population involved in international travel increasing from 3.5 percent in 1998 to 7 percent by 2020,

(3), Europe continuing to be the largest international tourism region, although by 2020 its market share being significantly eroded,

(4) by 2020 China being the largest receiver of international tourists,

(5) among the various international tourism market segments, eco-tourism, cultural tourism, theme based tourism, adventure tourism, and the cruise market growing in importance, and

(6) tourism as a sector growing at a faster rate than the global economy.

These predictions by the WTO suggest that the international tourism market will continue to expand at a rapid rate and become increasingly competitive. In this environment, the use of effective and efficient marketing strategies (based on appropriate usage of Multivariate Techniques) by NTOs as well as other international tourism organizations will therefore become increasingly important.

Submitted by;
Sumit Acharya
Roll - 12108
Finance Batch
SIBM Bangalore


Introduction to Logistic Regression


As we could not discuss the logistic regression in the class, below is some information about the same.

Meaning:

Logistic Regression is a type of predictive model that can be used when the target variable is a categorical variable with two categories – for example live/die, has disease/doesn’t have disease, purchases product/doesn’t purchase, wins race/doesn’t win, etc.

A logistic regression model does not involve decision trees and is more similar to nonlinear regression such as fitting a polynomial to a set of data values.

Logistic regression can be used only with two types of target variables:

1. A categorical target variable that has exactly two categories (i.e., a binary or dichotomous variable).

2. A continuous target variable that has values in the range 0.0 to 1.0 representing probability values or proportions.

As an example of logistic regression, consider a study whose goal is to model the response to a drug as a function of the dose of the drug administered. The target (dependent) variable, Response, has a value 1 if the patient is successfully treated by the drug and 0 if the treatment is not successful. Thus the general form of the model is:

Response = f(dose)

The input data for Response will have the value 1 if the drug is effective and 0 if the drug is not effective. The value of Response predicted by the model represents the probability of achieving an effective outcome, P(Response=1|Dose). As with all probability values, it is in the range 0.0 to 1.0.

One obvious question is “Why not simply use linear regression?” In fact, many studies have done just that, but there are two significant problems:

1. There are no limits on the values predicted by a linear regression, so the predicted response might be less than 0 or greater than 1 – clearly nonsensical as a response probability.

2. The response usually is not a linear function of the dosage. If a minute amount of the drug is administered, no patients will respond. Doubling the dose to a larger but still minute amount will not yield any positive response. But as the dosage is increases a threshold will be reached where the drug begins to become effective. Incremental increases in the dosage above the threshold usually will elicit an increasingly positive effect. However, eventually a saturation level is reached, and beyond that point increasing the dosage does not increase the response.

Submitted By

Kunal Shah

12029

Measuring the performance of MFIs: An application of factor analysis

The main applications of factor analytic techniques are:

(1) To reduce the number of variables.

(2) To detect structure in the relationships between variables, that is to classify variables.

Therefore, factor analysis is applied as a data reduction or structure detection.

Measuring the performance of microfinance institutions (MFIs) is not a small task. Indeed, looking at the financials of an MFI only gives its performance. As many MFIs primarily exist in order to help the poorest people, one also has to include aspects which influence their performance. Hence, MFIs' performance can be termed multidimensional.

This article which I’m sharing here, talks about how “Factor Analysis” as a statistical tools can offer new insights in the context of MFIs' performance evaluation. Factor analysis is used in a first step to construct performance indices based on several possible associations of variables without posing too many a prior restrictions.

Then the base variables are thus combined to produce different factors, each one representing a distinct dimension of performance. We then use the individual scores attributed to each MFI on each factor as the dependent variables of a simultaneous-equations model and present new evidence on the determinants of MFIs' performance.

Variables can be of different types which, according to me, can be classified as qualitative variables and quantitative as the two broad groups initially.

Qualitative variables like –

(i) Type of investors (are they PE investors or social investors)

ii) Motive of Investors (Highly Profit motive or Social Motive)

iii) Who Is investing

iv) Type of security accepted

v) Purpose of taking loan

etc. are the variables.

Qualitative variables like-

i. Profit on Every account

ii. Interest charged

iii. Amount of loan

Etc

Such Qualitative and Quantitative variables are then grouped into components on the basis of similarities. Such similar groups can be

1. investors 2. Security 3. Margins 4. Safety 5. Liquidity etc.

The correlation & variance within groups and across groups is then taken to know which are the variables are of same features and tackled in a same way. And also which are the most important and most influential ones.

Cluster of groups and their behaviour towards performance of MFIs will help the policymakers to solve the problems related to performance of MFIs.

In fact factor analysis is used for measuring the performances of various financial institutions and also for assessment of risks of banks & institutions.

Ref:- Sylvain Weber, University of Geneva, Department of Economics,Giovanni Ferro Luzzi, University of Geneva.

Submitted By

Akhil Parekh

12066



A NOTE ON CLUSTER ANALYSIS:

It is a multivariate technique used to determine group membership for cases or variables. In cluster analysis, the number of groups and the members of the groups are unknown. A cluster analysis of variables is like a factor analysis. A cluster analysis of cases is like a discriminant analysis. SPSS provides hierarchical cluster analysis and k-means cluster analysis. The hierarchical cluster analysis is for either cases or variables.

The k-means cluster analysis is for cases only when we have a large number (n >= 200) of cases. By using a distance measure the hierarchical algorithm combines closest pair of cases or variables to form a cluster. This technique continues to join pairs of cases (or variables) or cluster until the final step where all cases (or variables) are joined to form one cluster. Once two cases (or variables) are joined in earlier step, they remain together throughout the process. When some variables have large values or different scales of measurement, they should be standardized before using them in cluster analysis. Data for cluster analysis could be interval, count or binary. There are different distance measures depending on the type of data. Hierarchical cluster analysis provides way to automatically standardize the variables but one has to do standardization prior to using k-means cluster analysis.

This technique is common for reducing dimensionality among cases (case dimension reduction) or among variables (variable dimension reduction). Researchers usually take one step further by trying to make a meaningful interpretation of the common properties for each cluster, and by combining each cluster into a new single variable or a case, which are then used in other analyses.

A few important considerations for a successful cluster analysis:

Different measurement scales among different variables have dramatic influence on the clusters. It is important to make some kind of standardization of the measurement scales.

Outlying cases often dominate some clusters. It is important to take care of outliers.

Cluster analysis depends on the distance measures used for clustering. It is important to identify proper distance measure for the problem of interest. SPSS has a set of defaults. If you do not know what would be appropriate, the default options are usually the more commonly used.

The determination of the final number of clusters may differ from different criteria. It is a good idea to use several selection criteria to help you to choose the final number of clusters.

The context behind the problem of study is an important consideration in the choice of clustering techniques and the criteria for selecting the number of clusters.

METHODS OF CLUSTER ANALYSIS:

Two Step Cluster Analysis: This is an exploratory data analysis to determine clusters within a data set. This procedure works with both continuous and categorical variables. Cases are clustered based on the variables which could be continuous or categorical. To obtain a Two Step cluster analysis, go to Analyze, Classify, Two Step Cluster. This opens the main dialog box of Two Step Cluster Analysis. Select the distance measure, number of clusters and clustering criterion. The submenus are:

Options- This is where you select standardized variables and those to be standardized.

Plots- Enables you to select plots.

Output- Enables you to export final model.

K-Means Cluster Analysis: This is used to cluster cases when you have a large number of cases. The analysis requires one to specify the number of clusters. To run k-means cluster analysis, go to Analyze, Classify, k-Means Cluster. Select the variables and the classification method. There are three submenus:

Iterate- Allows you to select options for the iteration algorithm.

Save- Allows you to save results from the analysis.

Options- Allows you to select some results for display.

Hierarchical Cluster Analysis: This is used to cluster variables (or cases). One can analyze raw variables or use a variety of standardization to transform the variables. To run a Hierarchical Cluster Analysis, go to Analyze, Classify, Hierarchical Cluster. Select the variables for the analysis. Select cases or variables to cluster. By default, statistics and plots will be displayed. The main dialog box has the following submenus:

Statistics- Agglomeration schedule shows the cases or clusters combined at each state. Proximity matrix shows the distances/similarities between items. You can also display cluster membership by requesting a single solution or a range of solutions.

Plots- You can request dendrogram. By default, icicle of all clusters is displayed. You can turn this off. You can also select an orientation pattern.

Method- Here, one selects the cluster method, select the type of measure, choose whether to transform data values. Transform measures allows you to transform the distance measure values that are generated.

Save- Allows you to save cluster membership. This can be saved as a single solution or a range of solutions.

Submitted By

Neetu Rathod

Roll no:- 12118

SIBM Bangalore