Search This Blog

6.2.11

Cluster Analysis


Introduction

Cluster analysis is finding similarities between data according to the characteristics found in the data and grouping similar data objects into clusters. A cluster is a collection of data objects in which objects similar to one another are put within the same cluster and dissimilar to the objects in the other clusters. Clusters analysis falls into the area of unsupervised learning as there are no pre-specified groups or classes. These are used typically as a stand-alone tool to get insight into data distribution or used as a pre-processing step for other algorithm. It is used in Pattern Recognition, Spatial Data Analysis, Image Processing, Market Research, Documents classification on the internet. It helps in City planning, Earth Quake studies, Insurance, stock selection and marketing programs.

A good clustering method produces high intra-class similarity and low inter-class similarity. The quality of a clustering result depends on both similarity measure used by the method and its implementation. Quality is usually expressed in terms of a distance function or metric.

Categorization of Algorithms

Clustering algorithms can be categorized into partitioning method, hierarchical methods, density-based methods, hierarchical methods, density-based methods, grid-based methods and model based methods. The most common among are partitioning method and Hierarchical approach. In partitioning approach, the objects are partition into k clustered on basis of distance (similarity/dissimilarity). This approach is depicted in the below diagram.



Issues with Partitioning Method (K-mean)

The problem with this approach is that number of clusters has to be specified in advance. Secondly it is not able to handle noisy data and outliers. This problem can be easily handled by Hierarchical clustering. In hierarchical clustering distance matrix is used for clustering so it does not required pre-specifying the number of clusters in advance. It only requires the termination conditions. The below diagram depicts the hierarchical clustering formation in an top down (agglomerative) and bottom up approach (divisive).



Dendrogram

Dendrogram shows how the clusters are merged. It decomposes the objects into a several levels of nested partitioning (tree). The clustering of the data objects is obtained by cutting the dendrogram at the desired level and then each connected component forms a cluster. As shown in the below figure, left diagram show two clusters (below the magenta line) while right diagram shows three clusters.



Uses of Hierarchical clustering in Finance

The hierarchical clustering methodology can be applied in stock selection of selected companies. Since the stock returns are likely to be similar in a region due to geography and macroeconomic conditions so identification of similar clusters allows one to track similar returns but different risks. After the identification of a cluster, the investor can used the output (clusters) according to his interest – for instance, he can look for same return stock and then choose to minimize risk or else he can pick a cluster of same risk stocks and high return. In this way the investor can profit (or lose less) using cluster analysis to select stock.

References

  1. Stock selection based on cluster analysis” by Costa, Cunha, Silva.
  2. http://people.revoledu.com/kardi/tutorial/Clustering/dendogram.htm
  3. http://en.wikipedia.org/wiki/Dendrogram
  4. www.gersteinlab.org/courses/545/07-spr/slides/DM_clustering.ppt


Submitted By:

Mrityunjay Kapoor
12147
Group 5
Finance Batch of 2009-11
SIBM Bangalore.

Factor Analysis (Multivariate) of Ground Water Quality of Makhmor Plain/ North Iraq

INTRODUCTION

Classification of wells according to their water quality can provide useful information for the users. Complex processes control the distribution of water quality parameters in groundwater, which typically has a large range of chemical composition. The ground water quality depends not only on natural factors such as the litho logy of the aquifer, the quality of recharge water and the type of interaction between water and aquifer, but also on human activities, which can alter these groundwater systems either by polluting them or by changing the hydrological cycle.

Sophisticated data analysis techniques are required to interpret groundwater quality effectively. The univariate statistical analysis has been generally used to treat ground water quality. The simplicity of the univariate statistical analysis is obvious and likewise the fallacy of reductionism could be apparent. In order to avoid this problem, multivariate analysis was used to explain the correlation amongst a large number of variables in terms of small number of factors without losing much information. The intention underlying the use of multivariate analysis is to achieve great efficiency of data compression from the original data, and to gain some information useful in the interpretation of the environmental geochemical origin. This method can also help to indicate natural association between variables. Multivariate treatment of environmental data is widely successfully used to interpret the relationship among the variables, so that the environmental systems could be better.

Makhmor plain is an important area in Northern Iraq, well-known with agricultural activities. Irrigation with ground water in the plain had got more attention in the last years. This plain lies in south east Erbil city. It is surrounded by the upper Zab river from the north, lower Zab river from the south, Tigris river from the west and Qara Chauq Mountain from the east. It has an area of 2700 km. Many studies were conducted on the water quality of Makhmor plain. Ground water of the plain has bad quality. Sulfate is a principal component in the groundwater of Makhmor area due to high concentration in the soil and the rocks.

OBJECTIVE

This research aimed to apply multivariate statistical analysis on groundwater quality of Makhmor plain by using factor analysis to indicate natural association between variables. Also the research tries to classify the wells of the plain using cluster analysis into groups according to their water quality.

METHODOLOGY

Thirty five deep wells and 28 shallow wells lying in an area of 2700 km in Makhmor plain were included in the study. Ground water quality parameters were represented by pH, Ca2+, Mg2+, Na+, Boron, K+, Cl-, SO42-, CO3- + HCO3-, NO3-.

Factor analysis extracted two factors from the water quality parameters of the deep wells. Factor I accounted for more than 50% of the variance among water quality. Cations including Boron, Na, Mg and K with anions including Cl-, SO42- and NO3- were loaded significantly on it. It represented the variation in the geological formation of study area, inconsistent distribution of agricultural activities and wastewater. For shallow wells, factor analysis extracted three factors. Factor I accounted for more than the 50% of the variance in the water quality. Six of water quality parameters were loaded on factor I. These parameters included pH and cations represented by boron, Na+ and Mg2+ in addition to Cl- and SO42- as anions. Cluster analysis had divided the deep wells into three groups with 50% similarity. Cluster I included two wells with the worst water quality, while cluster II had the lowest concentrations of cations and anions in the area and includes 8 wells. Cluster III showed mid concentrations between I and II clusters. For shallow wells, three clusters were obtained with 37.5% similarity. Cluster I included 7 wells with worst water quality, while cluster II exhibited the lowest concentrations of ions. These results obtained from the multivariate analysis can be very useful for the farmers and the users of ground water in this area.

SOURCE

http://en.wikipedia.org/wiki/Factor_analysis

http://www.damascusuniversity.gov.sy/mag/eng/images/stories/19-26.pdf


SUBMITTED BY

Madhumita Das
12089
Finance Batch of 2009-11
SIBM Bangalore

Guide to Decision making in Factor Analysis

A. Decide how to treat missing data

1. List wise deletion may shrink the sample size too much.

2. Pairwise deletion may make the analysis impossible depending on missing data patterns.

3. Either list wise or pairwise deletion has the potential to bias results; consider imputation.

4. Be prepared to run multiple analyses employing different missing data strategies.

B. Assess suitability of data for factor analysis

1. Distributions should not be excessively far from normal; may need to transform.

2. Relationships need to be essentially linear; may need to transform.

3. Sampling adequacy (often optional). Desirable to have high RS relative to partial RS.

a. As a global measure, Kaiser-Meyer-Olkin statistic should be at least 0.5 and probably should be greater than 0.6.

b. Anti-image correlations: for individual variables, use the same standard as for KMO, and this can help decide inclusion/exclusion.

C. Extraction method

1. Use principal components only if looking to use all information in pure data reduction (i.e., if it is reasonable to assume ~0% of the info is measurement error).

2. Can use maximum likelihood extraction, and its associated goodness of fit test for number of factors, if

a. variable distributions are fairly close to normal, i.e., skewness less than 2 and kurtosis less than 7, or if willing to transform or drop non-normal variables.

b. sample size is moderate: goodness of fit test is sensitive to small or large n.

3. In most social science cases, principal axis factoring works well.

D. Number of factors in a principal axis factor solution

1. Plan on needing to try multiple solutions.

2. Begin by extracting factors from all initial eigenvalues of at least 0.7. (Much research has shown that including only initial eigenvalues greater than equal 1 can under- or over-estimate the number of factors.) Once you see how solutions turn out after extraction and rotation, you can specify a minimum-eigenvalue or number-of-factors criterion based on something more informed than just the Kaiser Guttman rule of "eigenvalues greater than equal 1."

3. Scree test, visual or numerical, is best performed on extracted eigenvalues, rather than initial (although initial is what is displayed in SPSS).

4. Factor correlations should not be too high. If they are, and if largest initial eigenvalue is more than 10 times as large as the next, consider a one-factor solution (Harman; Kriebeck; Tracey).

a. Assuming correlations are not too high, the variance explained by each factor after rotation should be at least 5 or 10%.

b. Consider Terence (T. J. G.) Tracey's method of partialing out the first un-rotated factor as a "general factor" attributable to bias, e.g., to social desirability or common method bias.

5. Each factor should have at the very least 2 and probably 3 variables with high loadings.

6. Only retain those factors that are interpretable

E. Rotation method

1. Use oblique rotation as the default. There is no nothing to lose: if correlations among factors are very low, you can always switch to varimax (in which case results should hardly change anyway). But by starting with varimax you risk ignoring meaningful correlations.

2. To display two factors' relationship after finalizing a solution, can use varimax to plot the two clusters of variables and show the factors' correlation (which will be represented by the angle between them, with r equal to the cosine).


Source: An article on factor analysis written by Roland B. Stark in 19th February, 2007

Website link: www.integrativestatistics.com

Submitted By

Dipayan Kabiraj
12133
Group 5
SIBM Bangalore

Discriminant Function Analysis

General Purpose:

Discriminant function analysis is used to determine which variables discriminate between two or more naturally occurring groups. For example, an educational researcher may want to investigate which variables discriminate between high school graduates who decide (1) to go to college, (2) to attend a trade or professional school, or (3) to seek no further training or education. For that purpose the researcher could collect data on numerous variables prior to students' graduation. After graduation, most students will naturally fall into one of the three categories. Discriminant Analysis could then be used to determine which variable(s) are the best predictors of students' subsequent educational choice.
A medical researcher may record different variables relating to patients' backgrounds in order to learn which variables best predict whether a patient is likely to recover completely (group 1), partially (group 2), or not at all (group 3). A biologist could record different characteristics of similar types (groups) of flowers, and then perform a discriminant function analysis to determine the set of characteristics that allows for the best discrimination between the types.

Computational Approach:

Computationally, discriminant function analysis is very similar to analysis of variance (ANOVA). Let us consider a simple example. Suppose we measure height in a random sample of 50 males and 50 females. Females are, on the average, not as tall as males, and this difference will be reflected in the difference in means (for the variable Height). Therefore, variable height allows us to discriminate between males and females with a better than chance probability: if a person is tall, then he is likely to be a male, if a person is short, then she is likely to be a female.
We can generalize this reasoning to groups and variables that are less "trivial." For example, suppose we have two groups of high school graduates: Those who choose to attend college after graduation and those who do not. We could have measured students' stated intention to continue on to college one year prior to graduation. If the means for the two groups (those who actually went to college and those who did not) are different, then we can say that intention to attend college as stated one year prior to graduation allows us to discriminate between those who are and are not college bound (and this information may be used by career counselors to provide the appropriate guidance to the respective students).
To summarize the discussion so far, the basic idea underlying discriminant function analysis is to determine whether groups differ with regard to the mean of a variable, and then to use that variable to predict group membership (e.g., of new cases).

Submitted By:
Ashish Bhatnagar
Roll No: 12128
Group 12

Factor Analysis- What It Can & Can't Do

Factor analysis is a statistical technique that is used to determine the extent to which a group of measures share common variance. Factor analysis is sometimes termed a "data reduction" technique because the method is frequently used to extract a few underlying components (or factors) from a large initial set of observed variables. It is extensively used in psychological research concerned with the construction of scales intended to measure attitudes, perceptions, motivations, and so forth. Business-related applications are numerous and examples include the development of scales used to measure customer satisfaction with products and employee work attitudes. Factor analysis, however, has applicability outside of the realm of psychological research. It may be used, for example, by financial analysts to identify groups of stocks whose prices fluctuate in similar ways. Factor analysis often plays a crucial role in establishing the validity of employment tests and performance appraisal methods, thus helping a firm defend itself against employment discrimination charges.
There are many different methods of factor analysis and the underlying mathematical theory is quite complex. The basic elements of factor analysis, though, are relatively simple to understand. An example of the use of factor analysis might involve research designed to construct a scale of employee job satisfaction. Initially, a researcher or consultant may assemble a large set of questionnaire items that seem to be related to job satisfaction. These items will generally be presented to subjects along with some type of numeric or verbal scale.
Factor analysis includes both component analysis and common factor analysis. More than other statistical techniques, factor analysis has suffered from confusion concerning its very purpose.
Disadvantages of Factor Analysis
1. Factor analysis is a statistical method for attempting to find what are known as latent variables when you have data on a great many questions. Latent variables are things that cannot be directly measured. For example, most aspects of personality are latent. Personality researchers often ask a sample of people a lot of questions that they think are related to personality, and then carry out factor analysis to determine what latent factors exist.
2. The factors that appear can only come from the answers to the questions you ask. If you do not ask about sleep habits, for example, then no factor related to sleep habits will appear. On the other hand, if you ask only about sleep habits, then nothing else can appear. Selecting a good set of questions is complicated, and different researchers will choose different sets of questions. Random Data Gives Factors
3. If you generate a lot of random numbers, a factor analysis may still find apparent structure in the data. It is difficult to tell if the factors that emerge reflect the data or are simply part of the power of factor analysis to find patterns. It Is Hard to Decide How Many Factors to Include.
4. One task of the factor analyst is deciding how many factors to keep. There are a variety of methods for determining this, and there is little agreement as to which is best. Interpretation of the Meaning of the Factors Is Subjective.
5. Factor analysis can tell you which variables in your dataset "go together" in ways that aren't always obvious. But interpreting what those sets of variables actually represent is up to the analyst, and reasonable people can disagree.

Submitted By:
Bipin Easow
Roll No:12094
Group 12

Factor Analysis-An Overview

Factor analysis is used mostly for data reduction purposes:
– To get a small set of variables (preferably uncorrelated) from a large set of
variables (most of which are correlated to each other)
– To create indexes with variables that measure similar things (conceptually).
Two types of factor analysis:
· Exploratory Factor Analysis(EFA)- It is exploratory when you do not have a pre-defined idea of the structure or how many dimensions are in a set of variables.
· Confirmatory Factor Analysis(CFA)- It is confirmatory when you want to test specific hypothesis about the structure or the number of dimensions underlying a set of variables (i.e. in your data you may think there are two dimensions and you want to verify that).
Factor analyses are performed by examining the pattern of correlations (or covariances) between the observed measures. Measures that are highly correlated (either positively or negatively) are likely influenced by the same factors, while those that are relatively uncorrelated are likely influenced by different factors.
In general, you want to use EFA if you do not have strong theory about the constructs underlying responses to your measures and CFA if you do. It is reasonable to use an EFA to generate a theory about the constructs underlying your measures and then follow this up with a CFA, but this must be done using seperate data sets. You are merely fitting the data (and not testing theoretical constructs) if you directly put the results of an EFA directly into a CFA on the same data. An acceptable procedure is to perform an EFA on one half of your data, and then test the generality of the extracted factors with a CFA on the second half of the data.
If you perform a CFA and get a significant lack of it, it is perfectly acceptable to follow this up with an EFA to try to locate inconsistencies between the data and your model. However, you should test any modifications you decide to make to your model on new data.

Submitted By:
Devina Singh
Roll No:12056
Group 12

Basic Idea of Factor Analysis as a Data Reduction Method... by Kathireshan R



The main applications of factor analytic techniques are:  to reduce the number of variables and to detect structure in the relationships between variables, that is to classify variables. Therefore, factor analysis is applied as a data reduction or structure detection method
Confirmatory factor analysis. allows you to test specific hypotheses about the factor structure for a set of variables, in one or several samples
Correspondence analysis. Correspondence analysis is a descriptive/exploratory technique designed to analyze two-way and multi-way tables containing some measure of correspondence between the rows and columns.
Combining Two Variables into a Single Factor. A regression line can then be fitted that represents the "best" summary of the linear relationship between the variables. If we could define a variable that would approximate the regression line in such a plot, then that variable would capture most of the "essence" of the two items. Subjects' single scores on that new factor, represented by the regression line, could then be used in future data analyses to represent that essence of the two items. In a sense we have reduced the two variables to one factor. Note that the new factor is actually a linear combination of the two variables.
Principal Components Analysis. The example described above, combining two correlated variables into one factor, illustrates the basic idea of factor analysis, or of principal components analysis to be precise (we will return to this later). If we extend the two-variable example to multiple variables, then the computations become more involved, but the basic principle of expressing two or more variables by a single factor remains the same.
Principal Factors Analysis
before we continue to examine the different aspects of the typical output from a principal components analysis, let us now introduce principal factors analysis. Let us return to our satisfaction questionnaire example to conceive of another "mental model" for factor analysis. We can think of subjects' responses as being dependent on two components. First, there are some underlying common factors, such as the "satisfaction-with-hobbies" factor we looked at before.
Miscellaneous Other Issues and Statistics Factor Scores. We can estimate the actual values of individual cases (observations) for the factors. These factor scores are particularly useful when you want to perform further analyses involving the factors that you have identified in the factor analysis. Reproduced and Residual Correlations. An additional check for the appropriateness of the respective number of factors that were extracted is to compute the correlation matrix that would result if those were indeed the only factors. That matrix is called the reproduced correlation matrix. To see how this matrix deviates from the observed correlation matrix, you can compute the difference between the two; that matrix is called the matrix of residual correlations. The residual matrix may point to "misfits," that is, to particular correlation coefficients that cannot be reproduced appropriately by the current number of factors

Factor Analysis in Economic development... by Ajay Amarnath A



Factor analysis is a means by which the regularity and order in phenomena can be discerned. As phenomena co-occur in space or in time, they are patterned; as these co-occurring phenomena are independent of each other, there are a number of distinct patterns.
Parsimony or data reduction. Factor analysis can be useful for reducing a mass of information to an economical description. For example, data on fifty characteristics for 300 nations are unwieldy to handle, descriptively or analytically. The management, analysis, and understanding of such data are facilitated by reducing them to their common factor patterns.
 Structure. Factor analysis may be employed to discover the basic structure of a domain. As a case in point, a scientist may want to uncover the primary independent lines or dimensions--such as size, leadership, and age--of variation in group characteristics and behavior. Data collected on a large sample of groups and factor analyzed can help disclose this structure.
Classification or description. Factor analysis is a tool for developing an empirical typology.7 It can be used to group interdependent variables into descriptive categories, such as ideology, revolution, liberal voting, and authoritarianism. It can be used to classify nation profiles into types with similar characteristics or behavior.
 Scaling. A scientist often wishes to develop a scale on which individuals, groups, or nations can be rated and compared. The scale may refer to such phenomena as political participation, voting behavior, or conflict. A problem in developing a scale is to weight the characteristics being combined. Factor analysis offers a solution by dividing the characteristics into independent sources of variation (factors).
 Hypothesis testing. Hypotheses abound regarding dimensions of attitude, personality, group, social behavior, voting, and conflict. Since the meaning usually associated with "dimension" is that of a cluster or group of highly intercorrelated characteristics or behavior, factor analysis may be used to test for their empirical existence.
Data transformation. Factor analysis can be used to transform data to meet the assumptions of other techniques. For instance, application of the multiple regression technique assumes If the predictor variables are correlated in violation of the assumption, factor analysis can be employed to reduce them to a smaller set of uncorrelated factor scores.
Exploration. In a new domain of scientific interest like peace research, the complex interrelations of phenomena have undergone little systematic investigation. The unknown domain may be explored through factor analysis. It can reduce complex interrelationships to a relatively simple linear expression and it can uncover unsuspected, perhaps startling, relationships.
Mapping. Besides facilitating exploration, factor analysis also enables a scientist to map the social terrain. By mapping I mean the systematic attempt to chart major empirical concepts and sources of variation.

Crosstab Advantage... by Kathireshan Rajiah



A Crosstab should never be mistaken for frequency distribution because the latter provides distribution of one variable only. A Cross Table has each cell showing the number of respondents which gives a particular combination of replies. An example of Cross Tabulation would be a 3 x 2 contingency table. One variable would be age group which has three age ranges: 12-20, 21-30, and 31-up. Another variable would be the choice of polo shirt or t-shirt. With a crosstab, it would be easy to for a company to see what the choices of shirts are for the three age groups. For instance, the table would show that 20% of those aged 12-20 prefer polo, while only 10% of those aged 31-up prefer t-shirts. With the information, they can up with moves which will be beneficial to the success of the business.
Cross Tabulations are popular choices for statistical reporting because they are very easy to understand and they are laid out in a clear format. They can be used with any level of data whether the data is ordinal, nominal, interval or ratio because the Crosstab will treat all of them as if they are nominal data. Crosstab tables are provide more detailed insights to a single statistics in a simple way and they solve the problem of empty or sparse cells.Since Cross Tabulation is widely used in statistics, there many statistical process and terms that are closely associated with it. Most of these processes are methods to test the strengths of Crosstabs which is needed to maintain consistency and come up with accurate data because data being laid out using Crosstabs may come from a wide variety of sources.
The Lambda Coefficient is a method of testing the strength of association of Crosstabs when the variables are measured at nominal level. Cramer’s V is another testing method that test the strength of Crosstabs which adjusts the number of rows and columns. Other ways to test the strength of Crosstabs associations include Chi-square, Contingency Coefficient, Phi Coefficient and the Kendall tau. Companies find the services of a data warehouse very indispensable. But inside the data warehouse can be found billions of data which most of them are unrelated. Without the aid of tools, these data will not make any sense to the company. These data are not homogenous. They may come from various sources, often from other data suppliers and other warehouses which may be coming from other branches in other geographical locations.
Software applications like relational database monitoring systems have Cross Tabulation functionalities which allow end users to correlate and compare any piece of data. Crosstab analysis engines can examine dozens of table very fast and efficiently and these engines can even create full statistical outputs by very clicks of the mouse or keyboards.

How Crosstabs different from Scatterplot and its limitations... By Ajay Amarnath A



Crosstabs is an SPSS procedure that cross-tabulates two variables, thus displaying their relationship in tabular form. In contrast to Frequencies, which summarizes information about one variable, Crosstabs generates information about bivariate relationships.
Crosstabs creates a table that contains a cell for every combination of categories in the two variables. Inside each cell is the number of cases that fit that particular combination of responses. SPSS can also report the row, column, and total percentages for each cell of the table.
Because Crosstabs creates a row for each value in one variable and a column for each value in the other, the procedure is not suitable for continuous variables that assume many values. Crosstabs is designed for discrete variables--usually those measured on nominal or ordinal scales.Like crosstabs, scatterplot portrays the joint distribution of two variables.Unlike crosstabs, scatterplot is designed for continuous variables which distribute cases across unique points in space.To underscore the difference, consider this scatterplot that plots "party identification" (dependent) by "ideology" (independent) for several hundred cases in the vote00 file:
 Crosstabs are usually presented with the independent variable across the top and the dependent along the side.By convention, the independent variable is arranged across the top of the table, unless number of categories or size of space prohibit. ALWAYS, percentages are computed within the categories of the independent variable -- as shown in the sample table.  
Percentages are computed by rows only if the layout of the data call for placing the independent variable in the rows.That may be needed if there are more categories in the independent variable than fit easily along the columns Only unique analytical needs invite calculating percentages by totals--avoid doing this unless you know why. SPSS offers the option of calculating percentages all three ways, but that produces a cluttered table.  avoid checking all three options for percentages.
Limitations of crosstabs print format crosstabs tables in SPSS can't handle more than seven categories in a column variable without "wrapping" over.There is no limitation on the number of categories for the dependent variable -- down the side. Consequences of the limitation