Search This Blog

7.2.11

Use of Factor Analysis in Market Research

1. Removing attributes from a studyIf you are asking respondents to rate service as “fast”, “speedy” and “slow”, these attributes all reflect a latent factor of service speed. Many questionnaires feature such redundancy; a factor analysis can shorten the survey.

2. As a first step to regression – When conducting key driver analysis, you want to use latent factors instead of attributes so that you are focusing on different drivers.

3. As a first step to clustering – As with regression, you want to conduct factor analysis first when conducting cluster analysis. Unlike with regression, you don’t want to use factor scores but want to use the attribute with the greatest correlation to the score: cluster analysis works best with variability and lumpiness, which the factor scores would smooth out.

4. Creating factor scores – Rather than report on customer satisfaction, likelihood to recommend and likelihood to continue purchasing, for instance, track and trend the factor score: Advocacy Loyalty.

When designing a questionnaire that will use factor analysis, require respondents to answer each attribute that will be used, otherwise you will have to delete incomplete records or impute missing answers. You should also avoid “Don’t Know” and “Not Applicable” responses in this case. If you can’t avoid them, then run two separate factor analyses, one which excludes attributes with such responses and one that doesn’t, then calculate the correlation analysis with the attributes you removed.

When redesigning a questionnaire to remove attributes, look at the verbatim comments to see what attributes people discuss. Don’t just remove attributes based on how they load to the factor score. “In a perfect world, with a normal distribution of values around the mean, the worst attribute is where nearly everyone rates it a 4 or 5,” Ray said. “I’ve even seen some studies where an attribute is 5 of 5 for almost everyone.” Do look at the correlation coefficients, but remove attributes with a low standard derivation. And, of course, which attributes to remove can be political: “Customers [of market researchers] have some questions that they are more in love with than others, and some annual compensation might be linked to specific attributes.”

Submitted by:

Aditiya Hakim

12004

Group 5

Finance

SIBM-Bangalore

6.2.11

DISCRIMINANT ANALYSIS

Discriminant analysis is a very important statistical tool used by researchers throughout. It helps in enhancing or deriving conclusions out of raw data.It is slightly complicated in parts but on the whole it gives good analytical results.
Discriminant analysis is a statistical method that is used by researchers to help them understand the relationship between a "dependent variable" and one or more "independent variables." A dependent variable is the variable that a researcher is trying to explain or predict from the values of the independent variables. Discriminant analysis is similar to regression analysis and analysis of variance (ANOVA). The principal difference between discriminant analysis and the other two methods is with regard to the nature of the dependent variable.

Discriminant analysis requires the researcher to have measures of the dependent variable and all of the independent variables for a large number of cases. In regression analysis and ANOVA, the dependent variable must be a "continuous variable." A numeric variable indicates the degree to which a subject possesses some characteristic, so that the higher the value of the variable, the greater the level of the characteristic. A good example of a continuous variable is a person's income.
In discriminant analysis, the dependent variable must be a "categorical variable."

The values of a categorical variable serve only to name groups and do not necessarily indicate the degree to which some characteristic is present. An example of a categorical variable is a measure indicating to which one of several different market segments a customer belongs; another example is a measure indicating whether or not a particular employee is a "high potential" worker. The categories must be mutually exclusive; that is, a subject can belong to one and only one of the groups indicated by the categorical variable. While a categorical variable must have at least two values (as in the "high potential" case), it may have numerous values (as in the case of the market segmentation measure). As the mathematical methods used in discriminant analysis are complex, they are described here only in general terms. We will do this by providing an example of a simple case in which the dependent variable has only two categories.
Discriminant analysis is most often used to help a researcher predict the group or category to which a subject belongs. For example, when individuals are interviewed for a job, managers will not know for sure how job candidates will perform on the job if hired. Suppose, however, that a human resource manager has a list of current employees who have been classified into two groups: "high performers" and "low performers." These individuals have been working for the company for some time, have been evaluated by their supervisors, and are known to fall into one of these two mutually exclusive categories. The manager also has information on the employees' backgrounds: educational attainment, prior work experience, participation in training programs, work attitude measures, personality characteristics, and so forth. This information was known at the time these employees were hired. The manager wants to be able to predict, with some confidence, which future job candidates are high performers and which are not. A researcher or consultant can use discriminant analysis, along with existing data, to help in this task.
There are two basic steps in discriminant analysis. The first involves estimating coefficients, or weighting factors, that can be applied to the known characteristics of job candidates (i.e., the independent variables) to calculate some measure of their tendency or propensity to become high performers. This measure is called a "discriminant function." Second, this information can then be used to develop a decision rule that specifies some cut-off value for predicting which job candidates are likely to become high performers.

The tendency of an individual to become a high performer can be written as a linear equation. The values of the various predictors of high performer status (i.e., independent variables) are multiplied by "discriminant function coefficients" and these products are added together to obtain a predicted discriminant function score. This score is used in the second step to predict the job candidates likelihood of becoming a high performer. Suppose that you were to use three different independent variables in the discriminant analysis. Then the discriminant function has the following form:

where D = discriminant function score,
B , = discriminant function coefficient relating independent variable i to the discriminant function score,
X = value of independent variable i.
The equation is quite similar to a regression equation. Conventional regression analysis should not be used in place of discriminant analysis. The dependent variable would have only two values (high performer and low performer) and would thus violate important assumptions of the regression model. Discriminant analysis does not have these limitations with respect to the dependent variable.
Estimation of the discriminant function coefficients requires a set of cases in which values of the independent variables and the dependent variables are known. In the case described above, the company has this information for a current group of employees. There are several different ways that can be used to estimate discriminant function coefficients, but all work on the same general principle: the values of the coefficients are selected so that differences between the groups defined by the dependent variable are maximized with regard to some objective function. One commonly used objective function is the F-ratio, which is defined as it is in ANOVA and regression problems. The coefficients are chosen to maximize the F-ratio when analysis of variance is performed on the resulting discriminant function, using the dependent variable (i.e., job performance) as the grouping variable. Most general statistical programs, such as the Statistical Package for the Social Sciences, contain discriminant analysis modules.

There are various tests of significance that can be used in discriminant analysis. One widely used test statistic is based on Wilks lambda, which provides an assessment of the discriminating power of the function derived from the analysis. If this value is found to be statistically significant, then the set of independent variables can be assumed to differentiate between the groups of the categorical variable. This test, which is analogous to the F-ratio test in ANOVA and regression, is useful in evaluating the overall adequacy of the analysis.
Unfortunately, discriminant analysis does not generate estimates of the standard errors of the individual coefficients, as in regression, so it is not quite so simple to assess the statistical significance of each coefficient. For example, most discriminant analysis programs have a stepwise option. Independent variables are entered into the equation one at a time. Again, Wilks lambda can be used to assess the potential contribution of each variable to the explanatory power of the model. Variables from the set of independent variables are added to the equation until a point is reached for which additional items provide no statistically significant increment in explanatory power.

Once the analysis is completed, the discriminant function coefficients can be used to assess the contributions of the various independent variables to the tendency of an employee to be a high performer. The discriminant function coefficients are analogous regression coefficients and they range between values of -1.0 and 1.0. The first box in Figure 1 (on the facing page) provides hypothetical results of the discriminant analysis. The second box provides the within-group averages for the discriminant function for the two categories of the dependent variable. Note that the high performers have an average score of 1.45 on the discriminant function, while the low performers have an average score of -.89. The discriminant function is treated as a standardized variable, so it has a mean of zero and a standard deviation of one. The average values of the discriminant function scores are meaningful only in that they help us interpret the coefficients. Since the high performers are at the upper end of the scale, all of the positive coefficients indicate that the greater the value of those variables, the greater the likelihood of a worker being a high performer (e.g., education, motivation). The magnitudes of the coefficients also tell us something about the relative contributions of the independent variables. The closer the value of a coefficient is to zero, the weaker it is as a predictor of the dependent variable. On the other hand, the closer the value of a coefficient is to either 1.0 or -1.0, the stronger it is as a predictor of the dependent variable. In this example, then, years of education and ability to handle stress both have positive coefficients, though the latter is quite weak. Finally, individuals who place high importance on family life are less likely to be high performers than those who do not.

The second step in discriminant analysis involves predicting to which group in the dependent variable a particular case belongs. A subject's discriminant score can be translated into a probability of being in a particular group by means of Bayes Rule. Separate probabilities are computed for each group and the subject is assigned to the group with the highest probability. Another test of the adequacy of a model is the degree to which known cases are correctly classified. As in other statistical procedures, it is generally preferable to test the model on a set of cases that were not used to estimate the model's parameters. This provides a more conservative test of the model. Thus, a set of cases should, if possible, be saved for this purpose. Having completed the analysis, the results can be used to predict the work potential of job candidates and hopefully serve to improve the selection process.
There are more complicated cases, in which the dependent variable has more than two categories. For example, workers might have been divided into three groups: high performers, average performers, low performers. Discriminant analysis allows for such a case, as well as many more categories. The interpretation, however, of the discriminant function scores and coefficients becomes more complex. The books included in the "Further Reading" section below explain in detail how to perform discriminant analysis with multiple categories and provide in-depth technical discussions.

KAPIL SHARMA
12087
MARKETING (2009-11)

Garima Singh 12081

23/1/11

The introduction session of BI-Workshop was very loaded inspite being an introductory session but it was definitely a very good session. The data looked never so good, meaningful and colourful.
We started with First Level Analysis where we were taught how to calculate Frequencies with the given data sheet. Then we covered Cross Tabulation both for 2 variables as well as 3 variables.
The third and the last technique for the day was Cluster Analysis in which two different types of clusters were covered- Divisive and Agglomerative, both these are types of Hierarchical Cluster.
Hierarchy type of cluster is done for Variables provided the number of objects is less than 50. K-Means type of cluster is done for Cases provided the number of objects is more than 50.

24/1/11

The second session was more relaxed. We started Perceptual Mapping a.k.a PerMap.
PerMap is of two types- Overall Similarity and Attribute Based. In Overall Similarity a thorough knowledge is required what goes in people's minds because of which they think two particular objects are similar/dissimilar. One disadvantage is that the hidden attributes might get exposed.
In Attribute-Based PerMap attributes are provided which gives rise to one major disadvantage, it being the chance missing of an important attribute.
We then learned how to make a PerMap in both the types in two as well as three dimensions and also the significance of Objective Function Value which is nothing but an error and how it gets reduced when PerMap display is changed from Two dimention to Three dimension.

25/1/11

The concluding session was on Factor Analysis wherein we learned how to divide a whole lot of given factors can be divided and then using factor
analysis (extraction method) less important factors can be eliminated.
Z-Scores were also covered in the last session of the workshop where one of the major things that we observed was that the graph of the variable (particularly Horse Power) was similar to the Z-Scores for Horse Power.

Thanks,

Garima Singh

Factor Analysis

Factor Analysis

Description

When trying to explain something via measuring a range of independent variables, factor analysis helps reduce the number of reported variables by determining significant variables and 'combining' these into a a single variable. It may be used in this way either to discover factors or to test a hypothesis that they exist.

The determination is statistical, but the output is a 'virtual' or unobserved variable which combines the measured variables in a linear formula which combines observed measurements with 'factor loading' constant numbers.

Factor Analysis also can be used to help demonstrate how a complex measurement instrument is really measuring one or a few bigger things.

Example

Perhaps the most well-known result of factor analysis is 'IQ', which is a 'virtual' variable based on variable measurements of ability in mathematics, language and logic.

Discussion

Factor Analysis originated in psychometrics by Charles Spearman (as in Spearman correlation) and Raymond Cattell (as in 16PF) with the assessment of intelligence and personality, and has since spread to many other fields, from marketing to operations research.

Factor Analysis is different to much research, which focuses on the relationships between independent and dependent variables. In contrast, Factor Analysis focuses on the relationship between multiple independent variables. 'Factor' basically means 'independent variable', although in this case the 'factors' are the new 'virtual' variables.

The use of Factor Analysis and its results is often an imprecise science and hence subject to debate (for example 'IQ' has been heavily criticized). Nevertheless it provides a simplification process that in practice can be very useful.

Principal component analysis

Principal Component Analysis is a variant of Factor Analysis and is equivalent when model 'errors' have the same variance. The difference lies in how Principal Component Analysis uses the total variance in the data and assume linear variable combinations, whilst Common Factor Analysis uses the common variance in the data and assumes latent variables.

Principal Component Analysis constructs as many components as there are original variables. The first component takes into account the greatest amount of variance between the variables, giving the weighting of these variables to form the single component. The second component is constructed to account for as much variance is left over that is not accounted for by the first component. The third component accounts for variance not accounted for by the first two components, and so on.

Normally, the first component is much larger than the rest, with a rapid drop off through the second and third components. If there is no useful way of reducing the matrix of variable correlations into a smaller number of factors then all components will be approximately equal.

Eigenvalues

Eigenvalues represent the proportion of variance explained by a given variable. With five variables, the sum of the eigenvalues will be 5. Sorting the factors by eigenvalue thus results in the first factor having the greatest importance (explaining the greatest amount of variance).

Eigenvalues can be used to help select the factors to select and carry forward for future use. Ways of doing this include:

· Greater than one: Select variables where the eigenvalue is >1.

· Fixed number: Select the top N variables only.

· Proportion explained: Select a proportion of variance that you want to explain, say 95%, then select those variables that fit within this.

· A Scree Plot is a simple sorted graph of Eigenvalues and can be used to select variables by taking only those on the steep part of the curve.

Orthogonal rotation

Factors selected through eigenvalue analysis may be refined further through orthogonal rotation. 'Orthogonal' factors are metaphorically at 'right angles' to one another, which means they do not correlate with one another and are so even more independent, making them useful measures.

For example if you were measuring types of intelligence, it would help if, say, creative intelligence was completely different from mathematical intelligence, so if you were working on one it would not 'leak' over into another area.



Vivek Kumar Singh

MBA – Operations Group 11

Discriminant Function Analysis


Discriminant Function Analysis

PURPOSE

The main purpose of a discriminant function analysis is to predict group membership based on a linear combination of the interval variables. The procedure begins with a set of observations where both group membership and the values of the interval variables are known. The end result of the procedure is a model that allows prediction of group membership when only the interval variables are known. A second purpose of discriminant function analysis is an understanding of the data set, as a careful examination of the prediction model that results from the procedure can give insight into the relationship between group membership and the variables used to predict group membership.

EXAMPLES

For example, a graduate admissions committee might divide a set of past graduate students into two groups: students who finished the program in five years or less and those who did not. Discriminant function analysis could be used to predict successful completion of the graduate program based on GRE score and undergraduate grade point average. Examination of the prediction model might provide insights into how each predictor individually and in combination predicted completion or non-completion of a graduate program.

Another example might predict whether patients recovered from a coma or not based on combinations of demographic and treatment variables. The predictor variables might include age, sex, general health, time between incident and arrival at hospital, various interventions, etc. In this case the creation of the prediction model would allow a medical practitioner to assess the chance of recovery based on observed variables. The prediction model might also give insight into how the variables interact in predicting recovery.

The Simplest Case

The simplest case of discriminant function analysis is the prediction of dichotomous group membership based on a single variable. An example of the simplest case is the prediction of successful completion of a graduate program based on the GRE verbal score. In this case, since the prediction model includes only a single variable, it gives little insight into how variables interact with each other in prediction. Thus prediction of group membership will be the major focus of the next section of this chapter.

With respect to the data file and purpose of analysis, this simplest case is identical to the case of linear regression with dichotomous dependent variables. As discussed previously, data of this type may be represented in any number of different forms: scatterplots, tables of means and standard deviations, and overlapping frequency polygons. Because overlapping frequency polygons have such an intuitive appeal, they will be used to describe how discriminant function analysis works.

Prediction Accuracy

A single interval variable might discriminate between groups in an almost perfect fashion, not at all, or somewhere in between. For example, if one wished to differentiate adult males and females, one could collect information on how many bras the person owned, score on the last statistics test, and height. In the case of the number of bras, the discrimination would be very good, but not perfect (some women don't own any bras, some men do). In the case of the score on the last statistics test, little discrimination would be possible because males and females generally score about the same. In the case of height, some discrimination between adult males and females would be possible, but it would be far from perfect.

In general, the larger the difference between the means of the two groups relative to the within groups variability, the better the discrimination between the groups. The following program allows the student to explore data sets with different degrees of discrimination ability.

Navneet Kumar

MBA -operations Group 11

Major Clustering Methods Explained with it Physical Significance in Statistics

Blog By: Milind Suryavanshi_Operations_Group 11_ (Roll No. 12031)


Major Clustering Methods Explained with it Physical Significance in Statistics

Note: Please refer to reference on images/diagrams.

A Categorization of Major Clustering Methods

  • Partitioning method. It classifies the data into k groups, which together satisfy the following requirements: (1) each group must contain at least one object, and (2) each object must belong to exactly one group. Like the algorithms k-means and k-medoids. The cons of it is, most partitioning methods cluster objects based on the distance between objects. Such methods can find only spherical-shaped clusters and encounter difficulty at discovering clusters of arbitrary shapes.
  • Hierarchical methods. A hierarchical method creates a hierarchical decomposition of the given set of data objects. Hierarchical methods suffer from the fact that once a step (merge or split) is done, it can never be undone.
  • Density-based methods. The general idea is to continue growing the given cluster as long as the density (number of objects or data points) in the “neighborhood” exceeds some threshold; that is, for each data point within a given cluster, the neighborhood of a given radius has to contain at least a minimum number of points. Such a method can be used to filter out noise (outliers) and discover clusters of arbitrary shape.
  • Grid-based methods. Grid-based methods quantize the object space into a finite number of cells that form a grid structure. The main advantage of this approach is its fast processing time, which is typically independent of the number of data objects and dependent only on the number of cells in each dimension in the quantized space.
  • Model-based methods. Model-based methods hypothesize a model for each of the clusters and find the best fit of the data to the given model. A model-based algorithm may locate clusters by constructing a density function that reflects the spatial distribution of the data points.
  • Other special method.
    • Clustering high-dimensional data
    • Constraint-based clustering

Partitioning Methods


Given D, a data set of n objects, and k, the number of clusters to form, a partitioning algorithm organizes the objects into k partitions (k ≤ n), where each partition represents a cluster. The usually used methods are k-Means and k-Medoids.

K-Means


The k-means algorithm takes the input parameter, k, and partitions a set of n objects into k clusters so that the resulting intracluster similarity is high but the intercluster similarity is low. Cluster similarity is measured in regard to the mean value of the objects in a cluster, which can be viewed as the cluster’s centroid or center of gravity.


It proceeds as follows:

1 Randomly selects k of the objects, each of which initially represents a cluster mean or center.
2 For each of the remaining objects, an object is assigned to the cluster to which it is the most similar, based on the distance between the object and the cluster mean.
3 It then computes the new mean for each cluster. This process iterates until the criterion function converges. Typically, the square-error criterion is used, defined as

where E is the sum of the square error for all objects in the data set; p is the point in space representing a given object; and mi is the mean of cluster Ci (both p and m_i are multidimensional). In other words, for each object in each cluster, the distance from the object to its cluster center is squared, and the distances are summed.
4 Iterate this process until no change happens.

The pros of this method:

  • It works well when the clusters are compact clouds that are rather well separated from one another.
  • The method is relatively scalable and efficient in processing large data sets because the computational complexity of the algorithm is O(nkt), where n is the total number of objects, k is the number of clusters, and t is the number of iterations.
  • The method often terminates at a local optimum

The cons of this method:

  • It can be applied only when the mean of a cluster is defined. This may not be the case in some applications.
  • It need users to specify k.
  • It is not suitable for discovering clusters with nonconvex shapes or clusters of very different size.
  • It is sensitive to noise and outlier data points because a small number of such data can substantially influence the mean value.
  • Often terminates at local optimal

Hierarchical Methods

Hierarchical clustering methods can be further classified as either agglomerative or divisive, depending on whether the hierarchical decomposition is formed in a bottom-up (merging) or top-down (splitting) fashion. The quality of a pure hierarchical clustering method suffers fromits inability to performadjustment once amerge or split decision has been executed. That is, if a particular merge or split decision later turns out to have been a poor choice, the method cannot backtrack and correct it.

Four widely used measures for distance between clusters are as follows, where |p-p'| is the distance between two objects or points, p and p'; m_ is the mean for cluster,C_i; and n_i is the number of objects in C_i.

Note: Refer to Image at the reference provided

One promising direction for improving the clustering quality of hierarchical methods is to integrate hierarchical clustering with other clustering techniques, resulting in multiple-phase clustering.

1. BIRCH

BIRCH is designed for clustering a large amount of numerical data by integration of hierarchical clustering (at the initial microclustering stage) and other clustering methods such as iterative partitioning (at the later macroclustering stage). It overcomes the two difficulties of agglomerative clustering methods: (1) scalability and (2) the inability to undo what was done in the previous step.

2. Rock

ROCK (RObust Clustering using linKs) is a hierarchical clustering algorithm that explores the concept of links (the number of common neighbors between two objects) for data with categorical attributes.

Two distinct clusters may have a few points or outliers that are close; therefore, relying on the similarity between points to make clustering decisions could cause the two clusters to be merged. ROCK takes a

more global approach to clustering by considering the neighborhoods of individual pairs of points. If two similar points also have similar neighborhoods, then the two points likely belong to the same cluster and so can be merged.

3. Chameleon

Chameleon is a hierarchical clustering algorithm that uses dynamic modeling to determine the similarity between pairs of clusters. In Chameleon, cluster similarity is assessed based on how well-connected objects are within a cluster and on the proximity of clusters. That is, two clusters are merged if their interconnectivity is high and they are close together.

Chameleon has been shown to have greater power at discovering arbitrarily shaped clusters of high quality than several well-known algorithms such as BIRCH and densitybased DBSCAN. However, the processing cost for high-dimensional data may require O(n2) time for n objects in the worst case.

4. Density-Based Methods

To discover clusters with arbitrary shape, density-based clustering methods have been developed. These typically regard clusters as dense regions of objects in the data space that are separated by regions of low density (representing noise).

5. DBSCAN

DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a density based clustering algorithm. The algorithm grows regions with sufficiently high density into clusters and discovers clusters of arbitrary shape in spatial databases with noise. It defines a cluster as a maximal set of density-connected points.

A density-based cluster is a set of density-connected objects that is maximal with respect to density-reachability. Every object not contained in any cluster is considered to be noise.

This method is sensitive to its parameter e and MinPts, and leaves the user with the responsibility of selecting parameter values that will lead to the discovery of acceptable clusters.

If a spatial index is used, the computational complexity of DBSCAN is O(nlogn), where n is the number of database objects. Otherwise, it is O(n2).

6. OPTICS

OPTICS computes an augmented cluster ordering for automatic and interactive cluster analysis.

Because of the structural equivalence of the OPTICS algorithm to DBSCAN, the OPTICS algorithm has the same runtime complexity as that of DBSCAN, that is, O(nlogn) if a spatial index is used, where n is the number of objects.

7. DENCLUE

DENCLUE (DENsity-based CLUstEring) is a clustering method based on a set of density distribution functions. The method is built on the following ideas: (1) the influence of each data point can be formally modeled using a mathematical function, called an influence function, which describes the impact of a data point within its neighborhood; (2) the overall density of the data space can be modeled analytically as the sum of the influence function applied to all data points; and (3) clusters can then be determined mathematically by identifying density attractors, where density attractors are local maxima of the overall density function.

The major advantages DENCLUE has are

  • It has a solid mathematical foundation and generalizes various clustering methods, including partitioning, hierarchical, and density-based methods;
  • It has good clustering properties for data sets with large amounts of noise;
  • It allows a compact mathematical description of arbitrarily shaped clusters in highdimensional data sets;
  • It uses grid cells, yet only keeps information about grid cells that actually contain data points. It manages these cells in a tree-based access structure, and thus is significantly faster than some influential algorithms, such as DBSCAN.

However, the method requires careful selection of the density parameter and noise threshold, as the selection of such parameters may significantly influence the quality of the clustering results.

8. Grid-Based Methods

The grid-based clustering approach uses a multiresolution grid data structure. It quantizes the object space into a finite number of cells that form a grid structure on which all of the operations for clustering are performed. The main advantage of the approach is its fast processing time, which is typically independent of the number of data objects, yet dependent on only the number of cells in each dimension in the quantized space.

9. STING

STING is a grid-based multiresolution clustering technique in which the spatial area is divided into rectangular cells. There are usually several levels of such rectangular cells corresponding to different levels of resolution, and these cells form a hierarchical structure: each cell at a high level is partitioned to form a number of cells at the next lower level. Statistical information regarding the attributes in each grid cell (such as the mean, maximum, and minimum values).

STING offers several advantages: (1) the grid-based computation is query-independent, because the statistical information stored in each cell represents the summary information of the data in the grid cell, independent of the query; (2) the grid structure facilitates parallel processing and incremental updating; and (3) the method’s efficiency is a major advantage: STING goes through the database once to compute the statistical parameters of the cells, and hence the time complexity of generating clusters is O(n), where n is the total number of objects. After generating the hierarchical structure, the query processing time is O(g), where g is the total number of grid cells at the lowest level, which is usually much smaller than n.

10. WaveCluster

WaveCluster is a multiresolution clustering algorithm that first summarizes the data by imposing a multidimensional grid structure onto the data space. It then uses a wavelet transformation to transformthe original feature space, finding dense regions in the transformed space.

11. Model-Based Clustering Methods

Model-based clustering methods attempt to optimize the fit between the given data and some mathematical model. Such methods are often based on the assumption that the data are generated by a mixture of underlying probability distributions.

12. Conceptual Clustering

Conceptual clustering is a form of clustering in machine learning that, given a set of unlabeled objects, produces a classification scheme over the objects. Unlike conventional clustering, which primarily identifies groups of like objects, conceptual clustering goes one step further by also finding characteristic descriptions for each group, where each group represents a concept or class. Hence, conceptual clustering is a two-step process: clustering is performed first, followed by characterization. Here, clustering quality is not solely a function of the individual objects. Rather, it incorporates factors such as thegenerality and simplicity of the derived concept descriptions.

Most methods of conceptual clustering adopt a statistical approach that uses probability measurements in determining the concepts or clusters. Probabilistic descriptions are typically used to represent each derived concept.

13. Clustering High-Dimensional Data

Most clusteringmethods are designed for clustering low-dimensional data and encounter challenges when the dimensionality of the data grows really high. This is because when the dimensionality increases, usually only a small number of dimensions are relevant to certain clusters, but data in the irrelevant dimensions may producemuch noise and mask the real clusters to be discovered.Moreover, when dimensionality increases, data usually become increasingly sparse because the data points are likely located in different dimensional subspaces.

To overcome this difficulty, we may consider using feature (or attribute) transformation and feature (or attribute) selection techniques. Feature transformation methods, such as principal component analysis and singular value decomposition, transformthe data onto a smaller space while generally preserving the original relative distance between objects. They summarize data by creating linear combinations of the attributes, and may discover hidden structures in the data. However, such techniques do not actually remove any of the original attributes from analysis. This is problematic when there are a large number of irrelevant attributes. The irrelevant information may mask the real clusters, even after transformation. Moreover, the transformed features (attributes) are often difficult to interpret, making the clustering results less useful. Thus, feature transformation is only suited to data sets where most of the dimensions are relevant to the clustering task. Unfortunately, real-world data sets tend to have many highly correlated, or redundant, dimensions.

Another way of tackling the curse of dimensionality is to try to remove some of the dimensions. Attribute subset selection (or feature subset selection7) is commonly used for data reduction by removing irrelevant or redundant dimensions (or attributes). Given a set of attributes, attribute subset selection finds the subset of attributes that are most relevant to the data mining task. Attribute subset selection involves searching through various attribute subsets and evaluating these subsets using certain criteria. It is most commonly performed by supervised learning—the most relevant set of attributes are found with respect to the given class labels. It can also be performed by an unsupervised process, such as entropy analysis, which is based on the property that entropy tends to be low for data that contain tight clusters.

Other evaluation functions, such as category utility, may also be used. Subspace clustering is an extension to attribute subset selection that has shown its strength at high-dimensional clustering. It is based on the observation that different subspaces may contain different, meaningful clusters. Subspace clustering searches for groups of clusters within different subspaces of the same data set. The problem becomes how to find such subspace clusters effectively and efficiently.

14. Constraint-Based Cluster Analysis


Users often have a clear view of the application requirements, which they would ideally like to use to guide the clustering process and influence the clustering results. Thus, in many applications, it is desirable to have the clustering process take user preferences and constraints into consideration.

15. Outlier Analysis

One person’s noise could be another person’s signal. Outlier detection and analysis are very useful for fraud detection, customized marketing, medical analysis, and many other tasks. Computer-based outlier analysis methods typically followeither a statistical distribution-based approach, a distance-based approach, a density-based local outlier detection approach, or a deviation-based approach.

Reference: https://sites.google.com/a/kingofat.com/wiki/data-mining/cluster-analysis#TOC-Hierarchical-Methods

Blog By: Milind Suryavanshi_Operations_Group 11_ (Roll No. 12031)