Search This Blog

24.1.11

CLUSTER ANALYSIS by vaibhav jaiswal


Cluster Analysis:
Cluster Analysis is a multivariate analysis technique that seeks to organize information about variables so that relatively homogeneous groups, or "clusters," can be formed. The clusters formed with this family of methods should be highly internally homogenous (members are similar to one another) and highly externally heterogeneous (members are not like members of other clusters.
In simple terms Cluster analysis or clustering is the assignment of a set of observations into subsets (called clusters) so that observations in the same cluster are similar in some sense.
Types of clustering:
Basically there are 2 types of clustering
  1. Hierarchical Clustering
  2. K-Means Clustering

Hierarchical clustering creates a hierarchy of clusters which may be represented in a tree structure called a dendrogram. Hierarchical algorithm find successive clusters using previously established clusters. They are either agglomerative or divisive. Agglomerative starts with each object as a separate cluster and then combine them into successively larger clusters. Divisive starts with the whole set and proceed to divide it into successively smaller clusters.
The main outcome of a cluster analysis is a dendrogram, which is also called a tree diagram.
Distance Measurement:
An important step in most clustering is to select a distance measure, which will determine how the similarity of two elements is calculated. This will influence the shape of the clusters, as some elements may be close to one another according to one distance and farther away according to another. The most commonly used distance function is the Euclidean distance which gives the shortest distance between two objects and is measured in a straight line.
Other commonly used methods are:
City block/Manhattan distance: It is the sum of absolute differences in values for each variable.
Chebychev distance: It is the maximum absolute difference in values for any variable.
Also the best way to measure the distance between two clusters or an object and a cluster is between groups method. This method calculates the distance by considering the distances between each pair of objects in two clusters.

How to determine the number of clusters to be formed?
The number of clusters to be formed is achieved through the cut-off point which is the point at which the next object joins the cluster at a much larger distance than the distance at which the previous object joined the cluster.

Statistics associated with cluster analysis:
Agglomeration schedule: An as gives information on the objects and cases being combined at each stage of hierarchical clustering process.
Cluster centroid: it is the mean value of the variable for all the cases or objects in a particular cluster centres. The cluster centres are the initial starting points in non hierarchical clustering.
Cluster membership: it indicates the cluster to which each object or case belongs.
Dendrogram: it is a graphical device for displaying the clustering results. Vertical line represents clusters that are joined together. The position if the line on the scale represents the distance at which the clusters are joined.
Icicle diagram: it is a graphical display of clustering results so called because it resembles the row of icicles hanging from the eaves of house. The columns correspond to the objects being clustered and the row corresponds to the no. of clusters.
Classification of clustering procedure:



Uses of Cluster Analysis in Marketing: Cluster analysis has been used in marketing for a variety of purposes.
  1. Segmenting the market: for e.g. customer may be clustered on the basis of benefits got from the purchase of the product. Each cluster would consist of consumers who are relatively homogeneous in terms of benefits they seek. This approach is called benefit segmentation
  2. Understanding buyer behaviour: CA can be used to identify homogeneous groups of buyers and then the buying behaviour of each group can be examined separately
  3. Identifying new product opportunities: by clustering brands and products, competitive sets within the market can be determined. Brands in the same cluster compete more fiercely with each other than with brands in other clusters. A firm can examine its current offering compared to those of its competitors to identify potential new product opportunity.
  4. Reducing data: it can be used as a general data reduction tool to develop clusters of data that are more manageable than individual observations.
Some Statistical software used to do cluster analysis:
1)      SPSS: In SPSS the main program for hierarchical clustering of objects or cases is hierarchical clusters. For non hierarchical the K-means cluster program can be used. This program is particularly helpful for clustering a large no. of cases.
2)      SAS:  In SAS the CLUSTER program can be used for the hierarchical clustering. Non hierarchical clustering can be accomplished using SASTCLUS. For clustering of variables the VARCLUS program is used.
3)      MINITAB: In MINITAB cluster analysis can be assessed in the multi variate > cluster observation function.
Cluster analysis is not available in Excel.

Submitted by:
Vaibhav Jaiswal
Roll No: 12109

SPSS – To cut a long story SHORT

SPSS- It is a statistical analysis and data management Software package. SPSS can take data from almost any type of file and use them to generate tabulated reports, charts, and plots of distributions and trends, descriptive statistics, and conduct complex statistical analyses.

Before start working on the actual software, it is necessary to go through some basics of the statistics involved. Basics of the statistics include terminologies like- Sample size, selection, mean, variance, standard deviation moreover it also includes frequency, cross tabulation, cluster analysis, chi square analysis, which we discussed in class.

First Day of the SPSS classes focused on some basic fundamental and how to use the SPSS software. When you open the software, two different windows can be opened while using SPSS,

The Data Editor –It is a spreadsheet in which you define your variables and enter data. Each row corresponds to a case while each column represents a variable. The title bar displays the name of the open data file or "Untitled" if the file has not yet been saved.

The Output Navigator- It displays the statistical results, tables, and charts from the analysis you performed. An Output Navigator window opens automatically when you run a procedure that generates output. In the Output Navigator windows, you can edit, move, delete and copy your results in a Microsoft Explorer-like environment.

Cluster Analysis-

Cluster analysis classifies a set of observations into two or more mutually exclusive unknown groups based on combinations of interval variables. The purpose of cluster analysis is to discover a system of organizing observations, usually people, into group, where members of the groups share properties in common. It is cognitively easier for people to predict behavior or properties of people or objects based on group membership, all of whom share similar properties. It is generally cognitively difficult to deal with individuals and predict behavior or properties based on observations of other behaviors or properties. The most important techniques for data collection are:

1. Cluster Analysis.

2. Discriminant Analysis.

Although both cluster and discriminant analysis classifies objects into categories, discriminant analysis requires one to know group membership for the cases used to decide the classification rule whereas in cluster analysis group membership for all the cases is unknown. In addition to membership, the number of groups is also generally unknown.

Steps in Cluster Analysis-

a) Select a measure of similarity.

b) Decision is to be made on the type of clustering technique to be used.

c) Type of clustering method for the selected technique is selected.

d) Decision regarding the number of cluster.

e) Cluster solution is interpreted.

Clustering Techniques in Wireless Sensor Networks-

Due to the communication devices on sensor nodes have limited battery capacity and transmission range, wireless sensor networks (WSNs) are considered to be energy constrained. In this paper, we propose a novel clustering algorithm called limiting member node clustering (LmC) algorithm to limit the number of member nodes for each cluster head by using a threshold value. The proposed clustering approach selects a cluster head based on a new cost function which considers the residual battery level, energy consumption and distance to the base station. In our experiments, we considered the transmission range of the base station in WSNs to improve the clustering performance. From experimental results, the proposed algorithm can efficiently achieve high number of successfully delivered packets as well as the longest network lifetime while give the shortest delay time and low energy consumption when compared with different existing algorithms.

Submitted By-

Prashant Vaish
Roll no: 12060
Finance Batch
SIBM, Bangalore

Cluster Analysis in Supply Chain Optimization

Cluster Analysis in Supply Chain Optimization

Cluster Analysis is a versatile tool for grouping objects based on their similarity to each other. It is applied in a wide variety of situations. Some example applications are similarity of faces, countries, and species of dogs. One starts by defining attributes of the individuals (countries, faces, ( length of nose, separation of eyes. One of several available metrics is used to define the ‘distance’ between the individuals. Individuals are then assigned CA)species…) that can be assigned a numerical value. CA produces some appealing results and gives useful insights that spark further investigation even in the exotic applications.

In the relatively straightforward context of grouping locations, cluster analysis performs equally well. This makes CA a useful tool in supply chain optimization. Designing a distribution network often involves planning of routes over regions or deciding on locations for warehouses. One could look at a map covered with dots representing locations and draw circles around groups. But this would be an arbitrary approach. CA offers a way to group locations in a systematic way and speeds up the process of exploring several different versions of the clusters. I am going to present a few examples and describe the mechanics of the CA process.

In the examples that follow we have a set of 400+ locations from an area surrounding Denver Colorado. The latitude and longitude of the locations are the attributes that will define the similarity of locations. In this case the distance between objects and clusters really is a distance. In CA you can generally use Euclidean distance or XY distance – as though you were on a sidewalk in a city. Which to use depends on the situation. If looking at a downtown area, the XY distance might be best. But over a broader area the Euclidean or Great Circle distance makes more sense.

There are really two broad categories of clustering – hierarchical and non-hierarchical. The hierarchical methods begin by considering all the individuals (locations) as separate groups. The nearest two are joined and now we have one less group. The location of that group can be defined in a few different ways. For now picture it as the centroid of the group members – which could be weighted by something like demand for a product. The next join depends on the remaining distances between members or between a member and the existing clusters. In either case we end up with one less group. CA keeps joining until there is one large group. This is agglomerative or hierarchical clustering.

Along the way there is a point at which we have 434, 433… 5, 4, 3, 2, 1 groups. For any given number of groups we can look at the distance between the groups and the distance of group members from the centroid of their own group. One could look for a small number of clusters where the distances within a cluster are small compared to the distances between clusters.

Most of the options within hierarchical clustering are based on the way we define the distance between an individual location and a cluster and the distance from one cluster to another. One example is nearest neighbor. In this case the distance of an individual to a cluster is the distance to the closest existing location in the cluster. Farthest neighbor is also a possibility. Neither of these is well-suited to most supply chain applications. The typical method is to use the centroid as the cluster location for defining distances. Let’s take a look at three methods, average distance, farthest neighbor, and Ward’s method (which will be described soon).

http://www.profitpt.com/wp-content/uploads/2010/11/New-Picture-300x200.png

Figure 1 – Clustering based on farthest neighbor (Complete Clustering)

In the complete method an individual or cluster is joined to another cluster based on the smallest maximum distance between any two individual locations. Figure 1 above is what it produces. If the groups were to be used to define clusters for delivery routes one problem might be the unequal sizes of the five groups that were created.

The average distance method joins individuals or clusters to another cluster based on the smallest average distance to the members of a cluster. This gives another result with unequal cluster membership. A method based on the median distance gives a result very similar to the complete method using the sample data, though it is less sensitive to outlying locations that are far from all others.

If reasonably equal group sizes are important, then another method called Ward’s method is one of the best choices. The results using Ward’s method are shown below. Ward’s method makes assignments that minimize the within clusters deviations from the centroid of the cluster. What you get depends on how many clusters you ask for. You can compare the results of five and six clusters.

http://www.profitpt.com/wp-content/uploads/2010/11/New-Picture-1-300x200.png

Figure 2 – Clustering based on average distance to the existing cluster members.

http://www.profitpt.com/wp-content/uploads/2010/11/New-Picture-2-300x200.png

Figure 3 – Ward’s Method with five clusters requested.

http://www.profitpt.com/wp-content/uploads/2010/11/New-Picture-3-300x200.png

Figure 4 – Ward’s Method with six clusters requested.

Another way to look at the results is with a type of graph called a dendogram. The dendogram gives a visual look at the way similarity of clusters changes as individuals are merged into the clusters. The dendogram is shown in Figure 5 below. The vertical scale is just a measure of similarity. It is linear and is often just a normalized scale from zero to one hundred.

What you take away from the dendogram is this: as fewer clusters are formed, the similarity of members within a cluster decreases. Picture a horizontal line that cuts across the figure where it will intersect just six of the vertical dendogram lines. Notice that with just a small change in the vertical scale we could cross four or five vertical lines of the dendogram. That implies there is not much change going from four to six clusters in terms of the nearness of individuals within the clusters. They are still very similar (close). Things change much more rapidly going below four clusters. Of course, at the bottom when there are two hundred clusters, members of each cluster are very similar (read close) to each other.

The last method to look at is different than hierarchical clustering. It is called K Means clustering. It is not a hierarchical method so a dendogram is not really possible – joining does not occur sequentially. Without too many details, K Means clustering works like this. We start with an initial division into a certain number of clusters – let’s say six. The initial division should be done in some sensible way, e.g. Ward’s method.

http://www.profitpt.com/wp-content/uploads/2010/11/New-Picture-4-300x200.png

Figure 5 – Dendogram based on Ward’s method.

K-Means clustering keeps examining all the observations to see if they are closer to the centroid of some group than to the group they are currently in. If so, they are moved and the centroids of both affected groups are recalculated. The method continues until no more improving moves can be made. It is important to start with a decent initial grouping. The result of K-Means clustering using Ward’s method as a starting point is shown in Figure 6.

http://www.profitpt.com/wp-content/uploads/2010/11/New-Picture-5-300x200.png

Figure 6 – K Means Clustering with an initial group.

If we had started without an initial group the result would be a bit different. It would look like Figure 7.

http://www.profitpt.com/wp-content/uploads/2010/11/New-Picture-6-300x200.png

Figure 7 – K Means Clustering without an initial group.

This is somewhat different than the result with an initial group.

Bottom line recommendations are these. Use K-Means clustering with Ward’s method as an initial seed. If K-Means clustering is not available, use Ward’s method. If you use a hierarchical method take a look at the dendogram as a way to decide on a number of clusters. You can reduce the number of clusters until the reduction of similarity is large. Looking at the scattergram of the points for various numbers of clusters should confirm the dendogram information. Methods like the farthest or nearest neighbor will very likely give poor results. All methods are affected by outliers to some degree. Consider managing a few far out points by hand. Minitab is a very good software product for doing cluster analysis.

CA is far from a mathematical curiosity. It can be a very useful tool for network or facility location analysis. It certainly lets the analyst explore more solutions than could be done manually. Rather than some arbitrary decisions on grouping, CA can contribute the analysis that leads to bottom line savings.

Blog by- Milind Suryavanshi (12031_Operation)

Source:

1. www.profitpoint.com

Discriminant analysis - An efficient tool in Research – Oriented Applications

The blog is on the basis of inputs given on Discriminant analysis and its relevance in Perceptual Mapping in Market Research projects.

Starting off with an example about, how a manufacturer of a salon brand hair care item, could find out if demographic variables, like, education level, ethnicity, personal income, sex and a number of other factors are useful in distinguishing purchasers of their products from purchasers of other salon hair care brands.

Discriminant Analysis is an efficient market segmentation technique because of its ability to classify individuals or experimental units into two or more uniquely defined populations. The discriminant score is the basis for predicting to which group (a purchaser of the manufacturer’s brand or a competitive brand) the particular individual belongs. The discriminant weights of each predictive variable (age, sex, income, etc) indicate the relative importance of each variable. For instance, if age has a low discriminant weight then it is less important than the other variables.

Following the example, which illustrated how discriminant analysis helped classify users and nonusers of salon brand hair care products based on independent variables, other uses of discriminant analysis include the following:
Product research – Distinguish between heavy, medium, and light users of a product in terms of their consumption habits and lifestyles
Perception/Image research – Distinguish between customers who exhibit favorable perceptions of a store or company and those who do not
Advertising research – Identify how market segments differ in media consumption habitsDirect marketing – Identify the characteristics of consumers who will respond to a direct marketing campaign and those who will not.

Submitted By:
Veena Viswanath
Roll No: 12055

Business Intelligence and Supply Chain

Business Intelligence and Supply Chain Management

In the 1980's, finance and telecommunication companies pioneered BI to support financial and market analysis of the large volumes of data that they had begun to accumulate electronically. The need for BI capabilities grew in the 80's and 90's in other industries as companies began capturing data electronically across the full range of their business activities. This need was further compounded by the growing interest in real time data access which required effective tools to mine and analyze dramatically increased data volumes. To support this growing need, large software and services providers like IBM and Oracle launched major initiatives to bring data warehousing capabilities to the marketplace. These data warehouses, or data marts, are the most common sources of data for BI applications. ERP systems have also been used to capture data and enforce consistency, but they tend to be too inflexible to support ad hoc exploration of data. Fortunately, better tools for access and analysis have emerged. These tools usually start with flexible query and reporting capabilities that are combined with some mix of online analytical processing (OLAP), statistical analysis, forecasting and data mining techniques

Why Is BI Useful in Supply Chain Management?

BI use is expanding from finance to other business functions because it provides a quick Return on Investment (ROI). It complements supply chain planning because BI applications provide incremental benefits while a business lays the foundation for more sophisticated tools and related business process changes. To reap some quick returns and support their supply chain projects, some companies are using BI tools to:

1. Improve data visibility so as to reduce inventory levels by 5% to 15% in some businesses.

2. Analyze customer service levels to identify specific problem areas.

3. Better understand the sources of variability in customer demand to improve forecast accuracy.

4. Analyze production variability to identify where corrective measures need to be taken.

5. Analyze transport performance to reduce costs by using the most efficient transport providers.

By providing wider visibility to plans and supporting data, BI tools increase the return on existing SCP applications because they help companies understand where and how they deviate from their plan objectives. In addition, they provide shared data availability that encourages a global perspective on business performance. As a result, people are more likely to make decisions based on their global impact.

Some Current Developments

BI capabilities are now being integrated into other products. In fact, Microsoft has vowed to bring BI to the workforce with their next release of SQL Server. Not surprisingly, its ownership of the desktop work environment gives Microsoft an edge over Oracle and IBM, who have also announced enhancements to their OLAP and data mining capabilities. All this puts pressure on traditional BI applications providers. The CRM (Customer Relationship Management) community has even come up with a name for BI analysis of customer behavior (CRM analytics).

Although SCP vendors are offering BI capabilities as well by adding layers of products from the traditional BI vendors, the resulting mix of applications can be cumbersome to implement and support. In addition, IT publications report that end users often have difficulty using generic tools that were not designed to support specific roles or job functions. New SCP products have an advantage because BI capabilities are easily incorporated as part of a single offering. These newer tools can also be configured to support specific business roles. BI applications have become increasingly cost effective because they utilize the connectivity provided by the Internet and by intranets and because component-based software development speeds implementation. Since implementation consists of installing the software and connecting the data feeds, BI tools with good user interfaces can be put into use within a few weeks. Some companies that haven't developed these capabilities in an organized fashion are seeing independent, often underground, local projects popping up in their businesses. While it is encouraging to see employees take the initiative in addressing business problems, this fragmented approach often produces a series of applications with overlapping functionality drawing on a Hodge podge of different technologies. Sustaining these applications becomes a headache, and effective support often hinges on the continued presence of a local super user.

Some Implications for the future of Supply Chain Management Applications

1. BI applications will become part of the standard technology set used by most businesses and will have a synergistic effect on current and future SCP applications.

2. The continued evolution of component based software development will lead to increased consideration of internal development, particularly around BI applications.

3. The option of internal development will put downward pressure on software prices. Vendors will move from pricing based on estimates of value added (which are often optimistic) to pricing based on development costs

Milind Suryavanshi (12031_Operations)

Source:

1. Supply Chain Consultants

2. www.supplychain.com

Overview of cluster analysis and perceptual mapping and its use in pharmaceutical industry: 12104 sejal solanki

Overview of cluster analysis and perceptual mapping and its use in pharmaceutical industry:  12104 sejal solanki

Cluster analysis:
Cluster analysis is a collection of statistical methods, which identifies groups of samples that behave similarly or show similar characteristics. In common parlance it is also called look-a-like groups. The simplest mechanism is to partition the samples using measurements that capture similarity or distance between samples. In this way, clusters and groups are interchangeable words.
Typically in clustering methods, all the samples with in a cluster is considered to be equally belonging to the cluster (as against belonging with certain probability). If each observation has its unique probability of belonging to a group(cluster) and the application is interested more about these probabilities than we have to use (binomial) multinomial models.

Difference between cluster analysis and segmentation techniques:
Segmentation method could be interpreted as a collection of methods, which identifies groups of entities or statistical samples (consumers/customers, markets, organizations, which generally do not have a good application definition to classify entities in one of many groups without significant errors) that share certain common characteristics such as attitudes, purchase propensities, media habits, and lifestyle etc. The sample characteristics are used to group the samples. Grouping can be arrived at, either hierarchically partitioning the sample or non-hierarchically partitioning the samples. Thus, segmentation methods include probability-based grouping of observations and cluster (grouping) based observations. It includes hierarchical (tree based method – divisive) and non-hierarchical (agglomerative) methods. Segmentation methods are thus very general category of methodology, which includes clustering methods also.

The clustering algorithms are broadly classified into two namely hierarchical and non-hierarchical algorithms.

Hierarchical Clustering:
There is a concept of ordering involved in this approach. The ordering is driven by how many observations could be combined at a time or what determines that the distance is not statistically different from 0 between two observations or two clusters. The clusters could be arrived at either from weeding out dissimilar observations (divisive method) or joining together similar observations (agglomerative method).

Perceptual mapping:

It is Marketing research technique in which consumer's views about a product are traced or plotted (mapped) on a chart. Respondents are asked questions about their experience with the product in terms of its performance, packaging, price, size, etc. Theses qualitative answers are transferred to a chart (called a perceptual map) using a suitable scale (such as the Likert scale), and the results are employed in improving the product or in developing a new one. It allows senior marketing planners to take a broad view of the strengths and weaknesses of their product or service offerings relative to the strengths and weaknesses of their competition. It allows the marketing planner to view the customer and the competitor simultaneously in the same realm.

Application of Cluster Analyses in Pharmaceutical Lead Discovery
High throughput screening (HTS) programs based on diverse collections of compounds can rapidly identify leads for potential drug candidates. In cases where the compound collection is truly diverse, one may only identify a few compounds of interest. However, where a large number of hits are identified, it becomes necessary to examine the structures to determine the true number of compound classes involved so that follow-up studies may be conducted as efficiently as possible. In this case, cluster analysis is applied to determine the structural relationship among HTS hits. To efficiently expand around the region of the hit (or a class of hits) in chemical space, we have applied nearest neighbors analysis1 to select additional compounds from collections of a large number of commercial vendors, achieving an average hit rate in excess of 15%. Applying these techniques in a number of different cases, we obtained results that are useful for subsequent investigations of hits from HTS and other relevant molecular structures from the literature
(David T. Stanton, Timothy W. Morris, Siddhartha Roychoudhury, and Christian N. Parker*

At the end its very important to know why all this techniques are used and why such a large data analyzed following picture will clarify this issue……..

Hierarchial Mapping – An Overview -12169 - SwapneelShetkar

Cluster Analysis: My understanding of the concept
Clustering – grouping things in a particular form e.g customers in a segment
Types - Hierarchical and KMeans clustering
Hierarchical clustering: Done when objects to be clustered are less than 50. If objects more than 50 we use KMeans
 We have 2 types of hierarchical clustering
Divisive clustering: All objects are considered as one big group and then divided into sub clusters and keep dividing until we have one object as one group
Agglomerative clustering:  All objects as individual clusters and combine until we get one big cluster
SPSS Follows agglomerative cluster

How to cluster:
Selection of variables to cluster – All people from a specialisations in one group
Distance Measurement – Objects which are at less distance get clustered first, then the ones farther. E.g. we form a group of every third person in the class. In business applications we use Euclidian distance i.e. the shortest distance between 2 points.
Clustering Criteria- Distance measurement gives distance between two objects. Clustering gives the distance between two clusters or cluster and object. For e.g distance taken from the midpoint of the cluster. There are different methods used for clustering
·         Between group linkages
·         Nearest Neighbour
·         Furthest Neighbour
·         Centroid
Statistically, we cut off clustering when the next object to join, does so at a relatively higher distance.

Dendogram: Visual representation of how the clusters happen, which combine first and what combine later.





Literature Review
http://www.norusis.com/pdf/SPC_v13.pdf
Although both cluster analysis and discriminant analysis classify objects (or
cases) into categories, discriminant analysis requires you to know group membership for the cases used to derive the classification rule. The goal of cluster analysis is to identify the actual groups.
Identifying groups of individuals or objects that are similar to each other but different
from individuals in other groups can be intellectually satisfying, profitable, or
sometimes both. Using your customer base, you may be able to form clusters of
customers who have similar buying habits or demographics. You can take advantage
of these similarities to target offers to subgroups that are most likely to be receptive
to them. Based on scores on psychological inventories, you can cluster patients into
subgroups that have similar response patterns. This may help you in targeting
appropriate treatment and studying typologies of diseases. By analyzing the mineral
contents of excavated materials, you can study their origins and spread.

Hierarchical clustering is one of the most straightforward methods. It can be either agglomerative or divisive. Agglomerative hierarchical clustering begins with every case being a cluster unto itself. At successive steps, similar clusters are merged. The algorithm ends with everybody in one jolly, but useless, cluster. Divisive clustering starts with everybody in one cluster and ends up with everyone in individual clusters. Obviously, neither the first step nor the last step is a worthwhile solution with either method. In agglomerative clustering, once a cluster is formed, it cannot be split; it can only be combined with other clusters. Agglomerative hierarchical clustering doesn’t let cases separate from clusters that they’ve joined.