Discriminant analysis is a technique for analysis data when the criterion or dependent variable is categorical in nature and predictor or independent variables are interval in nature.
The objectives of discriminant analysis are as follows
- Developments of discriminant function, which will best discriminate between the categories of the criterion or dependent variable (groups).
- Examination of whether significant difference exists among the groups, in terms of predictor variables
- Determination of whether the predicator variables contribute to most of the inter-group differences
- Classification of cases to one of the groups based on the values of the predictor variables
- Evaluation of accuracy of classification
Discriminant analysis techniques are described by the number of categories possessed by the criterion variable. When the criterion variable has two categories, the technique is called two group discriminant analysis. When three or more categories are involved, the technique is referred to as multiple discriminant analysis.
Discriminant analysis model:
The discriminant analysis model involves linear combination of the following form
L = b1x1 + b2x2 + …. + bnxn + c , where the b's are discriminant coefficients, the x's are the input variables or predictors, L is the discriminant score and c is a constant
The coefficients or weights (b) are estimated so that the groups differ as much as possible on the values of discriminant function. This occurs when the ration of between group sum of square to with-in group sum of squares for the discriminant function is at the maximum. Any other linear combination will result in smaller ratio.
I am giving a brief geometrical exposition of two group discriminant analysis. In the figure 1 we have two groups, A and B, and each member of these groups is measured on two variables X and Y which are the two axes. Members of A are denoted in ellipse A and members of B are denoted by ellipse B. the resultant ellipse encompasses some specified percentage of members in each group. A straight line is drawn through the two points where the ellipse intersects and then projected to new axes, I. The overlap between the univariate distributions A and B is smaller than would be obtained by any other line drawn through the ellipse representing the members of the groups. Thus the group differs maximum on I axes.
Some important statistics associated with discriminant analysis
- Canonical Correlation: it is the extent of association between the discriminant score and the groups.
- Classification matrix: it contains the number of correctly classified and misclassified cases.
- Discriminant function coefficient: these are the multipliers of the variables, when the variables are in the original units of measurement
- Eigen value: it is the ratio of between groups to with in group sum of squares. Larger eigen value means superior function
Use of discriminant analysis
Discriminant analysis is used in marketing at a very great extent in distinguishing the consumer behavior, store characteristics….etc.
But apart from being intensively in marketing discriminant analysis is also used in areas like finance and operations. I will give one example from operations where discriminant analysis is used effectively.
By using the discriminant analysis function within statistical software packages, managers can classify future production volume and work to separate defective products on the production line before they are delivered to customers. This study was conducted by David Lengacher.
Discriminant Analysis in Action
The following example outlines the creation of a data set and the use of discriminant analysis. A director of operations has been plagued with the same problem for months. The process is highly capable and internal scrap is low, but roughly 3 percent of the shipped products fail when the customer installs them into an assembly. Although the parts are within customer specifications, they are still failing. It may take months for a detailed chemical analysis to provide the director with a true root cause, but even then, there are no guarantees. However, upon graphing the last batch of customer returns against a random sample of parts ready for shipment, a clear pattern emerges (Figure 1).
It is much cheaper to scrap production inside the facility than it is to be charged for the customer’s complete assembly that failed because of a defective product. Therefore, the director’s goal is to be able to separate those units that are likely to fail as soon as they come off the production line. |
At first glance, however, it appears that the situation is hopeless. Customer returns and normal parts are overlapping in the scatter plot. In situations like this, statistical analysis software can be of great help. By using software with a discriminant analysis feature, practitioners can separate production with a high degree of accuracy and minimal cost.
The software will analyze the full data set and produce two group membership functions. These equations are constructed in a way that will minimize the misclassification of the sample data set. As shown in Figure 2, the functions are able to achieve a 91 percent accuracy level. Because each data point goes through both membership functions, practitioners can assign the data point to the function that returns the higher value. This solution can mitigate risk until the director of operations can identify the true root cause of the failures.
Thus the example stated above shows how discriminant analysis can be used in making better decisions.
References
1. Malhotra, N.K. (2007), “Marketing research an applied orientation.” 5th edition.
2. David Langacher, “Discriminant Analysis Can Minimize Returned Products”, Available online at http://www.isixsigma.com/index.php?option=com_k2&view=item&id=709:discriminant-analysis-can-minimize-returned-products&Itemid=204#
3. Wikipedia
By VAIBHAV JAISWAL
By VAIBHAV JAISWAL



No comments:
Post a Comment