Search This Blog

4.2.11

Insight about Discriminant Analysis

The main purpose of a discriminant function analysis is to predict group membership based on a linear combination of the interval variables. The procedure begins with a set of observations where both group membership and the values of the interval variables are known. The end result of the procedure is a model that allows prediction of group membership when only the interval variables are known. A second purpose of discriminant function analysis is an understanding of the data set, as a careful examination of the prediction model that results from the procedure can give insight into the relationship between group membership and the variables used to predict group membership.

Discriminant analysis is a statistical method that is used by researchers to help them understand the relationship between a "dependent variable" and one or more "independent variables." A dependent variable is the variable that a researcher is trying to explain or predict from the values of the independent variables. Discriminant analysis is similar to regression analysis and analysis of variance (ANOVA). The principal difference between discriminant analysis and the other two methods is with regard to the nature of the dependent variable.

Discriminant analysis requires the researcher to have measures of the dependent variable and all of the independent variables for a large number of cases. In regression analysis and ANOVA, the dependent variable must be a "continuous variable." A numeric variable indicates the degree to which a subject possesses some characteristic, so that the higher the value of the variable, the greater the level of the characteristic. A good example of a continuous variable is a person's income.

Discriminant analysis is most often used to help a researcher predict the group or category to which a subject belongs. For example, when individuals are interviewed for a job, managers will not know for sure how job candidates will perform on the job if hired. Suppose, however, that a human resource manager has a list of current employees who have been classified into two groups: "high performers" and "low performers." These individuals have been working for the company for some time, have been evaluated by their supervisors, and are known to fall into one of these two mutually exclusive categories. The manager also has information on the employees' backgrounds: educational attainment, prior work experience, participation in training programs, work attitude measures, personality characteristics, and so forth. This information was known at the time these employees were hired. The manager wants to be able to predict, with some confidence, which future job candidates are high performers and which are not. A researcher or consultant can use discriminant analysis, along with existing data, to help in this task.

There are two basic steps in discriminant analysis. The first involves estimating coefficients, or weighting factors, that can be applied to the known characteristics of job candidates (i.e., the independent variables) to calculate some measure of their tendency or propensity to become high performers. This measure is called a "discriminant function." Second, this information can then be used to develop a decision rule that specifies some cut-off value for predicting which job candidates are likely to become high performers.

Then the discriminant function has the following form:

where D = discriminant function score,
B = discriminant function coefficient relating independent variable i to the discriminant function score,
X = value of independent variable i.

The equation is quite similar to a regression equation. Conventional regression analysis should not be used in place of discriminant analysis. The dependent variable would have only two values (high performer and low performer) and would thus violate important assumptions of the regression model. Discriminant analysis does not have these limitations with respect to the dependent variable.

Submitted by-
Prashant Vaish
12060
SIBM, Bangalore

Factor Analysis - 360 Degree View ( Karanbir Singh , 12141 , Marketing)


INTRODUCTION


Factor analysis is a statistical technique to study the inter-relationships among the variables in an effort to find a new set of factors, fewer in number than the original variables so that the factors are common among the original variables. In factor analysis a small number of common factors are extracted so that these common factors are sufficient to study the relationships of original variables.

In many real-life applications, the number of independent variables used in predicting a response variable will be too many. The difficulties in having too many independent variables in such exercise are as follows,

•Increased computational time to get solution.
• Increased time in data collection.
• Too much expenditure in data collection.
•Presence of redundant independent variables.
•Difficulty in making inferences.
These can be avoided using factor analysis.

AIMS OF FACTOR ANALYSIS

• Factor analysis helps the researcher to reduce the number of variables to be analyzed, thereby making the analysis easier.
• Analysis based on a wide range of variables can be tedious and time consuming.
• For example, consider a market researcher at a credit card company who wants to evaluate the credit card usage and behaviour of customers, using various variables. The variables include age, gender, marital status, income level, education, employment status, credit history and family background.
•Using Factor Analysis, the researcher can reduce the large number of variables into a few dimensions called factors that summarize the available data.
• Its aims at grouping the original input variables into factors which underlying the input variables.
• For example, age, gender, marital status can be combined under a factor called demographic characteristics. The income level, education, employment status can be combined under a factor called socio-economic status. The credit card and family background can be combined under factor called background status.

BENEFITS OF FACTOR ANALYSIS

¨ To identify the hidden dimensions or construct which may not be apparent from direct analysis

¨ To identify relationships between variables

¨ It helps in data reduction

¨ It helps the researcher to cluster the product and population being analyzed.

TERMINOLOGY IN FACTOR ANALYSIS

•· Factor: A factor is an underlying construct or dimensions that represent a set of observed variables. In the credit card company example, the demographic characteristics, socio economic status and background status represent a set of variables.
•· Factor Loadings: Factor loading help in interpreting and labeling the factors. It measures how closely the variables in the factor are associated. It is also called factor-variable correlation. Factor loadings are correlation coefficients between the variables and the factors.
•· Eigen Values: Eigen values measure the variance in all the variables corresponding to the factor. Eigen values are calculated by adding the squares of factor loading of all the variables in the factor. It aid in explaining the importance of the factor with respect to variables. Generally factors with Eigen values more than 1.0 are considered stable. The factors that have low Eigen values (<1.0) may not explain the variance in the variables related to that factor.
•· Communalities: Communalities, denoted by h2, measure the percentage of variance in each variable explained by the factors extracted. It ranges from 0 to 1. A high communality value indicates that the maximum amount of the variance in the variable is explained by the factors extracted from the factor analysis.
•· Total Variance explained: The total variance explained is the percentage of total variance of the variables explained. This is calculating by adding all the communality values of each variable and dividing it by the number of variables.
•· Factor Variance explained: The factor variance explained is the percentage of total variance of the variables explained by the factors. This is calculating by adding the squared factor loadings of all the variables and dividing it by the number of variables.


PROCEDURE FOLLOWED FOR FACTOR ANALYSIS

• Define the problem
• Construct the correlation matrix that measures the relationship between the factors and the variables.
•Select an appropriate factor analysis method
•Determine the number of factors
•Rotation of factors
•Interpret the factors
•Determine the factor scores


APPLICATION AREAS

Factor analysis is by far the most often used multivariate technique of research studies, specially pertaining to social and behavioral science.

It is a technique applicable when there is a systematic interdependence among a set of observed or manifest variables and the researcher is finding out something more fundamental or latent which creates this commonality.



Example

Consider the problem of studying the customer's feedback about a two-wheeler produced by a company as explained below;

The marketing manager of a two-wheeler company designed a questionnaire to study the customer's feedback about its two-wheeler and in turn he is keen in identifying the factors of his study. He has identified six variables, which are:

1.Fuel Efficiency (X1)
2. Life of the Two-Wheeler (X2)
3.Handling Convenience (X3)
4.Quality of Original Spares (X4)
5.Breakdown Rate (X5)
6.Price (X6)

MEHTODS OF FACTOR ANALYSIS Centroid Method of Factor Analysis

This method of factor analysis, developed by L.L Thurstone, was quite frequently used until about 1950 before the advent of large capacity high-speed computers. The centroid method tends to maximize the sum of loadings, disregarding signs; it is method, which extracts the largest sum of absolute loadings for each factor in turn. It is defined by linear combinations in which all weights are either + 1.0 or -1.0. The main merit of this method is that it is relatively simple, can be easily understood and involves simpler computations. If one understands this method, it becomes easy to understand the mechanics involved in other method of factor analysis.

Principal Component Method

Principal-components method (or simply P.C. method) of factor analysis, developed by H.Hotelling, seeks to maximize the sum of squared loadings of each factor extracted in turn. Accordingly PC factor explains more variance than would the loadings obtained from any other method of factoring.

The aim of the principal components method is the construction out of a given set of variables X's (j=1, 2, 3… k) of new variables (pi) called principal components, which are linear combinations of the Xs.

ROTATION IN FACTOR ANALYSIS

One often talks about the rotated solutions in the context of factor analysis. This is done (i.e., a factor matrix is subjected to rotation) to attain what is technically called "simple structure" in data. Simple structure according to L.L Thurstone is obtained by rotating the axes** until:

(i) Each row of the factor matrix has one zero.

(ii) Each column of the factor matrix has p zeros, where p is the number of factors.

(iii) For each pair of factors, there are several variables for which the loadings on one is virtually zero and the loading on the other is substantial

(iv) If there are many factors, then for each pair of factors there are many variables for which both loadings are zero.

(v) For every pair of factors, the number of variables with non-vanishing loadings on both of them is small.

All these criteria simply that the factor analysis should reduce the complexity of all the variables.

R-TYPE AND Q-TYPE FACTOR ANALYSES

Factor analysis may be R-type factor analysis or it may be Q-type factor analysis. In R-type factor analysis, high correlations occur when respondents who score high on variable1 also score high on variable2 and respondents who score low on variable1 also score low on variable2. Factors emerge when there are high correlations within groups of variables.

In Q-type factor analysis, the correlations are computed between pairs of respondents instead of pairs of variables. High correlations occur when respondent 1's pattern of responses on all the variables is much like respondent 2's pattern of responses. Factor emerges when there are high correlations within groups of people. Q-type analysis is useful when the object is to sort out people into groups based on their simultaneous responses to all the variables.

Factor analysis is has been mainly used in developing psychological tests (such as IQ tests, personality tests, and the like) in the realm of psychology. In marketing, this technique has been used to look at media readership profiles of people.

MERITS

The main merits of factor analysis can be stated thus:

(i) The techniques of the factor analysis are quite useful when we want to condense and simplify the multivariate data.

(ii) The technique is helpful in pointing out important and interesting, relationship among observed data that were there all the time, but not easy to see from the data alone.

(iii) The technique can reveal the latent factors (i.e. underlying factors not directly observed) that determine relationships among several variable concerning a research study. For example, if people are asked to rate different cold drinks (say Limca, Nova-cola, and Gold Spot and so on) according to preference, a factor analysis may reveal some salient characteristics of cold drinks that underline the relative preferences.

(iv) The technique may be used in the context of empirical clustering of products, media or people i.e. providing a classification scheme when data scored on various rating scales have to be grouped together.

LIMITATION


One should also be aware of several limitations of factor analysis. Important ones are as follows:

(i) Factor analysis, like all multivariate techniques, involved laborious computations involving heavy cost burden.

(ii) The single factor analysis are considered generally less reliable and dependable for very often a factor analysis starts with a set of imperfect data. "the factors are nothing but blurred averages, difficult to be identified". To overcome this difficulty, it has been realized that analysis should at least be done twice. If we get more or less similar results from all rounds of analyses, our confidence concerning such results increases.

iii) Factor analysis is a complicated decision tool that can be used only when one has thorough knowledge and enough experience of handling this tool. Even then, at times it may not work well and may even disappoint the user.

To conclude, we can state that in spite of all the said limitations "When it works well, factor analysis help the investigator makes sense of large bodies of interwined data. When it works unusually well, it also points out some interesting relationships that might not have been obvious from examination of the input data alone".



CONCLUSION

Thus Factor analysis is an interdependence technique. The complete sets of interdependent relationships are examined. There is no specification of dependent variables, independent variables, or causality. Factor analysis assumes that all the rating data on different attributes can be reduced down to a few important dimensions. This reduction is possible because the attributes are related. The rating given to any one attribute is partially the result of the influence of other attributes.

Perceptual Maps – A tool that helps you get ahead of your competitor: - K.Sailesh, Roll No 12026

Perceptual Mapping is a data analysis tool that is used to represent the varying relationships that exist between one or many marketplace competitors.The mapping attempts to measure the factors that consumers use in making their purchase decisions. Using a perceptual map has become a popular tool for data analysts to use when briefing upper management on complex issues because it can represent complex relationships graphically.

Now that you have some idea on what exactly this wonderful tool is all about let me make your life simple by giving a 4 step process on how to use perceptual maps

Step1 : Select a product or service and pick three or more companies that produce the material or provide the service

Step 2: Take a piece of paper and draw one vertical line that splits the paper vertically and another line drawn horizontally which also splits the paper so that they visually look like a 'plus sign'. Now there are software like Permap to do the same, but anyways lets continue:

Step 3: Create two questions about the product or service to ask consumers. The question should directly ask questions about how company a's product compares to company b's to company c's. A minimum of six people should be asked the questions. The vertical axis of the paper drawn on in step two will represent one of the two questions and the horizontal axis the other question. This represents the perceptual map.

Step 4:

Plot the answers to the two questions on the perceptual map labeling the company's name where the consumer's answer to the question best fits the perceptual map questions.
For example, if you want to compare Ford and Cheverolet, a question might be, do you find Ford or Chevrolet vehicles more sporty? The top of the vertical axis would be labeled sporty and the bottom labeled less sporty. You would label the top of the axis Ford with the number of answers that were that manufacturer and put the number of Chevrolet answers in the applicable spot on the perceptual map.
That way you can put characteristics in each of the 4 quads and based on that you can come up with a decision on your product and beat the thinking of the competitors.
Many companies do perceptual mapping. Ive had the chance to interact with some companies and I have learnt from them that it is a very valuable tool.

Finance Business Intelligence – From Data to Information to Competitive Advantage.

Today's competitive and challenging economic environment places further emphasis on real-time information for decision-making purposes. Without relevant, reliable information provided in a timely manner, decision-makers are forced into suboptimal decisions based on assumptions and gut instinct. This lack of fact-based decision making places the company at risk with one or more constituencies that interact with the company on a daily basis –­ the implications of which are, in aggregate, substantial and negative to the enterprise.

The time for business intelligence (BI) is now, as it can be leveraged as a competitive weapon to improve efficiencies and better manage position in the value chain. If implemented as a comprehensive solution, BI will reduce the time spent on low- to non-value-added activities across the enterprise, as well as enhance the organization's ability to make profit-enhancing decisions.

Benefit targets and categories include:

Advertisement

  • 20 to 30 percent reduction in period-end closing activities through application automation and process redesign. Reconciliation activities can be reduced by as much as 50 to 60 percent as well.
  • 30 to 40 percent reduction in report generation activities (which we have seen account for as much as 30 percent of the total finance organization). Report sets are rationalized and standardized, and the user communities are enabled to access and produce their analyses without ongoing finance/information technology (IT) support.
  • 50 to 60 percent reduction in budgeting, planning and forecasting activities across the enterprise. The multiple process streams involving manual spreadsheets, reconciliation, rekeying and e-mails can be substantially reduced and create the opportunity to collapse the entire planning cycle.
  • Revenue enhancement, including the information to improve gross-to-net management, tracking of sales float from sales programs, customer penetration and trending.
  • Cost management by delivering accuracy in reporting major areas of spend to support benchmarking and/or other cost reduction initiatives.
  • Profitability management by viewing contribution at the customer, channel and SKU (stock-keeping unit) levels, thereby supporting the creation of profit plans to influence margins.
  • Avoidance of misrepresentations in external reporting and related market cap impacts, as well as the ability to build investor confidence through more fact-based support of financial performance.

The Business Intelligence Pyramid

We view business intelligence through an informational/technical "pyramid" containing the following elements: legacy systems, enterprise resource planning (ERP) solutions, data warehouse/marts, analytic applications and enterprise information system (EIS)/key performance indicator (KPI) portal (see Figure 1).


Figure 1: Business Intelligence Pyramid

As you work your way up the pyramid, the level of information correspondingly becomes more summarized and the views of information become more strategic.

Legacy Systems

The underlying legacy systems are used to support core processes and information capture (e.g., work management, time capture, etc.). These systems may require replacement over time, but are primarily viewed as a constant in the framework.

ERP Solutions

ERP software packages drive transaction processing efficiency and consistency. A company may have multiple packages, instances and/or configurations depending upon its operating model. Organizations have begun the shift from "implement" to "optimize" in this area, increasing enterprise-wide consistency in process and data definition.

Data Warehouse/Data Marts

This layer provides for the efficient extraction of data from multiple sources, the translation of disparate data into a common, rational data model and the loading of data into the analytic applications. The data structures in this layer are optimal for extracting, transforming and loading large volumes of data values from the legacy and ERP systems. The resulting data is ready to be analyzed by power users within the organizations or further manipulated within the analytic application layer.

The "bypass" concept within the pyramid represents an important decision for any BI implementation effort. In the interest of quick solution development, the data warehousing/marts layer is often overlooked as an unnecessary step. Our experience has proven, however, that consideration must be given to this decision up front to avoid the creation of a suboptimal or even unusable BI application. Furthermore, there are significant cost and timing ramifications associated with inserting an ex post facto data warehouse layer within a BI framework.

Analytic Applications

Analytic applications are purpose-built solutions meant to satisfy a specific analytic/decision-making set of requirements. These applications could be viewed as a series of workbenches designed to provide a given group within the enterprise the tools necessary to do their jobs. Specifically within the finance function, these include financial reporting, financial planning and financial analysis (see Figure 2).


Figure 2: Analytic Applications within the Finance Function

Financial Reporting

This component could be considered the required blocking and tackling owned by the finance organization. It includes core period-end activities, legal entity/SEC (Securities and Exchange Commission), tax and other regulatory reporting requirements. The opportunities here are to aggressively automate the topside adjustments and related activities, as well as standardize/automate the data submission, reconciliation and core reporting processes. By so doing, period- end windows will collapse, lower value-added activities will be eliminated and finance will have more time to truly review the numbers as well as ensure information quality.

Financial Planning

Financial planning covers the entire planning process across the enterprise, which finance typically owns from a submission, preparation and reporting perspective. In many organizations, the planning cycle has become an exercise in futility. The budget window can span eight-plus months and is plagued by sandbagging and stretch target setting. Forecasts become a manual nightmare that may consume as much as 25 percent of the finance organization's capacity each month, yet don't provide valid predictability to future earnings per share (EPS) or other finance metrics.

Financial Analysis

This is the component that enables business partnering with the organization, supporting customer, product and channel profitability analysis, predictive EVA/SVA value reporting and process-based performance and benchmarking. These analytics provide finance with multiple views of the business and create the environment to convert operational performance to financial results.

This analytic capability, however, can only be attained if finance organizational capacity is created through efficiency gains and automation in the other two components.

EIS/KPI Portal

The EIS/KPI portal is the decision-maker's desktop view and window to the various information sets and underlying analytics. This provides the enterprise with a single means of access to the information and related reporting. This means that the user can navigate through various analytics, ERP solutions and even underlying legacy systems through the same portal, thereby eliminating the need to toggle and determine what data is where within the framework.

Historically, internal Web portal efforts have often been plagued by content focused on employee self-service regarding individual human resources benefit plans, as a vehicle for mass communications, or at best, an executive-level dashboard. Today, they have evolved into a far more effective enabler for a given user to navigate the internal technological/informational landscape and complete various types of analysis.

Factors to Ensure Success

Following are key steps that can mean the difference between the success and failure of the BI endeavor.

  1. Start with the end in mind.

Lack of a clear, comprehensive BI strategy dramatically increases the implementation team's risk of failure. Dedicate anywhere from 6 to 12 weeks to complete the following strategic/vision-oriented components of a BI solution prior to detailed design and implementation.

    • Establish a clear set of conceptual requirements across multiple constituencies/functions within the organization.
    • Overtly apply the organization's strategic objectives to the business intelligence solution and communicate with key stakeholders.
    • Develop a detailed, resource-loaded and phased implementation plan.
    • Establish a solid business case and present same to executive management.
    • Assess the technical landscape, develop the technical road map, establish the short list of vendor products and develop the selection criteria process.
    • Identify the primary areas of incongruity that will require heavy analysis and decision making during the design phase (e.g., specific information standards).
    • Develop a strategy to manage the personal and organizational changes that will follow during implementation.

Without these steps, any solution risks falling into the trap of being labeled either an accounting project or another system implementation.

  1. Prioritize and incrementalize.

The implementation requirements surrounding BI solutions are inherently different from their ERP counterparts in that the implementations are far less time- and resource-consuming. This creates the ability for an organization to rapidly provide solutions for various organizational communities. However, this rapid deployment must be executed in a way that ties the core data elements of the organization together in an efficient manner.

Once the overall BI vision for the company has been articulated, the plan of attack should be comprised of component implementations that are prioritized by the degree of information pain being experienced within different areas of the organization. Implementations are recommended to hold a 90-day benchmark and be limited in scope to a manageable array of information pain (the amount of pain addressed within each 90-day scope should increase as the team becomes more experienced with the tools and processes). The benefits of an incremental approach to BI include:

    • Focusing on specific component information pains allows the implementation team to adequately address the complex functional and technological pieces of each implementation. When the entire organization is approached at once, the complexity of the components becomes unmanageable and the work unproductive.
    • 90-day windows provide the team with a means to demonstrate the value being created to the various constituencies on an incremental basis.
    • Momentum is maintained for the enterprise-wide deployment through a series of successes.
  1. Monitor and publish success.

Implementation teams can fall prey to viewing a phased solution as complete once the application set is live and the data is converted. The team must come back to the original business case, check performance to expectations and publish examples that have created value to the organization. These might include:

    • Making a profit-enhancing decision during a major customer negotiation based upon the improved information.
    • Supporting the analysis of a capital request.
    • What-if analysis decisions regarding where to make and what to make.
    • SKU rationalization/life-cycle profitability.
    • Predictive economic analysis using the improved tools and information.
    • Pure timeliness/capacity improvements in major areas such as financial closing, management reporting and budgeting.

While many companies have implemented some form of a BI solution, few have established organization-wide awareness of these capabilities, benefits and/or potential for future value- creation that these applications possess. By first confirming and then communicating these wins to the organization on at least a quarterly basis, the BI solution will become further indoctrinated into the organization and convert stakeholders from dreading to demanding additional BI applications for their respective areas.

Maximize the Return

In summary, remember the following:

  • Track your progress. Start with the end in mind. Don't let your organization fall prey to viewing BI as a software installation.
  • Begin smart by leveraging existing research, best practices, industry trends and diagnostics. Recognize and leverage the fact that you are not the first to travel the BI path.
  • Focus on a complete solution, but in a way that continually drives 90-day results. Assemble a strong cross-functional team that will produce a solid business case and build the trust and commitment from the organization overall.

Today, the technology has caught up. It won't be the software that prevents a company from maximizing the return on a BI solution, but rather the content, processes and people necessary to make the solution work.

Link for this: http://www.information-management.com/issues/20031201/7733-1.html

Submitted by;

Kichawele Diwani Msuya

MBA-Finance (2009-2011)

Roll No. 12157.

3.2.11

factor analysis

On 25th January we had our last workshop on Business intelligence. We dealt with two topics discriminant analysis and factor analysis.
Discriminant analysis is a technique for classifying a set of observations into predefined classes. The purpose is to determine the class of an observation based on a set of variables known as predictors or input variables. the technique constructs a set of linear functions of the predictors, known as discriminant functions, such that
L = b1x1 + b2x2 + ? + bnxn + c , where the b's are discriminant coefficients, the x's are the input variables or predictors and c is a constant.
These discriminant functions are used to predict the class of a new observation with unknown class. For a k class problem k discriminant functions are constructed. Given a new observation, all the k discriminant functions are evaluated and the observation is assigned to class i if the ith discriminant function has the highest value.

Factor analysis is a data reduction technique that tries to reduce a list of attributes or other measures to their essence; that is, a smaller set of “factors”that capture the paterns seen in the data. Marketers and researchers who study a product, service, or industry professionally sometimes perceive many more distinctions within their category than do their consumers. This can lead to questionnaires containing attribute lists that consumers see as somewhat or largely synonymous. Factor analysis tells you how many different core factors that consumers perceived out of the list of attributes thay rated.
The main benefits of factor analysis are that the analyst can focus their attention on the unique core elements instead of the redundant attributes, and as a data ‘pre-processor’for regression models.

Submitted by,
Pragya Mishra
12153
SIBM Bangalore

Cluster analysis Usage in various fields....... By Ajay Amarnath A


It is used mainly for software project planning control and management, an accurate estimate of software development cost is important. Past research has focused on using parametric models to predict development cost. The integration a neural network method with cluster analysis to estimate development cost.
Clustering is an economic development model signifying growth of similar kinds of industries at one geographical location. Locating near other similar firms provides numerous competitive advantages, including sharing a common labor pool, enhancing close working relationships between firms, reducing transaction costs and travel times between customers and suppliers, and enhancing the spread of technology through firms in the region.
As a cluster in a region takes root and expands, synergies often develop between firms and institutions, spurring additional growth and innovation. The existence of demand centre and concentration of Service Providers around the cluster also contributes to the growth of the cluster in terms of number of units. Other stakeholders like consultants, equipment manufacturers, Government Agencies etc also get concentrated in the cluster.
A cluster analysis was then performed to identify aspects of low, medium, and high risk projects. An examination of risk dimensions across the levels revealed that even low risk projects have a high level of complexity risk. For high risk projects, the risks associated with requirements, planning and control and the organization become more obvious. The influence of project scope, sourcing practices, and strategic orientation on project risk dimensions was also examined. Results suggested that project scope affects all dimensions of risk, whereas sourcing practices and strategic orientation had a more limited impact.
In marketing, cluster analysis is used for segmenting the market and determining target markets. Product positioning and New Product Development Selecting test markets the basic procedure. Formulate the problem - select the variables that you wish to apply the clustering technique.

How to solve human machine interface problem...... By KATHIRESHAN R


I have doubts regarding the human machine interface of the tool. i.e.
while we are evaluating the null hypothesis we are checking the significance value in the chi-square table and come to a conclusion whether to accept or reject the hypothesis. But the problem is, the chi-square value remains same and the way we define our hypothesis plays an important role. It can clearly be explained by 2 scenarios.

Scenario 1:
Null Hypothesis: Suppose there is no relationship between males who are educated are early married.
We found the significance value to be 0.05 from the chi-square table and hence we come to a conclusion that there is a relationship between males who are educated and early marriage.

Scenario 2:
Null Hypothesis: Suppose if we define there is relationships between male who are educated are early married.
Since the significance value which we obtained from the chi-square table is not changed then it implies that there is no relationships between male who are educated are early married.

These two scenarios contradict each other.

How is this possible and how to solve the human machine interface problem?And how to overcome the assumption made by the software? Is there a better way of defining the hypothesis in the tool?

28.1.11

Discriminant Analysis: By Vaibhav Khaparde


Discriminant Analysis (DA) undertakes the same task as multiple linear regressions by predicting an outcome. However, multiple linear regressions is limited to cases where the dependent variable on the Y axis is an interval variable so that the combination of predictors will, through the regression equation, produce estimated mean population numerical Y values for given values of weighted combinations of X values. But many interesting variables are categorical, such as political party voting intention, migrant/non-migrant status, making a profit or not, holding a particular credit card, owning, renting or paying a mortgage for a house, employed/unemployed, satisfied versus dissatisfied employees, which customers are likely to buy a product or not buy, what distinguishes Stellar Bean clients from
Gloria Beans clients whether a person is a credit risk or not, etc.
DA is used when:
The dependent is categorical with the predictor at interval level such as age, income, attitudes, perceptions, and years of education, although dummy variables can be used as predictors as in multiple regression. Logistic regression can be of any level of measurement. There are more than two DV categories, unlike logistic regression, which is limited to a dichotomous dependent variable.

The major underlying assumptions of DA are:

The observations are a random sample;

Each predictor variable is normally distributed;
Each of the allocations for the dependent categories in the initial classification are correctly classified;
There must be at least two groups or categories, with each case belonging to only one group so that the groups are mutually exclusive and collectively exhaustive (all cases can be placed in a group);

Each group or category must be well defined, clearly differentiated from any other group(s) and natural. Putting a median split on an attitude scale is not a natural way to form groups. Partitioning quantitative variables is only justifiable if there are easily identifiable gaps at the points of division;

For instance, three groups taking three available levels of amounts of housing loan; the groups or categories should be defined before collecting the data;

The attribute(s) used to separate the groups should discriminate quite clearly between the groups so that group or category overlap is clearly non-existent or minimal;

Group sizes of the dependent should not be grossly different and should be at least five times the number of independent variables.




There are several purposes of DA:

To investigate differences between groups on the basis of the attributes of the cases, indicating which attributes contribute most to group separation. The descriptive technique successively identifies the linear combination of attributes known as canonical discriminant functions (equations) which contribute maximally to group separation.
Predictive DA addresses the question of how to assign new cases to groups. The DA function uses a person’s scores on the predictor variables to predict the category to which the individual belongs.

To determine the most parsimonious way to distinguish between groups.

To classify cases into groups. Statistical significance tests using chi square enable you to see how well the function separates the groups.

To test theory whether cases are classified as predicted.

Methods implemented in this area are Multiple Discriminant Analysis, Fisher's Linear Discriminant Analysis, and K-Nearest Neighbors’ Discriminant Analysis.
Multiple Discriminant Analysis
(MDA) is also termed Discriminant Factor Analysis and Canonical Discriminant Analysis. It adopts a similar perspective to PCA: the rows of the data matrix to be examined constitute points in a multidimensional space, as also do the group mean vectors. Discriminating axes are determined in this space, in such a way that optimal separation of the predefined groups is attained. As with PCA, the problem becomes mathematically the eigen reduction of a real, symmetric matrix. The eigen values represent the discriminating power of the associated eigenvectors. The nYgroups lie in a space of dimension at most nY - 1. This will be the number of discriminant axes or factors obtainable in the most common practical case when n > m > nY(where n is the number of rows, and m the number of columns of the input data matrix).
Linear Discriminant Analysis
is the 2-group case of MDA. It optimally separates two groups, using the Mahalanobis metric or generalized distanceIt also gives the same linear separating decision surface as Bayesian maximum likelihood discrimination in the case of equal class covariance matrices.
K-NNs Discriminant Analysis
: Non-parametric (distribution-free) methods dispense with the need for assumptions regarding the probability density function. They have become very popular especially in the image processing area. The K-NNs method assigns an object of unknown affiliation to the group to which the majority of its K nearest neighbors’ belongs.

Submitted By:

Vaibhav Khaparde
12054
Marketing





Factor Analysis : Introduction

Factor analysis is a collection of methods used to examine how underlying constructs influence the responses on a number of measured variables.

There are basically two types of factor analysis: exploratory and confirmatory..

1>Exploratory factor analysis (EFA) attempts to discover the nature of the constructs influencing a set of responses.

2>Confirmatory factor analysis (CFA) tests whether a specified set of constructs is influencing responses in a predicted way.

Both types of factor analyses are based on the Common Factor Model, illustrated in figure 1.1. This model proposes that each observed response (measure 1 through measure 5) is influenced partially by underlying common factors (factor 1 and factor 2) and partially by underlying unique factors (E1 through E5). The strength of the link between each factor and each measure varies, such that a given factor influences some measures more than others. This is the same basic model as is used for LISREL analyses.

Factor analyses are performed by examining the pattern of correlations between the observed measures. Measures that are highly correlated (either positively or negatively) are likely influenced by the same factors, while those that are relatively uncorrelated are likely influenced by different factors.

Exploratory Factor Analysis:

The primary objectives of an EFA are to determine

1. The number of common factors influencing a set of measures.

2. The strength of the relationship between each factor and each observed measure.

Some common uses of EFA are to

Ø Identify the nature of the constructs underlying responses in a speci¯c content area.

Ø Determine what sets of items \hang together" in a questionnaire.

Ø Demonstrate the dimensionality of a measurement scale. Researchers often wish to develop scales that respond to a single characteristic.

Ø Determine what features are most important when classifying a group of items.

Ø Generate \factor scores" representing values of the underlying constructs for use in other analyses.

Miscellaneous notes on EFA:

1>To have acceptable reliability in your parameter estimates it is best to have data from at least 10 subjects for every measured variable in the model. This number should be increased if you expect that the influence of the common factors is relatively weak. You should also have measurements from at least three variables for every factor that you want to include in your model.

2> You should endeavour to have a wide variety of measurements for your EFA. The more accurately that your selection of measurements properly represents the population" of measurements that could be taken, the more generality you will have in your findings.

3> EFA can be performed in SAS using proc factor. Principal component analysis can be performed in SAS using proc princomp, while it can be performed in SPSS using the Analyze/Data reduction/Factor analysis menu selection. EFA cannot actually be performed in SPSS (despite the name of menu item used to perform PCA).

confirmatory factor analysis

Confirmatory factor analysis (CFA) is a special form of factor analysis. It is used to test whether measures of a construct are consistent with a researcher's understanding of the nature of that construct (or factor).

CFA is commonly used in social research. CFA is frequently used when developing a test, such as a personality test, intelligence test, or survey. CFA is also frequently used as a first step to assess the proposed measurement model in a structural equation model. Many of the rules of interpretation regarding assessment of model fit and model modification in structural equation modeling apply equally to CFA. CFA is distinguished from structural equation modeling by the fact that in CFA, there are no directed arrows between latent factors.[clarification needed] In the context of SEM ,the CFA often is called 'the measurement model', while the relations between the latent variables (with directed arrows) are called 'the structural model'


Source: Wikipedia, Asparouhov, T.;Muthén, B. (2009). "Exploratory structural equation modeling". Structural Equation Modeling, 16, 397-438., http://www.stat-help.com/factor.pdf


Blog By: Akhilesh Agarwal (Operations_SIBM)

Factor analysis titus raju 12053

Factor analysis
Factor analysis attempts to discover the nature of the constructs influencing a set of responses.
Both types of factor analyses are based on the Common Factor Model where model proposes that each observed response (measure 1 through measure 5) is influenced partially by underlying common factors (factor 1 and factor 2) and partially by underlying unique factors (E1 through E5). The strength of the link between each factor and each measure varies, such that a given factor influences some measures more than others.
Factor analysis helps you find these undetermined variables by looking at the variables you have actually collected.

An important element of the factor analysis output is the standardized factor score coefficients (Output-Table), which gives location of each product on each factor.
The vectors are obtained based on the amount of correlation the original attitudes possess with the factor scores (represented as factors). The direction of the vectors indicates the factor with which each attribute is associated, and the length of the vector indicates the strength of association. Thus, on the left map the “filling” attribute has little association with any factor, whereas on the right map the “filling” attribute is strongly associated with “refreshing” factor.Although a factor is not observable like the other original variables, it is still a variable. One output of most factor analysis programs is the values for each factor for all respondents. These values are termed factor scores and are shown above for three factors that were found to underlie the five input variables. Thus, each beverage has a factor score on each factor, in addition to the beverage’s rating on the original eight attributes. In subsequent stages while doing perceptual mapping this factor scores are used to position beverages in the perceptual map.

My learning from class:
Factor analysis is basically what is done when you have a large number of variables to work with. Such a large number of variables makes it very difficult to organize data, and analyze it. Thus, we can use factor analysis, to reduce the number of variables so that it becomes much easier to work.
Factor analysis uses correlation between variables to see which ones are related, and can be eliminated, without causing significant impact on the output of analysis.
For this purpose, we take a default eigen value >1 and use verimax rotation so that the first few variables have the maximum effect.
Using Rotated Component Matrix, we can very well determine which variables are to be combined into factors, and keeping a limit on the eigen values to be accepted, we can determine what percentage of the data is contained in the factors we are choosing. According to this, we can increase or decrease the number of factors to simplify our calculations.
We can look at the correlation between the variables in scatter plot, in which visually clustered patterns can be made out and combined.
The remaining factors can be combined to give more meaning to the analysis.



Regards
Titus Raju