Search This Blog

26.1.11

Factor Analysis as a Classification Method : Minjal Desai – 12146, Marketing


Let us assume that we are at a point in our analysis where we basically know how many factors to extract. We may now want to know the meaning of the factors, that is, whether and how we can interpret them in a meaningful manner. To illustrate how this can be accomplished, let us work "backwards," that is, begin with a meaningful structure and then see how it is reflected in the results of a factor analysis. Let us return to our satisfaction example; shown below is the correlation matrix for items pertaining to satisfaction at work and items pertaining to satisfaction at home.

The work satisfaction items are highly correlated amongst themselves, and the home satisfaction items are highly intercorrelated amongst themselves. The correlations across these two types of items (work satisfaction items with home satisfaction items) is comparatively small. It thus seems that there are two relatively independent factors reflected in the correlation matrix, one related to satisfaction at work, the other related to satisfaction at home.

Factor Loadings. Let us now perform a principal components analysis and look at the two-factor solution. Specifically, let us look at the correlations between the variables and the two factors (or "new" variables), as they are extracted by default; these correlations are also called factor loadings.

Apparently, the first factor is generally more highly correlated with the variables than the second factor. This is to be expected because, as previously described, these factors are extracted successively and will account for less and less variance overall.

Rotating the Factor Structure. We could plot the factor loadings shown above in a scatterplot. In that plot, each variable is represented as a point. In this plot we could rotate the axes in any direction without changing the relative locations of the points to each other; however, the actual coordinates of the points, that is, the factor loadings would of course change. In this example, if you produce the plot it will be evident that if we were to rotate the axes by about 45 degrees we might attain a clear pattern of loadings identifying the work satisfaction items and the home satisfaction items.

Rotational strategies. There are various rotational strategies that have been proposed. The goal of all of these strategies is to obtain a clear pattern of loadings, that is, factors that are somehow clearly marked by high loadings for some variables and low loadings for others. This general pattern is also sometimes referred to as simple structure (a more formalized definition can be found in most standard textbooks). Typical rotational strategies arevarimax, quartimax, and equamax.

We want to find a rotation that maximizes the variance on the new axes; put another way, we want to obtain a pattern of loadings on each factor that is as diverse as possible, lending itself to easier interpretation. Below is the table of rotated factor loadings.

Interpreting the Factor Structure. Now the pattern is much clearer. As expected, the first factor is marked by high loadings on the work satisfaction items, the second factor is marked by high loadings on the home satisfaction items. We would thus conclude that satisfaction, as measured by our questionnaire, is composed of those two aspects; hence we have arrived at a classification of the variables.

Consider another example, this time with four additional Hobby/Misc variables added to our earlier example.

In the plot of factor loadings above, 10 variables were reduced to three specific factors, a work factor, a home factor and a hobby/misc. factor. Note that factor loadings for each factor are spread out over the values of the other two factors but are high for its own values. For example, the factor loadings for the hobby/misc variables (in green) have both high and low "work" and "home" values, but all four of these variables have high factor loadings on the "hobby/misc" factor.

Oblique Factors. Some authors (e.g., Cattell & Khanna; Harman, 1976; Jennrich & Sampson, 1966; Clarkson & Jennrich, 1988) have discussed in some detail the concept of oblique (non-orthogonal) factors, in order to achieve more interpretable simple structure. Specifically, computational strategies have been developed to rotate factors so as to best represent "clusters" of variables, without the constraint of orthogonality of factors. However, the oblique factors produced by such rotations are often not easily interpreted. To return to the example discussed above, suppose we would have included in the satisfaction questionnaire above four items that measured other, "miscellaneous" types of satisfaction. Let us assume that people's responses to those items were affected about equally by their satisfaction at home (Factor 1) and at work (Factor 2). An oblique rotation will likely produce two correlated factors with less-than- obvious meaning, that is, with many cross-loadings.

Hierarchical Factor Analysis. Instead of computing loadings for often difficult to interpret oblique factors, you can use a strategy first proposed by Thompson (1951) and Schmid and Leiman (1957), which has been elaborated and popularized in the detailed discussions by Wherry (1959, 1975, 1984). In this strategy, you first identify clusters of items and rotate axes through those clusters; next the correlations between those (oblique) factors is computed, and that correlation matrix of oblique factors is further factor-analyzed to yield a set of orthogonal factors that divide the variability in the items into that due to shared or common variance (secondary factors), and unique variance due to the clusters of similar variables (items) in the analysis (primary factors). To return to the example above, such a hierarchical analysis might yield the following factor loadings:

Careful examination of these loadings would lead to the following conclusions:
  1. There is a general (secondary) satisfaction factor that likely affects all types of satisfaction measured by the 10 items;
  2. There appear to be two primary unique areas of satisfaction that can best be described as satisfaction with work and satisfaction with home life.

Confirmatory Factor Analysis. Over the past 15 years, so-called confirmatory methods have become increasingly popular (e.g., see Jöreskog and Sörbom, 1979). In general, you can specify a priori, a pattern of factor loadings for a particular number of orthogonal or oblique factors, and then test whether the observed correlation matrix can be reproduced given these specifications. Confirmatory factor analyses can be performed via Structural Equation Modeling (SEPATH).

Miscellaneous Other Issues and Statistics

Factor Scores. We can estimate the actual values of individual cases (observations) for the factors. These factor scores are particularly useful when you want to perform further analyses involving the factors that you have identified in the factor analysis.

Reproduced and Residual Correlations. An additional check for the appropriateness of the respective number of factors that were extracted is to compute the correlation matrix that would result if those were indeed the only factors. That matrix is called the reproduced correlation matrix. To see how this matrix deviates from the observed correlation matrix, you can compute the difference between the two; that matrix is called the matrix of residual correlations. The residual matrix may point to "misfits," that is, to particular correlation coefficients that cannot be reproduced appropriately by the current number of factors.

Matrix Ill-conditioning. If, in the correlation matrix there are variables that are 100% redundant, then the inverse of the matrix cannot be computed. For example, if a variable is the sum of two other variables selected for the analysis, then the correlation matrix of those variables cannot be inverted, and the factor analysis can basically not be performed. In practice this happens when you are attempting to factor analyze a set of highly intercorrelated variables, as it, for example, sometimes occurs in correlational research with questionnaires. Then you can artificially lower all correlations in the correlation matrix by adding a small constant to the diagonal of the matrix, and then restandardizing it. This procedure will usually yield a matrix that now can be inverted and thus factor-analyzed; moreover, the factor patterns should not be affected by this procedure. However, note that the resulting estimates are not exact.

Principal Components and Factor Analysis - Srikrishnan V, 12165, Marketing

General Purpose

The main applications of factor analytic techniques are: (1) to reduce the number of variables and (2) to detect structure in the relationships between variables, that is to classify variables. Therefore, factor analysis is applied as a data reduction or structure detection method .

Basic Idea of Factor Analysis as a Data Reduction Method

Suppose we conducted a (rather "silly") study in which we measure 100 people's height in inches and centimeters. Thus, we would have two variables that measure height. If in future studies, we want to research, for example, the effect of different nutritional food supplements on height, would we continue to use both measures? Probably not; height is one characteristic of a person, regardless of how it is measured.

Let's now extrapolate from this "silly" study to something that you might actually do as a researcher. Suppose we want to measure people's satisfaction with their lives. We design a satisfaction questionnaire with various items; among other things we ask our subjects how satisfied they are with their hobbies (item 1) and how intensely they are pursuing a hobby (item 2). Most likely, the responses to the two items are highly correlated with each other. Given a high correlation between the two items, we can conclude that they are quite redundant.

Combining Two Variables into a Single Factor. You can summarize the correlation between two variables in a scatterplot. Regression line can then be fitted that represents the "best" summary of the linear relationship between the variables. If we could define a variable that would approximate the regression line in such a plot, then that variable would capture most of the "essence" of the two items. Subjects' single scores on that new factor, represented by the regression line, could then be used in future data analyses to represent that essence of the two items. In a sense we have reduced the two variables to one factor. Note that the new factor is actually a linear combination of the two variables.

Principal Components Analysis. The example described above, combining two correlated variables into one factor, illustrates the basic idea of factor analysis, or of principal components analysis to be precise (we will return to this later). If we extend the two-variable example to multiple variables, then the computations become more involved, but the basic principle of expressing two or more variables by a single factor remains the same.

Extracting Principal Components. The extraction of principal components amounts to a variance maximizing (varimax) rotation of the original variable space. For example, in a scatterplot we can think of the regression line as the original X axis, rotated so that it approximates the regression line. This type of rotation is called variance maximizing because the criterion for (goal of) the rotation is to maximize the variance (variability) of the "new" variable (factor), while minimizing the variance around the new variable

Generalizing to the Case of Multiple Variables. When there are more than two variables, we can think of them as defining a "space," just as two variables defined a plane. Thus, when we have three variables, we could plot a three- dimensional scatterplot, and, again we could fit a plane through the data.
With more than three variables it becomes impossible to illustrate the points in a scatterplot, however, the logic of rotating the axes so as to maximize the variance of the new factor remains the same.

Multiple orthogonal factors. After we have found the line on which the variance is maximal, there remains some variability around this line. In principal components analysis, after the first factor has been extracted, that is, after the first line has been drawn through the data, we continue and define another line that maximizes the remaining variability, and so on. In this manner, consecutive factors are extracted. Because each consecutive factor is defined to maximize the variability that is not captured by the preceding factor, consecutive factors are independent of each other. Put another way, consecutive factors are uncorrelated or orthogonal to each other.

How many Factors to Extract? Remember that, so far, we are considering principal components analysis as a data reduction method, that is, as a method for reducing the number of variables. The question then is, how many factors do we want to extract? Note that as we extract consecutive factors, they account for less and less variability. The decision of when to stop extracting factors basically depends on when there is only very little "random" variability left. The nature of this decision is arbitrary; however, various guidelines have been developed, and they are reviewed in Reviewing the Results of a Principal Components Analysis under Eigenvalues and the Number-of- Factors Problem.

Reviewing the Results of a Principal Components Analysis. Let us now look at some of the standard results from a principal components analysis. To reiterate, we are extracting factors that account for less and less variance. To simplify matters, you usually start with the correlation matrix, where the variances of all variables are equal to 1.0. Therefore, the total variance in that matrix is equal to the number of variables. For example, if we have 10 variables each with a variance of 1 then the total variability that can potentially be extracted is equal to 10 times 1. Suppose that in the satisfaction study introduced earlier we included 10 items to measure different aspects of satisfaction at home and at work. The variance accounted for by successive factors would be summarized as follows:



Eigenvalues
In the second column (Eigenvalue) above, we find the variance on the new factors that were successively extracted. In the third column, these values are expressed as a percent of the total variance (in this example, 10). As we can see, factor 1 accounts for 61 percent of the variance, factor 2 for 18 percent, and so on. As expected, the sum of the eigenvalues is equal to the number of variables. The third column contains the cumulative variance extracted. The variances extracted by the factors are called the eigenvalues. This name derives from the computational issues involved.

Eigenvalues and the Number-of-Factors Problem
Now that we have a measure of how much variance each successive factor extracts, we can return to the question of how many factors to retain. As mentioned earlier, by its nature this is an arbitrary decision. However, there are some guidelines that are commonly used, and that, in practice, seem to yield the best results.

The Kaiser criterion. First, we can retain only factors with eigen values greater than 1. In essence this is like saying that, unless a factor extracts at least as much as the equivalent of one original variable, we drop it. This criterion was proposed by Kaiser (1960), and is probably the one most widely used. In our example above, using this criterion, we would retain 2 factors (principal components).

The scree test. A graphical method is the scree test first proposed by Cattell (1966). We can plot the eigenvalues shown above in a simple line plot.


Cattell suggests to find the place where the smooth decrease of eigen values appears to level off to the right of the plot. To the right of this point, presumably, you find only "factorial scree" - "scree" is the geological term referring to the debris which collects on the lower part of a rocky slope. According to this criterion, we would probably retain 2 or 3 factors in our example.

Which criterion to use. Both criteria have been studied in detail (Browne, 1968; Cattell & Jaspers, 1967; Hakstian, Rogers, & Cattell, 1982; Linn, 1968; Tucker, Koopman & Linn, 1969). Theoretically, you can evaluate those criteria by generating random data based on a particular number of factors. You can then see whether the number of factors is accurately detected by those criteria. Using this general technique, the first method (Kaiser criterion) sometimes retains too many factors, while the second technique (scree test) sometimes retains too few; however, both do quite well under normal conditions, that is, when there are relatively few factors and many cases. In practice, an additional important aspect is the extent to which a solution is interpretable. Therefore, you usually examines several solutions with more or fewer factors, and chooses the one that makes the best "sense." We will discuss this issue in the context of factor rotations below.

Principal Factors Analysis
Before we continue to examine the different aspects of the typical output from a principal components analysis, let us now introduce principal factors analysis. Let us return to our satisfaction questionnaire example to conceive of another "mental model" for factor analysis. We can think of subjects' responses as being dependent on two components. First, there are some underlying common factors, such as the "satisfaction-with-hobbies" factor we looked at before. Each item measures some part of this common aspect of satisfaction. Second, each item also captures a unique aspect of satisfaction that is not addressed by any other item.

Communalities. If this model is correct, then we should not expect that the factors will extract all variance from our items; rather, only that proportion that is due to the common factors and shared by several items. In the language of factor analysis, the proportion of variance of a particular item that is due to common factors (shared with other items) is called communality. Therefore, an additional task facing us when applying this model is to estimate the communalities for each variable, that is, the proportion of variance that each item has in common with other items. The proportion of variance that is unique to each item is then the respective item's total variance minus the communality. A common starting point is to use the squared multiple correlation of an item with all other items as an estimate of the communality. Some authors have suggested various iterative "post-solution improvements" to the initial multiple regression communality estimate; for example, the so-called MINRES method (minimum residual factor method; Harman & Jones, 1966) will try various modifications to the factor loadings with the goal to minimize the residual (unexplained) sums of squares.

Principal factors vs. principal components. The defining characteristic then that distinguishes between the two factor analytic models is that in principal components analysis we assume that all variability in an item should be used in the analysis, while in principal factors analysis we only use the variability in an item that it has in common with the other items. In most cases, these two methods usually yield very similar results. However, principal components analysis is often preferred as a method for data reduction, while principal factors analysis is often preferred when the goal of the analysis is to detect structure

Factor Rotation: Its principles and types By: Arunava Guha- Roll: 12016


In the year 1984 Thompson demonstrated how the un-rotated pattern/structure actually misrepresents the true nature of the factors & how factor rotation resolves this misrepresentation. Interpretation of the factor analytical results is therefore almost always aided by factor rotation, as it’s possible to redistribute the common variance across the factors to achieve a more parsimonious solution. After the factor solution is rotated, the first un-rotated factor may not account for the largest portion of the variance & thus may not have the largest variance accounted for the value. Since the variance has been redistributed through-out the factors, any of the factors could account for the largest portion of the total variance.

Five principles of Factor Rotation explained by Gorsuch in the year 1983:
a>    
Each variable should have at-least one 0 loadings
b>    Each variable should have a set of linearity independent variables whose factor loadings are 0.
c>     For every pair of factors, there should be several variables whose loadings are 0 for one factor but not for other
d>    For every pair of factors, a large portion of variables should have loading 0 on both factors whenever more than about four factors are extracted
e>    For every pair of factors, there should only be a small no. of variables with non-zero loadings on both.
Thus, factor rotation is devised to shift the factors in their factor space so that each variable in the analysis has a large factor pattern coefficient on only one factor & has very small or 0 factor pattern coefficients on the other extracted latent constructs.

Types of Factor Rotation

Orthogonal Factor Rotation: It shifts the factor in the factor space maintaining 90degree angles of the factors to one another to achieve the best simple structure. Since the cosine of the angles between vectors of unit length equals r, & the cosine of a 90 degree angle is 0, this rotation strategy maintains the perfectly uncorrelated nature of the factors after the solution is rotated. Less sampling error occurs at this place due to less capitalization on chance that would occur if more parameters were estimated, as is the case in oblique rotation.
                Varimax Rotation: Of the Orthogonal Rotation Technique, it is one of the most popular rotation techniques. In this technique the factors are cleaned up so that every observed variable has a large factor pattern/structure coefficients for a small no. of variables & near-zero or very low pattern coefficients with the other group of variables.
                Quartimax Rotation: Another popular orthogonal Rotation Technique, in this technique the factor pattern of a variable is simplified by forcing the variable to correlate highly with one main factor the so called G-factor & very little with other factors. The variables are much easy to interpret in this case, but factors are more difficult to interpret since all variables are primarily associated with one factor.
Oblique Factor Rotation: The second type of factor rotation is Oblique Rotation. This method of rotation provides for correlations among the latent constructs. This is termed as oblique because the angles between the factors become more than 90 degree.
                Direct Oblimin: One of the popular Oblique Rotation techniques. This is moderated by a delta value, in which higher value of delta represents higher correlations between factors & negative value represents lesser correlations between the factors. This technique more closely honors the nature of reality & demands careful consideration by the researcher, as the correlation between factors must be set prior to analysis.
                Promax : Another most popular techniques of Oblique Rotation. In this technique researchers attempt to achieve the most parsimonious simple structure given that the factors are allowed to be correlated with one another.
It has three distinct steps :
a>    - Rotate the factor orthogonally
b>    -Target matrix is contrived by raising  the factor pattern coefficients to an exponent greater than 2(Typically exponent 3 or 4 are used). The coefficients in the target matrix become smaller, but the absolute distance between them actually increases.
c>     -The final step is Promax Rotation involves the “Procrustean” rotation of the original matrix to a best fit position with the target matrix. Promax is often the oblique rotation strategy of choice, as it’s relatively easy to use, typically provides good solutions & tends to produce more replicable results than the direct oblimin rotations

Reference:
“Orthogonal versus Oblique Rotation: A Review of the literature regarding the pros and cons”
by Kieffer, Kevin M published on 11/04/98

Exploratory Factor Analysis - Deepika Gnanasekar, 12075


1.1 Objectives
 The primary objectives of an EFA are to determine
1. The number of common factors infuencing a set of measures.
2. The strength of the relationship between each factor and each observed measure.

   Some common uses of EFA are to
1.     Identify the nature of the constructs underlying responses in a specific content area.
2.      Determine what sets of items hang together in a questionnaire.
3.     Demonstrate the dimensionality of a measurement scale. Researchers often wish to develop scales that respond to a single characteristic.
4.     Determine what features are most important when classifying a group of items.
      5.  Generate \factor scores" representing values of the underlying constructs for use in other    analyses.
   1.2 Performing EFA
There are seven basic steps to performing an EFA:
1.Collect measurements. You need to measure your variables on the same (or matched) experimental units.
2. Obtain the correlation matrix. You need to obtain the correlations (or covariances) between each of your variables.
3. Select the number of factors for inclusion. Sometimes you have a specifc hypothesis that will
determine the number factors you will include, while other times you simply want your final model
to account for as much of the covariance in your data with as few factors as possible. If you have k
measures, then you can at most extract k factors. There are a number of methods to determine the
optimal" number of factors by examining your data. The Kaiser criterion states that you should use
a number of factors equal to the number of the eigen values of the correlation matrix that are greater than one. The Scree test states that you should plot the eigenvalues of the correlation matrix in descending order, and then use a number of factors equal to the number of eigenvalues that occur prior to the last major drop in eigen value magnitude.
4. Extract your initial set of factors. You must submit your correlations or covariance into a computer  program to extract your factors. This step is too complex to reasonably be done by hand. Thereare a number of different extraction methods, including maximum likelihood, principal component,and principal axis extraction. The best method is generally maximum likelihood extraction, unless you seriously lack multivariate normality in your measures.
5. Rotate your factors to a final solution. For any given set of correlations and number of factors
there are actually an infinite number of ways that you can define your factors and still account for the same amount of covariance in your measures. Some of these definitions, however, are easier to interpret theoretically than others. By rotating your factors you attempt to and a factor solution that is equal to that obtained in the initial extraction but which has the simplest interpretation.
There are many diferent types of rotation, but they all try make your factors each highly responsive
to a small subset of your items (as opposed to being moderately responsive to a broad set). There
are two major categories of rotations, orthogonal rotations, which produce uncorrelated factors, and oblique rotations, which produce correlated factors. The best orthogonal rotation is widely believed to be Varimax. Oblique rotations are less distinguishable, with the three most commonly used being Direct Quartimin, Promax, and Harris-Kaiser Orthoblique.
6. Interpret your factor structure. Each of your measures will be linearly related to each of your
factors. The strength of this relationship is contained in the respective factor loading, produced by
your rotation. This loading can be interpreted as a standardized regression coefficient, regressing the factor on the measures. You define a factor by considering the possible theoretical constructs that could be responsible for the observed pattern of positive and negative loadings. To ease interpretation you have the option of multiplying all of the loadings for a given factor by -1. This essentially reverses the scale of the factor, allowing you, for example, to turn an unfriendliness factor into a friendliness factor.
7. Construct factor scores for further analysis. If you wish to perform additional analyses using
the factors as variables you will need to construct factor scores.  The score for a given factor is    a
linear combination of all of the measures, weighted by the corresponding factor loading. Sometimes factor scores are idealized, assigning a value of 1 to strongly positive loadings, a value of -1 to strongly negative loadings, and a value of 0 to intermediate loadings. These factor scores can then be used in analyses just like any other variable, although you should remember that they will be strongly collinear with the measures used to generate them.

Posted By:
Deepika Gnanasekar,
12075,
Marketing Batch
Source: “http://www.stat-help.com/factor.pdf”

How predictive analytics mints money for social networks - Gomathi Shankar K


Enough has already been said about the scary amount of information about us that others know through social networks. Admittedly, being the social animals that we are, we tend to take this scare with a pinch of salt and probably think hard to write a witty one-liner and post it as a status message aiming to get 20 'like's. Hence this article does not in any way force you to forbid social networking. All it tried to do is to make us aware of how the information we let loose can be used (read sold). I, being one of the so-called social animals myself, start this article with my witty one-liner. 

Social network indulgence is subject to privacy risks. Please read data mining carefully before updating status !

Social networks strive to influence and arrange contact and information sharing of many people. Meanwhile they also develop statistical tools to eavesdrop and sell information to the giant ears of the businessmen who eternally want that elusive root cause. This is normal statistical analysis. These sites also sell information to another set of giant business ears which want fodder to run their predictive machines. Fed by information input about us through social networks, marketers decide what to say to whom and whom to say what. 

Let us look at five ways in which social networks can mint money through predictive analytics. 

Matchmaking/Recruiting : Some HR professionals say that finding the right employee is like finding the right spouse. In that light, companies are serial polygamists and the society wants it that way for the greater good too. Social networks act as the perfect matchmakers to match right employees with right bosses. Information is available from both parties and all the sites have to do is run an intelligent match probability for the applicants and tell employers! LinkedIn does it and rightly so we flock there. As recruiting happens this way, the literal matchmaking happens in two ways through social networks, though only one will earn revenues for them. The matrimony sites do the same by predicting the matching accuracy for both parties. Social networks can feed them the vital information about the parties for the matrimony sites to enhance their 'happily married through us' rate. The other non-revenue generating mode will be well known to all social animals who got a date from a person who they just met online. In this case, the networks can only hope that they both have a great date and update the world about the experience soon after for them to sell their status messages. 

Sentiment Analysis: Social networks are just the right sources to analyse public sentiments now. Every major happening goes through a round of social network discussions before dying from public memory (which is happening at a very rapid case now). Who will be interested to know these sentiments? Politicians, New product launchers, Policy makers etc. Customers with big purses, really ! Social networks make merry. These customers can also identify popular influencers through these predictions and strategise their actions to gain their favour. Twitter is a good example. From the days of celebrities in Twitter, now we see Twitter celebrities who rose to fame only through their 140 caharacters. Social networks are also trying their hand at predicting election results. This link will explain it interestingly. 

Market fluctuations: Sentiments do not only play a major role in politics and product launches. They rule the stock market! Social networks simply adopt the same predictive power to mint money in the stock market. Tweets and updates by traders who are directly facing the full impact of sentiments offer the information for the networks to predict which share will be bear/bull the next day. 

Recommendation Engines: We all have studied in eight standard physics that an ideal engine is not practically possible. Well, recommendation engines aided by predictive powers of social networks are threatening to break that rule by their scarily close-to-thought suggestions. They are fast becoming mind readers aided by predictive statistics. 

Location Based Marketing: They know what we speak, what we think and what we feel. So naturally they know where we are. Social networks can also predict where we will be as individuals or groups. Marketers can use this predictions to reach before us there and wait for us to roll out their irresistible offers. Imagine you take great pains to gather information online and fix a perfect secluded honeymoon in the Bahamas, only to end up being greeted by a South-Indian restaurant waiting for you there at the airport with his "I knew you will come here now and I knew you will your favourite Masala dosa. Why don't you try our Dosa?" dialogue. Not totally impossible !


Thanks to Dr Rodo Kotorov's article for providing good references. 

_ Gomathi Shankar K
  12022




25.1.11

AN OVER VIEW AND APPLICATION BUSINESS INTELLIGENCE AND ANALYTICS IN INSURANCE MARKET- INDIA.

In today's ever evolving insurance business landscape, successful strategies of the past may not prove to be sustainable, let alone profitable in the future.

In today’s ever evolving insurance business landscape, successful strategies of the past may not prove to be sustainable, let alone profitable in the future. Over the last 50 years, the Indian insurance industry underwent three radical transformations.

  • The first transformation was when the private insurance industry with the usual level of government supervision, was nationalized. This was carried out in two stages - life insurance in 1956 and general insurance in 1972. This transformation created a monopoly in life insurance sector and a terrified oligopoly in the non-life sectors.
  • The second transformation in 2000 saw the re-entry of a new breed of private companies and also foreign players (with a restricted stake of up to 26 per cent).
  • The third transformation in 2007 saw the safety net of tariff based pricing being withdrawn in general insurance, leading to increased competition. With abiding speculations about the restrictions on foreign ownership being relaxed and the introduction of the new risk based capital norms along with various other changes, the industry is all set to metamorphose into a new landscape.

So, what does this mean in terms of business strategy for Indian insurance? What will be the impact of this anticipated future on today’s strategies? While in the long run, insurers may have to reinvent their present methods of doing business and adopt new and unforeseen initiatives, it would serve them well to optimize their current strategies to adapt to the business environment in the near future. In a data driven industry such as insurance, companies will not only need to compete in terms of their product offerings, but will also need to leverage business intelligence enabled analytics for a competitive edge.

In today’s increasingly competitive scenario, profitable growth is an elusive goal for the insurance industry. Rapid development and deployment of new products and product features, balancing broader distribution channel opportunities, managing risks across the organization, responding to increasingly demanding regulatory and reporting agency demands, and providing more precise pricing levels require effective decisions to be made with greater accuracy, efficiency and transparency. Insurers will need to analyze the impact of optimum non-financial performance in driving optimum financial performance. Instead of reacting to sudden changes, they would need to make accurate forecasts of future performance to plan business strategy.

Optimizing strategies: A driven by intelligent business analytics

In order to optimize their business processes, insurers will have to stabilize their existing operations and then accelerate profitability by understanding their optimum shape, product mix and operations. This translates into well defined strategies for the core areas of Marketing, Cost Management and Risk Management. Their strategies should empower them to: offer the most profitable product mix through the most cost-efficient channels to the most profitable customer; reduce operational losses through efficient claims fraud management and economic claims settlement; and effective risk based capital management. To dramatically improve efficiency and enhance business performance, strategic decisions should be aided by better executive and operational insight, by having the right infrastructure to support, analyze and understand the underlying complexities. The strategic insights gained into these core areas for insurance companies, both in the life as well as the non life sectors, by the use of intelligent business analytics are discussed in more detail in this article.

Marketing needs to focus on customer segmentation and effective distribution networks. Insurers experiencing poorer customer loyalty levels and increased costs, will need to become more customer-centric rather than product-focused. It will become extremely important for insurers to segment customers based on their behavior and potential profitability. Analytics can help them select more accurately, which policies and services to offer to which customers. To increase market share, insurers will have to employ cross-sell and up-sell techniques, facilitated by market basket analysis carried out on the basis of historical data of policies held by clients, as well as details of active policies, customer demographics, claims propensity and other key variables. Data from not only the traditional sources such as the insurers own records, but also sources such as records of the parent organizations can be used to identify various customer segment attributes. Apart from customer segmentation techniques, data mining can also be used to predict the likelihood of policy cancellation in advance, to aid in customer retention analysis by analyzing previous cancellation trends. Simulation models can also be used to generate the probable cash flows and compute the present value of a customer, by estimating the difference between the total amount of revenues from the customer and the expenses for the customer during the whole relationship period, known as the Customer Lifetime Value.

Increased competition is already putting downward pressure on premiums. This coupled with high customer churn and high customer acquisition costs has led insurance companies to implement multi-channel integration strategies. Distribution analytics can enable insurers to conduct “deep dives” into causal factors to answer a variety of questions. Issues such as unmet demand due to improper market assessment, segmentation, positioning and sales support, can be addressed. Similarly compensation and recognition programs influencing leading / lagging sales productivity and agent retention can be designed. All this can help in developing more effective sales channels.

New product development could also be positively influenced by better data analysis. Insurers’ portfolios of new products could be shaped by the new market opportunities which evolve from recent natural catastrophes, the latest technology, feedback from agents and customers, the availability of capital, the actions of competitors and changes in laws. With the use of data mining and predictive analytics, insurers can identify characteristics of individual risks and this will change how insurers see their market. New market opportunities can be identified by evaluating four main parameters. Firstly, by paying attention to the new and emerging needs of the customers. Secondly, by understanding the trends in the global marketplace, both demographically and geographically, and what the insurance implications of those trends might be. Thirdly, by developing new products in accordance with changes in legislation or regulatory environment. Fourthly, identifying new market opportunities by zeroing in on market needs that follow catastrophes.

Getting the right mix of price, claim, channels to market, operations and product differentiation is crucial for success and this can be achieved using business analytics.

Cost Management: Focus on claims fraud prediction and intervention.

As claims are prone to fraud or value inflation, the handling of claims affects the long-term sustainability of the company’s profits. Thus analytical solutions which can help manage the complex claims process effectively and help detect claims fraud by accurately forecasting likely outcomes in order to mitigate the severity of the claim, will be invaluable inputs in decision making and forecasting the loss reserves. Analytics can be used to benchmark claims to detect where padding might have occurred. Most companies have sufficient data to create claims benchmarks and claim value models. To manage fraud, companies can adopt a hybrid approach to detection, using a combination of profiling, rules to filter out fraudulent transactions and advanced analytics software.

Health insurers may develop predictive models to decrease claims costs by analyzing claim characteristics from their own claim files. This can
help them develop benchmark costs for various diseases by regions or by provider type or by severity. Such models can be used to identify highest-cost providers, claims having a higher propensity for fraud or enrollee with significant exposure. A step by step approach towards achieving this:

a) Acquire, load, and cleanse data from both internal and external sources, such as enrollment and TPA records.
b) Develop a methodology for carrying out effective analysis of variables, such as cluster diagnosis codes into easy to analyze disease categories.
c) Analyze and report on correlation of claim variables by geography, age, provider etc.
d) Utilize output to create package rates for treatments, set limits on benefits or undertake focus negotiations with providers.

As a consequence of the above actions, an insurer would be able to identify cost saving opportunities and develop tiered networks or disease specific limits to help achieve targeted loss ratios.

Risk Management: Focus on economic capital and solvency.

As every insurer has limited capital upon which to write new business, profitability is ultimately tied to effective risk management of capital. To adapt to a risk based approach, insurers will have to implement an economic capital regime, where they predict and evaluate the risk profiles of the underwritten business under both best case and worst case scenarios, and thus determine the prudent level of economic capital for sufficient reserves. Simulation techniques and stochastic models for various risks, such as credit, market, liability, group, underwriting and operational risk can be used for scenario and stress testing to determine the optimum economic capital level. This will not only facilitate risk-based capital but also enable real time solvency monitoring. Such a model would be adaptive to the insurer’s evolving environment, by adjusting its parameters as the economic conditions and liabilities change. The capability to accurately project its capital requirements can also be used to continuously monitor solvency since it can also estimate the change in the value of the time zero balance sheet as market conditions move.

An alternative to this approach is to calculate the solvency position on a statutory basis, but this would not measure the true likely response of the insurer to market events. Another possible approach is to use closed-form calculations for the cost of guarantees; however, these would be inaccurate because they do not allow for management actions and would thus result in only very approximate results. Such complex specifications cannot be met by traditional methods and require a state-of-the-art stochastic modeling approach, which is smart, fast and flexible.

Overcoming obstacles in implementation: Data quality and its capabilities

The issue of data has presented several challenges to all insurance industry stakeholders. Being able to conduct an in-depth analysis to provide critical strategic inputs directly depends on the quality of data. Over the past couple of years the awareness about utility of data in building rating structures, product design and flexible pricing systems has increased. However the industry has not fully transferred this knowledge to action. This was partly due to the fact that earlier the data was hard to use, it was frequently of poor quality and incomplete. In addition, different data was captured in different formats, thus making data aggregation and analysis at the insurance company’s end a very daunting task. This has changed over the past few years.

Suitability of data for analytical purposes can be ensured by focusing on two aspects - completeness and standardization. To ensure that the data is complete and accurate, validation features need to be introduced in the business front end and claims processing software tools. In addition to validation checks, it is imperative that the data must be collected in a standardized format. This can be achieved by standardizing the various forms used across the industry.

Insurance Regulatory and Development Authority (IRDA) has recognized that collection and dissemination of reliable and accurate data is important for the insurance industry and has formed the nodal Insurance Information Bureau (IIB). IIB has created a data repository to enable insurance companies, other stake holders and researchers to have easy access to validated data from one source.

However, data quality is only one side of the coin. The other side, supporting IT infrastructure, is again an area where insurers will have to work upon. An insurance IT system should capture and analyze information across all lines of business and risks. In addition, the system should provide analytic capabilities and produce reports for a variety of users. In general, a technology platform for deployment of a business intelligence analytical application should include the following:

  • Data integration technology that can capture data, clean, transform and load data from operational systems into a data warehouse in near real-time to enable offline analysis as well as on-line analytical processing (OLAP).
  • Data warehouse data management technology to manage information held in an operational data store, historical data marts for dimensional analysis, and analytical data stores for scoring and predictive data mining.
  • Dedicated servers to analyze data for multiple user reporting and analytical applications.
  • Tools to allow power users, managers and executives to build reports, slice, dice and drill, and mine data to model, score and predict.

Insurance is a data-rich industry and unfortunately, most of that data is underutilized. The key to gaining a competitive advantage in the insurance industry is found in analyzing this data and getting a greater business insight. Insurers can unlock the intelligence contained in their operational applications - like policy administration, claims management and CRM solutions - through modern data mining and analysis technology. Data mining uses predictive modeling, database segmentation, cluster analysis, neural networks and combinations thereof to quickly answer crucial business questions with greater accuracy. Business needs and business strategy must drive decisions about the structure and functionality of the business intelligence platform including the data warehouse and the data mart to enable this.

In summary, insurers should harness the power of extensive data provided by their business environment, conduct an extensive analysis of the data to model that environment and predict the consequences of alternative actions to guide executive decision making.

Business analytics solutions could be used by insurers to take rapid, effective and precise operational decisions in highly competitive markets, which maximize organizational value and minimize risks.

Submitted By,

Kichawele D. Msuya.

Roll Number: 12157.

MBA-Finance (2009-2011).