Search This Blog

26.1.11

Principal Components and Factor Analysis - Srikrishnan V, 12165, Marketing

General Purpose

The main applications of factor analytic techniques are: (1) to reduce the number of variables and (2) to detect structure in the relationships between variables, that is to classify variables. Therefore, factor analysis is applied as a data reduction or structure detection method .

Basic Idea of Factor Analysis as a Data Reduction Method

Suppose we conducted a (rather "silly") study in which we measure 100 people's height in inches and centimeters. Thus, we would have two variables that measure height. If in future studies, we want to research, for example, the effect of different nutritional food supplements on height, would we continue to use both measures? Probably not; height is one characteristic of a person, regardless of how it is measured.

Let's now extrapolate from this "silly" study to something that you might actually do as a researcher. Suppose we want to measure people's satisfaction with their lives. We design a satisfaction questionnaire with various items; among other things we ask our subjects how satisfied they are with their hobbies (item 1) and how intensely they are pursuing a hobby (item 2). Most likely, the responses to the two items are highly correlated with each other. Given a high correlation between the two items, we can conclude that they are quite redundant.

Combining Two Variables into a Single Factor. You can summarize the correlation between two variables in a scatterplot. Regression line can then be fitted that represents the "best" summary of the linear relationship between the variables. If we could define a variable that would approximate the regression line in such a plot, then that variable would capture most of the "essence" of the two items. Subjects' single scores on that new factor, represented by the regression line, could then be used in future data analyses to represent that essence of the two items. In a sense we have reduced the two variables to one factor. Note that the new factor is actually a linear combination of the two variables.

Principal Components Analysis. The example described above, combining two correlated variables into one factor, illustrates the basic idea of factor analysis, or of principal components analysis to be precise (we will return to this later). If we extend the two-variable example to multiple variables, then the computations become more involved, but the basic principle of expressing two or more variables by a single factor remains the same.

Extracting Principal Components. The extraction of principal components amounts to a variance maximizing (varimax) rotation of the original variable space. For example, in a scatterplot we can think of the regression line as the original X axis, rotated so that it approximates the regression line. This type of rotation is called variance maximizing because the criterion for (goal of) the rotation is to maximize the variance (variability) of the "new" variable (factor), while minimizing the variance around the new variable

Generalizing to the Case of Multiple Variables. When there are more than two variables, we can think of them as defining a "space," just as two variables defined a plane. Thus, when we have three variables, we could plot a three- dimensional scatterplot, and, again we could fit a plane through the data.
With more than three variables it becomes impossible to illustrate the points in a scatterplot, however, the logic of rotating the axes so as to maximize the variance of the new factor remains the same.

Multiple orthogonal factors. After we have found the line on which the variance is maximal, there remains some variability around this line. In principal components analysis, after the first factor has been extracted, that is, after the first line has been drawn through the data, we continue and define another line that maximizes the remaining variability, and so on. In this manner, consecutive factors are extracted. Because each consecutive factor is defined to maximize the variability that is not captured by the preceding factor, consecutive factors are independent of each other. Put another way, consecutive factors are uncorrelated or orthogonal to each other.

How many Factors to Extract? Remember that, so far, we are considering principal components analysis as a data reduction method, that is, as a method for reducing the number of variables. The question then is, how many factors do we want to extract? Note that as we extract consecutive factors, they account for less and less variability. The decision of when to stop extracting factors basically depends on when there is only very little "random" variability left. The nature of this decision is arbitrary; however, various guidelines have been developed, and they are reviewed in Reviewing the Results of a Principal Components Analysis under Eigenvalues and the Number-of- Factors Problem.

Reviewing the Results of a Principal Components Analysis. Let us now look at some of the standard results from a principal components analysis. To reiterate, we are extracting factors that account for less and less variance. To simplify matters, you usually start with the correlation matrix, where the variances of all variables are equal to 1.0. Therefore, the total variance in that matrix is equal to the number of variables. For example, if we have 10 variables each with a variance of 1 then the total variability that can potentially be extracted is equal to 10 times 1. Suppose that in the satisfaction study introduced earlier we included 10 items to measure different aspects of satisfaction at home and at work. The variance accounted for by successive factors would be summarized as follows:



Eigenvalues
In the second column (Eigenvalue) above, we find the variance on the new factors that were successively extracted. In the third column, these values are expressed as a percent of the total variance (in this example, 10). As we can see, factor 1 accounts for 61 percent of the variance, factor 2 for 18 percent, and so on. As expected, the sum of the eigenvalues is equal to the number of variables. The third column contains the cumulative variance extracted. The variances extracted by the factors are called the eigenvalues. This name derives from the computational issues involved.

Eigenvalues and the Number-of-Factors Problem
Now that we have a measure of how much variance each successive factor extracts, we can return to the question of how many factors to retain. As mentioned earlier, by its nature this is an arbitrary decision. However, there are some guidelines that are commonly used, and that, in practice, seem to yield the best results.

The Kaiser criterion. First, we can retain only factors with eigen values greater than 1. In essence this is like saying that, unless a factor extracts at least as much as the equivalent of one original variable, we drop it. This criterion was proposed by Kaiser (1960), and is probably the one most widely used. In our example above, using this criterion, we would retain 2 factors (principal components).

The scree test. A graphical method is the scree test first proposed by Cattell (1966). We can plot the eigenvalues shown above in a simple line plot.


Cattell suggests to find the place where the smooth decrease of eigen values appears to level off to the right of the plot. To the right of this point, presumably, you find only "factorial scree" - "scree" is the geological term referring to the debris which collects on the lower part of a rocky slope. According to this criterion, we would probably retain 2 or 3 factors in our example.

Which criterion to use. Both criteria have been studied in detail (Browne, 1968; Cattell & Jaspers, 1967; Hakstian, Rogers, & Cattell, 1982; Linn, 1968; Tucker, Koopman & Linn, 1969). Theoretically, you can evaluate those criteria by generating random data based on a particular number of factors. You can then see whether the number of factors is accurately detected by those criteria. Using this general technique, the first method (Kaiser criterion) sometimes retains too many factors, while the second technique (scree test) sometimes retains too few; however, both do quite well under normal conditions, that is, when there are relatively few factors and many cases. In practice, an additional important aspect is the extent to which a solution is interpretable. Therefore, you usually examines several solutions with more or fewer factors, and chooses the one that makes the best "sense." We will discuss this issue in the context of factor rotations below.

Principal Factors Analysis
Before we continue to examine the different aspects of the typical output from a principal components analysis, let us now introduce principal factors analysis. Let us return to our satisfaction questionnaire example to conceive of another "mental model" for factor analysis. We can think of subjects' responses as being dependent on two components. First, there are some underlying common factors, such as the "satisfaction-with-hobbies" factor we looked at before. Each item measures some part of this common aspect of satisfaction. Second, each item also captures a unique aspect of satisfaction that is not addressed by any other item.

Communalities. If this model is correct, then we should not expect that the factors will extract all variance from our items; rather, only that proportion that is due to the common factors and shared by several items. In the language of factor analysis, the proportion of variance of a particular item that is due to common factors (shared with other items) is called communality. Therefore, an additional task facing us when applying this model is to estimate the communalities for each variable, that is, the proportion of variance that each item has in common with other items. The proportion of variance that is unique to each item is then the respective item's total variance minus the communality. A common starting point is to use the squared multiple correlation of an item with all other items as an estimate of the communality. Some authors have suggested various iterative "post-solution improvements" to the initial multiple regression communality estimate; for example, the so-called MINRES method (minimum residual factor method; Harman & Jones, 1966) will try various modifications to the factor loadings with the goal to minimize the residual (unexplained) sums of squares.

Principal factors vs. principal components. The defining characteristic then that distinguishes between the two factor analytic models is that in principal components analysis we assume that all variability in an item should be used in the analysis, while in principal factors analysis we only use the variability in an item that it has in common with the other items. In most cases, these two methods usually yield very similar results. However, principal components analysis is often preferred as a method for data reduction, while principal factors analysis is often preferred when the goal of the analysis is to detect structure

Factor Rotation: Its principles and types By: Arunava Guha- Roll: 12016


In the year 1984 Thompson demonstrated how the un-rotated pattern/structure actually misrepresents the true nature of the factors & how factor rotation resolves this misrepresentation. Interpretation of the factor analytical results is therefore almost always aided by factor rotation, as it’s possible to redistribute the common variance across the factors to achieve a more parsimonious solution. After the factor solution is rotated, the first un-rotated factor may not account for the largest portion of the variance & thus may not have the largest variance accounted for the value. Since the variance has been redistributed through-out the factors, any of the factors could account for the largest portion of the total variance.

Five principles of Factor Rotation explained by Gorsuch in the year 1983:
a>    
Each variable should have at-least one 0 loadings
b>    Each variable should have a set of linearity independent variables whose factor loadings are 0.
c>     For every pair of factors, there should be several variables whose loadings are 0 for one factor but not for other
d>    For every pair of factors, a large portion of variables should have loading 0 on both factors whenever more than about four factors are extracted
e>    For every pair of factors, there should only be a small no. of variables with non-zero loadings on both.
Thus, factor rotation is devised to shift the factors in their factor space so that each variable in the analysis has a large factor pattern coefficient on only one factor & has very small or 0 factor pattern coefficients on the other extracted latent constructs.

Types of Factor Rotation

Orthogonal Factor Rotation: It shifts the factor in the factor space maintaining 90degree angles of the factors to one another to achieve the best simple structure. Since the cosine of the angles between vectors of unit length equals r, & the cosine of a 90 degree angle is 0, this rotation strategy maintains the perfectly uncorrelated nature of the factors after the solution is rotated. Less sampling error occurs at this place due to less capitalization on chance that would occur if more parameters were estimated, as is the case in oblique rotation.
                Varimax Rotation: Of the Orthogonal Rotation Technique, it is one of the most popular rotation techniques. In this technique the factors are cleaned up so that every observed variable has a large factor pattern/structure coefficients for a small no. of variables & near-zero or very low pattern coefficients with the other group of variables.
                Quartimax Rotation: Another popular orthogonal Rotation Technique, in this technique the factor pattern of a variable is simplified by forcing the variable to correlate highly with one main factor the so called G-factor & very little with other factors. The variables are much easy to interpret in this case, but factors are more difficult to interpret since all variables are primarily associated with one factor.
Oblique Factor Rotation: The second type of factor rotation is Oblique Rotation. This method of rotation provides for correlations among the latent constructs. This is termed as oblique because the angles between the factors become more than 90 degree.
                Direct Oblimin: One of the popular Oblique Rotation techniques. This is moderated by a delta value, in which higher value of delta represents higher correlations between factors & negative value represents lesser correlations between the factors. This technique more closely honors the nature of reality & demands careful consideration by the researcher, as the correlation between factors must be set prior to analysis.
                Promax : Another most popular techniques of Oblique Rotation. In this technique researchers attempt to achieve the most parsimonious simple structure given that the factors are allowed to be correlated with one another.
It has three distinct steps :
a>    - Rotate the factor orthogonally
b>    -Target matrix is contrived by raising  the factor pattern coefficients to an exponent greater than 2(Typically exponent 3 or 4 are used). The coefficients in the target matrix become smaller, but the absolute distance between them actually increases.
c>     -The final step is Promax Rotation involves the “Procrustean” rotation of the original matrix to a best fit position with the target matrix. Promax is often the oblique rotation strategy of choice, as it’s relatively easy to use, typically provides good solutions & tends to produce more replicable results than the direct oblimin rotations

Reference:
“Orthogonal versus Oblique Rotation: A Review of the literature regarding the pros and cons”
by Kieffer, Kevin M published on 11/04/98

Exploratory Factor Analysis - Deepika Gnanasekar, 12075


1.1 Objectives
 The primary objectives of an EFA are to determine
1. The number of common factors infuencing a set of measures.
2. The strength of the relationship between each factor and each observed measure.

   Some common uses of EFA are to
1.     Identify the nature of the constructs underlying responses in a specific content area.
2.      Determine what sets of items hang together in a questionnaire.
3.     Demonstrate the dimensionality of a measurement scale. Researchers often wish to develop scales that respond to a single characteristic.
4.     Determine what features are most important when classifying a group of items.
      5.  Generate \factor scores" representing values of the underlying constructs for use in other    analyses.
   1.2 Performing EFA
There are seven basic steps to performing an EFA:
1.Collect measurements. You need to measure your variables on the same (or matched) experimental units.
2. Obtain the correlation matrix. You need to obtain the correlations (or covariances) between each of your variables.
3. Select the number of factors for inclusion. Sometimes you have a specifc hypothesis that will
determine the number factors you will include, while other times you simply want your final model
to account for as much of the covariance in your data with as few factors as possible. If you have k
measures, then you can at most extract k factors. There are a number of methods to determine the
optimal" number of factors by examining your data. The Kaiser criterion states that you should use
a number of factors equal to the number of the eigen values of the correlation matrix that are greater than one. The Scree test states that you should plot the eigenvalues of the correlation matrix in descending order, and then use a number of factors equal to the number of eigenvalues that occur prior to the last major drop in eigen value magnitude.
4. Extract your initial set of factors. You must submit your correlations or covariance into a computer  program to extract your factors. This step is too complex to reasonably be done by hand. Thereare a number of different extraction methods, including maximum likelihood, principal component,and principal axis extraction. The best method is generally maximum likelihood extraction, unless you seriously lack multivariate normality in your measures.
5. Rotate your factors to a final solution. For any given set of correlations and number of factors
there are actually an infinite number of ways that you can define your factors and still account for the same amount of covariance in your measures. Some of these definitions, however, are easier to interpret theoretically than others. By rotating your factors you attempt to and a factor solution that is equal to that obtained in the initial extraction but which has the simplest interpretation.
There are many diferent types of rotation, but they all try make your factors each highly responsive
to a small subset of your items (as opposed to being moderately responsive to a broad set). There
are two major categories of rotations, orthogonal rotations, which produce uncorrelated factors, and oblique rotations, which produce correlated factors. The best orthogonal rotation is widely believed to be Varimax. Oblique rotations are less distinguishable, with the three most commonly used being Direct Quartimin, Promax, and Harris-Kaiser Orthoblique.
6. Interpret your factor structure. Each of your measures will be linearly related to each of your
factors. The strength of this relationship is contained in the respective factor loading, produced by
your rotation. This loading can be interpreted as a standardized regression coefficient, regressing the factor on the measures. You define a factor by considering the possible theoretical constructs that could be responsible for the observed pattern of positive and negative loadings. To ease interpretation you have the option of multiplying all of the loadings for a given factor by -1. This essentially reverses the scale of the factor, allowing you, for example, to turn an unfriendliness factor into a friendliness factor.
7. Construct factor scores for further analysis. If you wish to perform additional analyses using
the factors as variables you will need to construct factor scores.  The score for a given factor is    a
linear combination of all of the measures, weighted by the corresponding factor loading. Sometimes factor scores are idealized, assigning a value of 1 to strongly positive loadings, a value of -1 to strongly negative loadings, and a value of 0 to intermediate loadings. These factor scores can then be used in analyses just like any other variable, although you should remember that they will be strongly collinear with the measures used to generate them.

Posted By:
Deepika Gnanasekar,
12075,
Marketing Batch
Source: “http://www.stat-help.com/factor.pdf”

How predictive analytics mints money for social networks - Gomathi Shankar K


Enough has already been said about the scary amount of information about us that others know through social networks. Admittedly, being the social animals that we are, we tend to take this scare with a pinch of salt and probably think hard to write a witty one-liner and post it as a status message aiming to get 20 'like's. Hence this article does not in any way force you to forbid social networking. All it tried to do is to make us aware of how the information we let loose can be used (read sold). I, being one of the so-called social animals myself, start this article with my witty one-liner. 

Social network indulgence is subject to privacy risks. Please read data mining carefully before updating status !

Social networks strive to influence and arrange contact and information sharing of many people. Meanwhile they also develop statistical tools to eavesdrop and sell information to the giant ears of the businessmen who eternally want that elusive root cause. This is normal statistical analysis. These sites also sell information to another set of giant business ears which want fodder to run their predictive machines. Fed by information input about us through social networks, marketers decide what to say to whom and whom to say what. 

Let us look at five ways in which social networks can mint money through predictive analytics. 

Matchmaking/Recruiting : Some HR professionals say that finding the right employee is like finding the right spouse. In that light, companies are serial polygamists and the society wants it that way for the greater good too. Social networks act as the perfect matchmakers to match right employees with right bosses. Information is available from both parties and all the sites have to do is run an intelligent match probability for the applicants and tell employers! LinkedIn does it and rightly so we flock there. As recruiting happens this way, the literal matchmaking happens in two ways through social networks, though only one will earn revenues for them. The matrimony sites do the same by predicting the matching accuracy for both parties. Social networks can feed them the vital information about the parties for the matrimony sites to enhance their 'happily married through us' rate. The other non-revenue generating mode will be well known to all social animals who got a date from a person who they just met online. In this case, the networks can only hope that they both have a great date and update the world about the experience soon after for them to sell their status messages. 

Sentiment Analysis: Social networks are just the right sources to analyse public sentiments now. Every major happening goes through a round of social network discussions before dying from public memory (which is happening at a very rapid case now). Who will be interested to know these sentiments? Politicians, New product launchers, Policy makers etc. Customers with big purses, really ! Social networks make merry. These customers can also identify popular influencers through these predictions and strategise their actions to gain their favour. Twitter is a good example. From the days of celebrities in Twitter, now we see Twitter celebrities who rose to fame only through their 140 caharacters. Social networks are also trying their hand at predicting election results. This link will explain it interestingly. 

Market fluctuations: Sentiments do not only play a major role in politics and product launches. They rule the stock market! Social networks simply adopt the same predictive power to mint money in the stock market. Tweets and updates by traders who are directly facing the full impact of sentiments offer the information for the networks to predict which share will be bear/bull the next day. 

Recommendation Engines: We all have studied in eight standard physics that an ideal engine is not practically possible. Well, recommendation engines aided by predictive powers of social networks are threatening to break that rule by their scarily close-to-thought suggestions. They are fast becoming mind readers aided by predictive statistics. 

Location Based Marketing: They know what we speak, what we think and what we feel. So naturally they know where we are. Social networks can also predict where we will be as individuals or groups. Marketers can use this predictions to reach before us there and wait for us to roll out their irresistible offers. Imagine you take great pains to gather information online and fix a perfect secluded honeymoon in the Bahamas, only to end up being greeted by a South-Indian restaurant waiting for you there at the airport with his "I knew you will come here now and I knew you will your favourite Masala dosa. Why don't you try our Dosa?" dialogue. Not totally impossible !


Thanks to Dr Rodo Kotorov's article for providing good references. 

_ Gomathi Shankar K
  12022




25.1.11

AN OVER VIEW AND APPLICATION BUSINESS INTELLIGENCE AND ANALYTICS IN INSURANCE MARKET- INDIA.

In today's ever evolving insurance business landscape, successful strategies of the past may not prove to be sustainable, let alone profitable in the future.

In today’s ever evolving insurance business landscape, successful strategies of the past may not prove to be sustainable, let alone profitable in the future. Over the last 50 years, the Indian insurance industry underwent three radical transformations.

  • The first transformation was when the private insurance industry with the usual level of government supervision, was nationalized. This was carried out in two stages - life insurance in 1956 and general insurance in 1972. This transformation created a monopoly in life insurance sector and a terrified oligopoly in the non-life sectors.
  • The second transformation in 2000 saw the re-entry of a new breed of private companies and also foreign players (with a restricted stake of up to 26 per cent).
  • The third transformation in 2007 saw the safety net of tariff based pricing being withdrawn in general insurance, leading to increased competition. With abiding speculations about the restrictions on foreign ownership being relaxed and the introduction of the new risk based capital norms along with various other changes, the industry is all set to metamorphose into a new landscape.

So, what does this mean in terms of business strategy for Indian insurance? What will be the impact of this anticipated future on today’s strategies? While in the long run, insurers may have to reinvent their present methods of doing business and adopt new and unforeseen initiatives, it would serve them well to optimize their current strategies to adapt to the business environment in the near future. In a data driven industry such as insurance, companies will not only need to compete in terms of their product offerings, but will also need to leverage business intelligence enabled analytics for a competitive edge.

In today’s increasingly competitive scenario, profitable growth is an elusive goal for the insurance industry. Rapid development and deployment of new products and product features, balancing broader distribution channel opportunities, managing risks across the organization, responding to increasingly demanding regulatory and reporting agency demands, and providing more precise pricing levels require effective decisions to be made with greater accuracy, efficiency and transparency. Insurers will need to analyze the impact of optimum non-financial performance in driving optimum financial performance. Instead of reacting to sudden changes, they would need to make accurate forecasts of future performance to plan business strategy.

Optimizing strategies: A driven by intelligent business analytics

In order to optimize their business processes, insurers will have to stabilize their existing operations and then accelerate profitability by understanding their optimum shape, product mix and operations. This translates into well defined strategies for the core areas of Marketing, Cost Management and Risk Management. Their strategies should empower them to: offer the most profitable product mix through the most cost-efficient channels to the most profitable customer; reduce operational losses through efficient claims fraud management and economic claims settlement; and effective risk based capital management. To dramatically improve efficiency and enhance business performance, strategic decisions should be aided by better executive and operational insight, by having the right infrastructure to support, analyze and understand the underlying complexities. The strategic insights gained into these core areas for insurance companies, both in the life as well as the non life sectors, by the use of intelligent business analytics are discussed in more detail in this article.

Marketing needs to focus on customer segmentation and effective distribution networks. Insurers experiencing poorer customer loyalty levels and increased costs, will need to become more customer-centric rather than product-focused. It will become extremely important for insurers to segment customers based on their behavior and potential profitability. Analytics can help them select more accurately, which policies and services to offer to which customers. To increase market share, insurers will have to employ cross-sell and up-sell techniques, facilitated by market basket analysis carried out on the basis of historical data of policies held by clients, as well as details of active policies, customer demographics, claims propensity and other key variables. Data from not only the traditional sources such as the insurers own records, but also sources such as records of the parent organizations can be used to identify various customer segment attributes. Apart from customer segmentation techniques, data mining can also be used to predict the likelihood of policy cancellation in advance, to aid in customer retention analysis by analyzing previous cancellation trends. Simulation models can also be used to generate the probable cash flows and compute the present value of a customer, by estimating the difference between the total amount of revenues from the customer and the expenses for the customer during the whole relationship period, known as the Customer Lifetime Value.

Increased competition is already putting downward pressure on premiums. This coupled with high customer churn and high customer acquisition costs has led insurance companies to implement multi-channel integration strategies. Distribution analytics can enable insurers to conduct “deep dives” into causal factors to answer a variety of questions. Issues such as unmet demand due to improper market assessment, segmentation, positioning and sales support, can be addressed. Similarly compensation and recognition programs influencing leading / lagging sales productivity and agent retention can be designed. All this can help in developing more effective sales channels.

New product development could also be positively influenced by better data analysis. Insurers’ portfolios of new products could be shaped by the new market opportunities which evolve from recent natural catastrophes, the latest technology, feedback from agents and customers, the availability of capital, the actions of competitors and changes in laws. With the use of data mining and predictive analytics, insurers can identify characteristics of individual risks and this will change how insurers see their market. New market opportunities can be identified by evaluating four main parameters. Firstly, by paying attention to the new and emerging needs of the customers. Secondly, by understanding the trends in the global marketplace, both demographically and geographically, and what the insurance implications of those trends might be. Thirdly, by developing new products in accordance with changes in legislation or regulatory environment. Fourthly, identifying new market opportunities by zeroing in on market needs that follow catastrophes.

Getting the right mix of price, claim, channels to market, operations and product differentiation is crucial for success and this can be achieved using business analytics.

Cost Management: Focus on claims fraud prediction and intervention.

As claims are prone to fraud or value inflation, the handling of claims affects the long-term sustainability of the company’s profits. Thus analytical solutions which can help manage the complex claims process effectively and help detect claims fraud by accurately forecasting likely outcomes in order to mitigate the severity of the claim, will be invaluable inputs in decision making and forecasting the loss reserves. Analytics can be used to benchmark claims to detect where padding might have occurred. Most companies have sufficient data to create claims benchmarks and claim value models. To manage fraud, companies can adopt a hybrid approach to detection, using a combination of profiling, rules to filter out fraudulent transactions and advanced analytics software.

Health insurers may develop predictive models to decrease claims costs by analyzing claim characteristics from their own claim files. This can
help them develop benchmark costs for various diseases by regions or by provider type or by severity. Such models can be used to identify highest-cost providers, claims having a higher propensity for fraud or enrollee with significant exposure. A step by step approach towards achieving this:

a) Acquire, load, and cleanse data from both internal and external sources, such as enrollment and TPA records.
b) Develop a methodology for carrying out effective analysis of variables, such as cluster diagnosis codes into easy to analyze disease categories.
c) Analyze and report on correlation of claim variables by geography, age, provider etc.
d) Utilize output to create package rates for treatments, set limits on benefits or undertake focus negotiations with providers.

As a consequence of the above actions, an insurer would be able to identify cost saving opportunities and develop tiered networks or disease specific limits to help achieve targeted loss ratios.

Risk Management: Focus on economic capital and solvency.

As every insurer has limited capital upon which to write new business, profitability is ultimately tied to effective risk management of capital. To adapt to a risk based approach, insurers will have to implement an economic capital regime, where they predict and evaluate the risk profiles of the underwritten business under both best case and worst case scenarios, and thus determine the prudent level of economic capital for sufficient reserves. Simulation techniques and stochastic models for various risks, such as credit, market, liability, group, underwriting and operational risk can be used for scenario and stress testing to determine the optimum economic capital level. This will not only facilitate risk-based capital but also enable real time solvency monitoring. Such a model would be adaptive to the insurer’s evolving environment, by adjusting its parameters as the economic conditions and liabilities change. The capability to accurately project its capital requirements can also be used to continuously monitor solvency since it can also estimate the change in the value of the time zero balance sheet as market conditions move.

An alternative to this approach is to calculate the solvency position on a statutory basis, but this would not measure the true likely response of the insurer to market events. Another possible approach is to use closed-form calculations for the cost of guarantees; however, these would be inaccurate because they do not allow for management actions and would thus result in only very approximate results. Such complex specifications cannot be met by traditional methods and require a state-of-the-art stochastic modeling approach, which is smart, fast and flexible.

Overcoming obstacles in implementation: Data quality and its capabilities

The issue of data has presented several challenges to all insurance industry stakeholders. Being able to conduct an in-depth analysis to provide critical strategic inputs directly depends on the quality of data. Over the past couple of years the awareness about utility of data in building rating structures, product design and flexible pricing systems has increased. However the industry has not fully transferred this knowledge to action. This was partly due to the fact that earlier the data was hard to use, it was frequently of poor quality and incomplete. In addition, different data was captured in different formats, thus making data aggregation and analysis at the insurance company’s end a very daunting task. This has changed over the past few years.

Suitability of data for analytical purposes can be ensured by focusing on two aspects - completeness and standardization. To ensure that the data is complete and accurate, validation features need to be introduced in the business front end and claims processing software tools. In addition to validation checks, it is imperative that the data must be collected in a standardized format. This can be achieved by standardizing the various forms used across the industry.

Insurance Regulatory and Development Authority (IRDA) has recognized that collection and dissemination of reliable and accurate data is important for the insurance industry and has formed the nodal Insurance Information Bureau (IIB). IIB has created a data repository to enable insurance companies, other stake holders and researchers to have easy access to validated data from one source.

However, data quality is only one side of the coin. The other side, supporting IT infrastructure, is again an area where insurers will have to work upon. An insurance IT system should capture and analyze information across all lines of business and risks. In addition, the system should provide analytic capabilities and produce reports for a variety of users. In general, a technology platform for deployment of a business intelligence analytical application should include the following:

  • Data integration technology that can capture data, clean, transform and load data from operational systems into a data warehouse in near real-time to enable offline analysis as well as on-line analytical processing (OLAP).
  • Data warehouse data management technology to manage information held in an operational data store, historical data marts for dimensional analysis, and analytical data stores for scoring and predictive data mining.
  • Dedicated servers to analyze data for multiple user reporting and analytical applications.
  • Tools to allow power users, managers and executives to build reports, slice, dice and drill, and mine data to model, score and predict.

Insurance is a data-rich industry and unfortunately, most of that data is underutilized. The key to gaining a competitive advantage in the insurance industry is found in analyzing this data and getting a greater business insight. Insurers can unlock the intelligence contained in their operational applications - like policy administration, claims management and CRM solutions - through modern data mining and analysis technology. Data mining uses predictive modeling, database segmentation, cluster analysis, neural networks and combinations thereof to quickly answer crucial business questions with greater accuracy. Business needs and business strategy must drive decisions about the structure and functionality of the business intelligence platform including the data warehouse and the data mart to enable this.

In summary, insurers should harness the power of extensive data provided by their business environment, conduct an extensive analysis of the data to model that environment and predict the consequences of alternative actions to guide executive decision making.

Business analytics solutions could be used by insurers to take rapid, effective and precise operational decisions in highly competitive markets, which maximize organizational value and minimize risks.

Submitted By,

Kichawele D. Msuya.

Roll Number: 12157.

MBA-Finance (2009-2011).

How to Lie with Statistics!!! harvinder singh 12083

How to Lie with Statistics!!!
“How to Lie with Statistics is a book written by Darrell Huff in 1954 presenting an introduction to statistics for the general reader. Huff was a journalist who wrote many "how to" articles as a freelancer, but was not a statistician.
The book is a brief, breezy, illustrated volume outlining common errors, both intentional and unintentional, associated with the interpretation of statistics, and how these errors can lead to inaccurate conclusions. It has become one of the best-selling statistics books in history, with over one and a half million copies sold in the English-language edition, even though the monetary examples have become dated because of inflation. It has also been widely translated.
Themes of the book include "Correlation does not imply causation" and "Using Random Sampling". It also shows how statistical graphs can be used to distort reality, for example by truncating the bottom of a line or bar chart, so that differences seem larger than they are, or by representing one-dimensional quantities on a pictogram by two- or three-dimensional objects to compare their sizes, so that the reader forgets that the images don't scale the same way the quantities do.
The original edition contained humorous illustrations by artist Irving Geis. In a UK edition these were replaced with cartoons by Mel Calman.
Source: “http://en.wikipedia.org/wiki/How_to_Lie_with_Statistics

Correlation does not imply causation

"Correlation does not imply causation" (also called "Ignoring a Common Cause" and "Questionable Cause") is a phrase used in science and statistics to emphasize that correlation between two variables does not automatically imply that one causes the other (though correlation is necessary for linear causation, and can indicate possible causes or areas for further investigation... in other words, correlation can be a hint).
The opposite belief, correlation proves causation, is a logical fallacy by which two events that occur together are claimed to have a cause-and-effect relationship. The fallacy is also known as cum hoc ergo propter hoc (Latin for "with this, therefore because of this") and false cause. By contrast, the fallacy post hoc ergo propter hoc requires that one event occur before the other and so may be considered a type of cum hoc fallacy.
In a widely-studied example, numerous epidemiological studies showed that women who were taking combined hormone replacement therapy (HRT) also had a lower-than-average incidence of coronary heart disease (CHD), leading doctors to propose that HRT was protective against CHD. But randomized controlled trials showed that HRT caused a small but statistically significant increase in risk of CHD. Re-analysis of the data from the epidemiological studies showed that women undertaking HRT were more likely to be from higher socio-economic
groups (ABC1), with better than average diet and exercise regimes. The use of HRT and decreased incidence of coronary heart disease were coincident effects of a common cause (i.e., the benefits associated with a higher socioeconomic status), rather than cause and effect as had been supposed.

Usage

In logic, the technical use of the word "implies" means "to be a sufficient circumstance". This is the meaning intended by statisticians when they say causation is not certain. Indeed, p implies q has the technical meaning of logical implication: if p then q symbolized as p → q. That is "if circumstance p is true, then q necessarily follows." In this sense, it is always correct to say "Correlation does not imply causation".
However, in casual use, the word "imply" loosely means suggests rather than requires. The idea that correlation and causation are connected is certainly true; where there is causation, there is likely to be correlation. Indeed, correlation is used when inferring causation; the important point is that such inferences are not always correct because there are other possibilities, as explained later in this article.
Edward Tufte, in a criticism of the brevity of Microsoft PowerPoint presentations, deprecates the use of "is" to relate correlation and causation (as in "Correlation is not causation"), citing its inaccuracy as incomplete.[1] While it is not the case that correlation is causation, simply stating their nonequivalence omits information about their relationship. Tufte suggests that the shortest true statement that can be made about causality and correlation is one of the following:[4]
  • "Empirically observed covariation is a necessary but not sufficient condition for causality."
  • "Correlation is not causation but it sure is a hint."

General pattern

The cum hoc ergo propter hoc logical fallacy can be expressed as follows:
  1. A occurs in correlation with B.
  2. Therefore, A causes B.
In this type of logical fallacy, one makes a premature conclusion about causality after observing only a correlation between two or more factors. Generally, if one factor (A) is observed to only be correlated with another factor (B), it is sometimes taken for granted that A is causing B even when no evidence supports it. This is a logical fallacy because there are at least five possibilities:
  1. A may be the cause of B.
  2. B may be the cause of A.
  3. some unknown third factor C may actually be the cause of both A and B.
  4. there may be a combination of the above three relationships. For example, B may be the cause of A at the same time as A is the cause of B (contradicting that the only relationship between A and B is that A causes B). This describes a self-reinforcing system.
  5. the "relationship" is a coincidence or so complex or indirect that it is more effectively called a coincidence (i.e. two events occurring at the same time that have no direct relationship to each other besides the fact that they are occurring at the same time). A larger sample size helps to reduce the chance of a coincidence, unless there is a systematic error in the experiment.
In other words, there can be no conclusion made regarding the existence or the direction of a cause and effect relationship only from the fact that A and B are correlated. Determining whether there is an actual cause and effect relationship requires further investigation, even when the relationship between A and B is statistically significant, a large effect size is observed, or a large part of the variance is explained.

Examples

B causes A (reverse causation)

The more firemen fighting a fire, the bigger the fire is going to be.
Therefore firemen cause fire.
The above example is simple and easy to understand. The strong correlation between the number of firemen at a scene and the size of the fire that is present does not imply that the firemen cause the fire. Firemen are sent according to the severity of the fire and if there is a large fire, a greater number of firemen are sent; therefore it is rather that fire causes firemen to arrive at the scene.

A causes B and B causes A (bidirectional causation)

Increased pressure results in increased temperature.
Therefore pressure causes temperature.
The ideal gas law, PV = nRT, describes the direct relationship between pressure and temperature (along with other factors) to show that there is a direct correlation between the two properties. For a fixed volume, an increase in temperature will cause an increase in pressure; likewise, increased pressure will cause an increase in temperature. This demonstrates (4) in that the two are directly proportional to each other and not independent functions.

Third factor C (the common-causal variable) causes both A and B

Main article: Spurious relationship
All these examples deal with a lurking variable, which is simply a hidden third variable that affects both clauses of the correlation; for example, the fact that it is summer in Example 3. A difficulty often also arises where the third factor, though fundamentally different from A and B, is so closely related to A and/or B as to be confused with them or very difficult to scientifically disentangle from them (see Example 4).
Example 1
Sleeping with one's shoes on is strongly correlated with waking up with a headache.
Therefore, sleeping with one's shoes on causes headache.
The above example commits the correlation-implies-causation fallacy, as it prematurely concludes that sleeping with one's shoes on causes headache. A more plausible explanation is that both are caused by a third factor, in this case going to bed drunk, which thereby gives rise to a correlation.
Example 2
Young children who sleep with the light on are much more likely to develop myopia in later life.
The former is a recent scientific example that resulted from a study at the University of Pennsylvania Medical Center. Published in the May 13, 1999 issue of Nature,[5] the study received much coverage at the time in the popular press.[6] However, a later study at The Ohio State University did not find that infants sleeping with the light on caused the development of myopia. It did find a strong link between parental myopia and the development of child myopia, also noting that myopic parents were more likely to leave a light on in their children's bedroom.[7][8][9][10] In this case, the cause of both conditions is parental myopia.
Example 3
As ice cream sales increase, the rate of drowning deaths increases sharply.
Therefore, ice cream causes drowning.
The aforementioned example fails to recognize the importance of time in relationship to ice cream sales. Ice cream is sold during the summer months at a much greater rate, and it is during the summer months that people are more likely to engage in activities involving water, such as swimming. The increased drowning deaths are simply caused by more exposure to water based activities, not ice cream.
Example 4
A hypothetical study shows a relationship between test anxiety scores and shyness scores, with a statistical r value (strength of correlation) of +.59.[11] Therefore, it may be simply concluded that shyness, in some part, causally influences test anxiety. However, as encountered in many psychological studies, another variable, a "self-consciousness score," is discovered which has a sharper correlation (+.73) with shyness. This suggests a possible "third variable" problem, however, when three such closely related measures are found, it further suggests that each may have bidirectional tendencies (see "bidirectional variable," above), being a cluster of correlated values each influencing one another to some extent.

Coincidence

With a decrease in the number of pirates, there has been an increase in global warming over the same period.
Therefore, global warming is caused by a lack of pirates.
The example above is used satirically by the parody religion Pastafarianism to illustrate the logical fallacy of assuming that correlation equals causation.
Since the 1950s, both the atmospheric CO2 level and crime levels have increased sharply.
Hence, atmospheric CO2 causes crime.
The above example arguably makes the mistake of prematurely concluding a causal relationship where the relationship between the variables, if any, is so complex it may be labeled coincidental. The two events have no simple relationship to each other beside the fact that they are occurring at the same time. Another possible example is the somewhat jocular Mierscheid Law.

Determining causation

David Hume argued that causality is based on experience, and experience similarly based on the assumption that the future models the past, which in turn can only be based on experience – leading to circular logic. In conclusion he asserted that causality is not based on actual reasoning: only correlation can actually be perceived.[12]
Intuitively, causation seems to require not just a correlation, but a counterfactual dependence. Suppose that a student performed poorly on a test and guesses that the cause was his not studying. To prove this, one thinks of the counterfactual – the same student writing the same test under the same circumstances but having studied the night before. If one could rewind history, and change only one small thing (making the student study for the exam), then causation could be observed (by comparing version 1 to version 2). Because one cannot rewind history and replay events after making small controlled changes, causation can only be inferred, never exactly known. This is referred to as the Fundamental Problem of Causal Inference – it is impossible to directly observe causal effects.[13]
A major goal of scientific experiments and statistical methods is to approximate as best as possible the counterfactual state of the world.[14] For example, one could run an experiment on identical twins who were known to consistently get the same grades on their tests. One twin is sent to study for six hours while the other is sent to the amusement park. If their test scores suddenly diverged by a large degree, this would be strong evidence that studying (or going to the amusement park) had a causal effect on test scores. In this case, correlation between studying and test scores would almost certainly imply causation.
Well-designed experimental studies replace equality of individuals as in the previous example by equality of groups. This is achieved by randomization of the subjects to two or more groups. Although not a perfect system, the likeliness of being equal in all aspects rises with the number of subjects placed randomly in the treatment/placebo groups. From the significance of the difference of the effect of the treatment vs. the placebo, one can conclude the likeliness of the treatment having a causal effect on the disease. This likeliness can be quantified in statistical terms by the P-value.
When experimental studies are impossible and only pre-existing data are available, as is usually the case for example in economics, regression analysis can be used. Factors other than the potential causative variable of interest are controlled for by including them as regressors in addition to the regressor representing the variable of interest. False inferences of causation due to reverse causation (or wrong estimates of the magnitude of causation due the presence of bidirectional causation) can be avoided by using explanators (regressors) that are necessarily exogenous, such as physical explanators like rainfall amount (as a determinant of, say, futures prices), lagged variables whose values were determined before the dependent variable's value was determined, instrumental variables for the explanators (chosen based on their known exogeneity), etc. See Causality#Economics. Spurious correlation due to mutual influence from a third, common causative, variable, is harder to avoid: the model must be specified such that there is a theoretical reason to believe that no such underlying causative variable has been omitted from the model; in particular, underlying time trends of both the dependent variable and the independent (potentially causative) variable must be controlled for by including time as another independent variable.

http://en.wikipedia.org/wiki/Correlation_does_not_imply_causation