Search This Blog

25.1.11

AN OVER VIEW AND APPLICATION BUSINESS INTELLIGENCE AND ANALYTICS IN INSURANCE MARKET- INDIA.

In today's ever evolving insurance business landscape, successful strategies of the past may not prove to be sustainable, let alone profitable in the future.

In today’s ever evolving insurance business landscape, successful strategies of the past may not prove to be sustainable, let alone profitable in the future. Over the last 50 years, the Indian insurance industry underwent three radical transformations.

  • The first transformation was when the private insurance industry with the usual level of government supervision, was nationalized. This was carried out in two stages - life insurance in 1956 and general insurance in 1972. This transformation created a monopoly in life insurance sector and a terrified oligopoly in the non-life sectors.
  • The second transformation in 2000 saw the re-entry of a new breed of private companies and also foreign players (with a restricted stake of up to 26 per cent).
  • The third transformation in 2007 saw the safety net of tariff based pricing being withdrawn in general insurance, leading to increased competition. With abiding speculations about the restrictions on foreign ownership being relaxed and the introduction of the new risk based capital norms along with various other changes, the industry is all set to metamorphose into a new landscape.

So, what does this mean in terms of business strategy for Indian insurance? What will be the impact of this anticipated future on today’s strategies? While in the long run, insurers may have to reinvent their present methods of doing business and adopt new and unforeseen initiatives, it would serve them well to optimize their current strategies to adapt to the business environment in the near future. In a data driven industry such as insurance, companies will not only need to compete in terms of their product offerings, but will also need to leverage business intelligence enabled analytics for a competitive edge.

In today’s increasingly competitive scenario, profitable growth is an elusive goal for the insurance industry. Rapid development and deployment of new products and product features, balancing broader distribution channel opportunities, managing risks across the organization, responding to increasingly demanding regulatory and reporting agency demands, and providing more precise pricing levels require effective decisions to be made with greater accuracy, efficiency and transparency. Insurers will need to analyze the impact of optimum non-financial performance in driving optimum financial performance. Instead of reacting to sudden changes, they would need to make accurate forecasts of future performance to plan business strategy.

Optimizing strategies: A driven by intelligent business analytics

In order to optimize their business processes, insurers will have to stabilize their existing operations and then accelerate profitability by understanding their optimum shape, product mix and operations. This translates into well defined strategies for the core areas of Marketing, Cost Management and Risk Management. Their strategies should empower them to: offer the most profitable product mix through the most cost-efficient channels to the most profitable customer; reduce operational losses through efficient claims fraud management and economic claims settlement; and effective risk based capital management. To dramatically improve efficiency and enhance business performance, strategic decisions should be aided by better executive and operational insight, by having the right infrastructure to support, analyze and understand the underlying complexities. The strategic insights gained into these core areas for insurance companies, both in the life as well as the non life sectors, by the use of intelligent business analytics are discussed in more detail in this article.

Marketing needs to focus on customer segmentation and effective distribution networks. Insurers experiencing poorer customer loyalty levels and increased costs, will need to become more customer-centric rather than product-focused. It will become extremely important for insurers to segment customers based on their behavior and potential profitability. Analytics can help them select more accurately, which policies and services to offer to which customers. To increase market share, insurers will have to employ cross-sell and up-sell techniques, facilitated by market basket analysis carried out on the basis of historical data of policies held by clients, as well as details of active policies, customer demographics, claims propensity and other key variables. Data from not only the traditional sources such as the insurers own records, but also sources such as records of the parent organizations can be used to identify various customer segment attributes. Apart from customer segmentation techniques, data mining can also be used to predict the likelihood of policy cancellation in advance, to aid in customer retention analysis by analyzing previous cancellation trends. Simulation models can also be used to generate the probable cash flows and compute the present value of a customer, by estimating the difference between the total amount of revenues from the customer and the expenses for the customer during the whole relationship period, known as the Customer Lifetime Value.

Increased competition is already putting downward pressure on premiums. This coupled with high customer churn and high customer acquisition costs has led insurance companies to implement multi-channel integration strategies. Distribution analytics can enable insurers to conduct “deep dives” into causal factors to answer a variety of questions. Issues such as unmet demand due to improper market assessment, segmentation, positioning and sales support, can be addressed. Similarly compensation and recognition programs influencing leading / lagging sales productivity and agent retention can be designed. All this can help in developing more effective sales channels.

New product development could also be positively influenced by better data analysis. Insurers’ portfolios of new products could be shaped by the new market opportunities which evolve from recent natural catastrophes, the latest technology, feedback from agents and customers, the availability of capital, the actions of competitors and changes in laws. With the use of data mining and predictive analytics, insurers can identify characteristics of individual risks and this will change how insurers see their market. New market opportunities can be identified by evaluating four main parameters. Firstly, by paying attention to the new and emerging needs of the customers. Secondly, by understanding the trends in the global marketplace, both demographically and geographically, and what the insurance implications of those trends might be. Thirdly, by developing new products in accordance with changes in legislation or regulatory environment. Fourthly, identifying new market opportunities by zeroing in on market needs that follow catastrophes.

Getting the right mix of price, claim, channels to market, operations and product differentiation is crucial for success and this can be achieved using business analytics.

Cost Management: Focus on claims fraud prediction and intervention.

As claims are prone to fraud or value inflation, the handling of claims affects the long-term sustainability of the company’s profits. Thus analytical solutions which can help manage the complex claims process effectively and help detect claims fraud by accurately forecasting likely outcomes in order to mitigate the severity of the claim, will be invaluable inputs in decision making and forecasting the loss reserves. Analytics can be used to benchmark claims to detect where padding might have occurred. Most companies have sufficient data to create claims benchmarks and claim value models. To manage fraud, companies can adopt a hybrid approach to detection, using a combination of profiling, rules to filter out fraudulent transactions and advanced analytics software.

Health insurers may develop predictive models to decrease claims costs by analyzing claim characteristics from their own claim files. This can
help them develop benchmark costs for various diseases by regions or by provider type or by severity. Such models can be used to identify highest-cost providers, claims having a higher propensity for fraud or enrollee with significant exposure. A step by step approach towards achieving this:

a) Acquire, load, and cleanse data from both internal and external sources, such as enrollment and TPA records.
b) Develop a methodology for carrying out effective analysis of variables, such as cluster diagnosis codes into easy to analyze disease categories.
c) Analyze and report on correlation of claim variables by geography, age, provider etc.
d) Utilize output to create package rates for treatments, set limits on benefits or undertake focus negotiations with providers.

As a consequence of the above actions, an insurer would be able to identify cost saving opportunities and develop tiered networks or disease specific limits to help achieve targeted loss ratios.

Risk Management: Focus on economic capital and solvency.

As every insurer has limited capital upon which to write new business, profitability is ultimately tied to effective risk management of capital. To adapt to a risk based approach, insurers will have to implement an economic capital regime, where they predict and evaluate the risk profiles of the underwritten business under both best case and worst case scenarios, and thus determine the prudent level of economic capital for sufficient reserves. Simulation techniques and stochastic models for various risks, such as credit, market, liability, group, underwriting and operational risk can be used for scenario and stress testing to determine the optimum economic capital level. This will not only facilitate risk-based capital but also enable real time solvency monitoring. Such a model would be adaptive to the insurer’s evolving environment, by adjusting its parameters as the economic conditions and liabilities change. The capability to accurately project its capital requirements can also be used to continuously monitor solvency since it can also estimate the change in the value of the time zero balance sheet as market conditions move.

An alternative to this approach is to calculate the solvency position on a statutory basis, but this would not measure the true likely response of the insurer to market events. Another possible approach is to use closed-form calculations for the cost of guarantees; however, these would be inaccurate because they do not allow for management actions and would thus result in only very approximate results. Such complex specifications cannot be met by traditional methods and require a state-of-the-art stochastic modeling approach, which is smart, fast and flexible.

Overcoming obstacles in implementation: Data quality and its capabilities

The issue of data has presented several challenges to all insurance industry stakeholders. Being able to conduct an in-depth analysis to provide critical strategic inputs directly depends on the quality of data. Over the past couple of years the awareness about utility of data in building rating structures, product design and flexible pricing systems has increased. However the industry has not fully transferred this knowledge to action. This was partly due to the fact that earlier the data was hard to use, it was frequently of poor quality and incomplete. In addition, different data was captured in different formats, thus making data aggregation and analysis at the insurance company’s end a very daunting task. This has changed over the past few years.

Suitability of data for analytical purposes can be ensured by focusing on two aspects - completeness and standardization. To ensure that the data is complete and accurate, validation features need to be introduced in the business front end and claims processing software tools. In addition to validation checks, it is imperative that the data must be collected in a standardized format. This can be achieved by standardizing the various forms used across the industry.

Insurance Regulatory and Development Authority (IRDA) has recognized that collection and dissemination of reliable and accurate data is important for the insurance industry and has formed the nodal Insurance Information Bureau (IIB). IIB has created a data repository to enable insurance companies, other stake holders and researchers to have easy access to validated data from one source.

However, data quality is only one side of the coin. The other side, supporting IT infrastructure, is again an area where insurers will have to work upon. An insurance IT system should capture and analyze information across all lines of business and risks. In addition, the system should provide analytic capabilities and produce reports for a variety of users. In general, a technology platform for deployment of a business intelligence analytical application should include the following:

  • Data integration technology that can capture data, clean, transform and load data from operational systems into a data warehouse in near real-time to enable offline analysis as well as on-line analytical processing (OLAP).
  • Data warehouse data management technology to manage information held in an operational data store, historical data marts for dimensional analysis, and analytical data stores for scoring and predictive data mining.
  • Dedicated servers to analyze data for multiple user reporting and analytical applications.
  • Tools to allow power users, managers and executives to build reports, slice, dice and drill, and mine data to model, score and predict.

Insurance is a data-rich industry and unfortunately, most of that data is underutilized. The key to gaining a competitive advantage in the insurance industry is found in analyzing this data and getting a greater business insight. Insurers can unlock the intelligence contained in their operational applications - like policy administration, claims management and CRM solutions - through modern data mining and analysis technology. Data mining uses predictive modeling, database segmentation, cluster analysis, neural networks and combinations thereof to quickly answer crucial business questions with greater accuracy. Business needs and business strategy must drive decisions about the structure and functionality of the business intelligence platform including the data warehouse and the data mart to enable this.

In summary, insurers should harness the power of extensive data provided by their business environment, conduct an extensive analysis of the data to model that environment and predict the consequences of alternative actions to guide executive decision making.

Business analytics solutions could be used by insurers to take rapid, effective and precise operational decisions in highly competitive markets, which maximize organizational value and minimize risks.

Submitted By,

Kichawele D. Msuya.

Roll Number: 12157.

MBA-Finance (2009-2011).

How to Lie with Statistics!!! harvinder singh 12083

How to Lie with Statistics!!!
“How to Lie with Statistics is a book written by Darrell Huff in 1954 presenting an introduction to statistics for the general reader. Huff was a journalist who wrote many "how to" articles as a freelancer, but was not a statistician.
The book is a brief, breezy, illustrated volume outlining common errors, both intentional and unintentional, associated with the interpretation of statistics, and how these errors can lead to inaccurate conclusions. It has become one of the best-selling statistics books in history, with over one and a half million copies sold in the English-language edition, even though the monetary examples have become dated because of inflation. It has also been widely translated.
Themes of the book include "Correlation does not imply causation" and "Using Random Sampling". It also shows how statistical graphs can be used to distort reality, for example by truncating the bottom of a line or bar chart, so that differences seem larger than they are, or by representing one-dimensional quantities on a pictogram by two- or three-dimensional objects to compare their sizes, so that the reader forgets that the images don't scale the same way the quantities do.
The original edition contained humorous illustrations by artist Irving Geis. In a UK edition these were replaced with cartoons by Mel Calman.
Source: “http://en.wikipedia.org/wiki/How_to_Lie_with_Statistics

Correlation does not imply causation

"Correlation does not imply causation" (also called "Ignoring a Common Cause" and "Questionable Cause") is a phrase used in science and statistics to emphasize that correlation between two variables does not automatically imply that one causes the other (though correlation is necessary for linear causation, and can indicate possible causes or areas for further investigation... in other words, correlation can be a hint).
The opposite belief, correlation proves causation, is a logical fallacy by which two events that occur together are claimed to have a cause-and-effect relationship. The fallacy is also known as cum hoc ergo propter hoc (Latin for "with this, therefore because of this") and false cause. By contrast, the fallacy post hoc ergo propter hoc requires that one event occur before the other and so may be considered a type of cum hoc fallacy.
In a widely-studied example, numerous epidemiological studies showed that women who were taking combined hormone replacement therapy (HRT) also had a lower-than-average incidence of coronary heart disease (CHD), leading doctors to propose that HRT was protective against CHD. But randomized controlled trials showed that HRT caused a small but statistically significant increase in risk of CHD. Re-analysis of the data from the epidemiological studies showed that women undertaking HRT were more likely to be from higher socio-economic
groups (ABC1), with better than average diet and exercise regimes. The use of HRT and decreased incidence of coronary heart disease were coincident effects of a common cause (i.e., the benefits associated with a higher socioeconomic status), rather than cause and effect as had been supposed.

Usage

In logic, the technical use of the word "implies" means "to be a sufficient circumstance". This is the meaning intended by statisticians when they say causation is not certain. Indeed, p implies q has the technical meaning of logical implication: if p then q symbolized as p → q. That is "if circumstance p is true, then q necessarily follows." In this sense, it is always correct to say "Correlation does not imply causation".
However, in casual use, the word "imply" loosely means suggests rather than requires. The idea that correlation and causation are connected is certainly true; where there is causation, there is likely to be correlation. Indeed, correlation is used when inferring causation; the important point is that such inferences are not always correct because there are other possibilities, as explained later in this article.
Edward Tufte, in a criticism of the brevity of Microsoft PowerPoint presentations, deprecates the use of "is" to relate correlation and causation (as in "Correlation is not causation"), citing its inaccuracy as incomplete.[1] While it is not the case that correlation is causation, simply stating their nonequivalence omits information about their relationship. Tufte suggests that the shortest true statement that can be made about causality and correlation is one of the following:[4]
  • "Empirically observed covariation is a necessary but not sufficient condition for causality."
  • "Correlation is not causation but it sure is a hint."

General pattern

The cum hoc ergo propter hoc logical fallacy can be expressed as follows:
  1. A occurs in correlation with B.
  2. Therefore, A causes B.
In this type of logical fallacy, one makes a premature conclusion about causality after observing only a correlation between two or more factors. Generally, if one factor (A) is observed to only be correlated with another factor (B), it is sometimes taken for granted that A is causing B even when no evidence supports it. This is a logical fallacy because there are at least five possibilities:
  1. A may be the cause of B.
  2. B may be the cause of A.
  3. some unknown third factor C may actually be the cause of both A and B.
  4. there may be a combination of the above three relationships. For example, B may be the cause of A at the same time as A is the cause of B (contradicting that the only relationship between A and B is that A causes B). This describes a self-reinforcing system.
  5. the "relationship" is a coincidence or so complex or indirect that it is more effectively called a coincidence (i.e. two events occurring at the same time that have no direct relationship to each other besides the fact that they are occurring at the same time). A larger sample size helps to reduce the chance of a coincidence, unless there is a systematic error in the experiment.
In other words, there can be no conclusion made regarding the existence or the direction of a cause and effect relationship only from the fact that A and B are correlated. Determining whether there is an actual cause and effect relationship requires further investigation, even when the relationship between A and B is statistically significant, a large effect size is observed, or a large part of the variance is explained.

Examples

B causes A (reverse causation)

The more firemen fighting a fire, the bigger the fire is going to be.
Therefore firemen cause fire.
The above example is simple and easy to understand. The strong correlation between the number of firemen at a scene and the size of the fire that is present does not imply that the firemen cause the fire. Firemen are sent according to the severity of the fire and if there is a large fire, a greater number of firemen are sent; therefore it is rather that fire causes firemen to arrive at the scene.

A causes B and B causes A (bidirectional causation)

Increased pressure results in increased temperature.
Therefore pressure causes temperature.
The ideal gas law, PV = nRT, describes the direct relationship between pressure and temperature (along with other factors) to show that there is a direct correlation between the two properties. For a fixed volume, an increase in temperature will cause an increase in pressure; likewise, increased pressure will cause an increase in temperature. This demonstrates (4) in that the two are directly proportional to each other and not independent functions.

Third factor C (the common-causal variable) causes both A and B

Main article: Spurious relationship
All these examples deal with a lurking variable, which is simply a hidden third variable that affects both clauses of the correlation; for example, the fact that it is summer in Example 3. A difficulty often also arises where the third factor, though fundamentally different from A and B, is so closely related to A and/or B as to be confused with them or very difficult to scientifically disentangle from them (see Example 4).
Example 1
Sleeping with one's shoes on is strongly correlated with waking up with a headache.
Therefore, sleeping with one's shoes on causes headache.
The above example commits the correlation-implies-causation fallacy, as it prematurely concludes that sleeping with one's shoes on causes headache. A more plausible explanation is that both are caused by a third factor, in this case going to bed drunk, which thereby gives rise to a correlation.
Example 2
Young children who sleep with the light on are much more likely to develop myopia in later life.
The former is a recent scientific example that resulted from a study at the University of Pennsylvania Medical Center. Published in the May 13, 1999 issue of Nature,[5] the study received much coverage at the time in the popular press.[6] However, a later study at The Ohio State University did not find that infants sleeping with the light on caused the development of myopia. It did find a strong link between parental myopia and the development of child myopia, also noting that myopic parents were more likely to leave a light on in their children's bedroom.[7][8][9][10] In this case, the cause of both conditions is parental myopia.
Example 3
As ice cream sales increase, the rate of drowning deaths increases sharply.
Therefore, ice cream causes drowning.
The aforementioned example fails to recognize the importance of time in relationship to ice cream sales. Ice cream is sold during the summer months at a much greater rate, and it is during the summer months that people are more likely to engage in activities involving water, such as swimming. The increased drowning deaths are simply caused by more exposure to water based activities, not ice cream.
Example 4
A hypothetical study shows a relationship between test anxiety scores and shyness scores, with a statistical r value (strength of correlation) of +.59.[11] Therefore, it may be simply concluded that shyness, in some part, causally influences test anxiety. However, as encountered in many psychological studies, another variable, a "self-consciousness score," is discovered which has a sharper correlation (+.73) with shyness. This suggests a possible "third variable" problem, however, when three such closely related measures are found, it further suggests that each may have bidirectional tendencies (see "bidirectional variable," above), being a cluster of correlated values each influencing one another to some extent.

Coincidence

With a decrease in the number of pirates, there has been an increase in global warming over the same period.
Therefore, global warming is caused by a lack of pirates.
The example above is used satirically by the parody religion Pastafarianism to illustrate the logical fallacy of assuming that correlation equals causation.
Since the 1950s, both the atmospheric CO2 level and crime levels have increased sharply.
Hence, atmospheric CO2 causes crime.
The above example arguably makes the mistake of prematurely concluding a causal relationship where the relationship between the variables, if any, is so complex it may be labeled coincidental. The two events have no simple relationship to each other beside the fact that they are occurring at the same time. Another possible example is the somewhat jocular Mierscheid Law.

Determining causation

David Hume argued that causality is based on experience, and experience similarly based on the assumption that the future models the past, which in turn can only be based on experience – leading to circular logic. In conclusion he asserted that causality is not based on actual reasoning: only correlation can actually be perceived.[12]
Intuitively, causation seems to require not just a correlation, but a counterfactual dependence. Suppose that a student performed poorly on a test and guesses that the cause was his not studying. To prove this, one thinks of the counterfactual – the same student writing the same test under the same circumstances but having studied the night before. If one could rewind history, and change only one small thing (making the student study for the exam), then causation could be observed (by comparing version 1 to version 2). Because one cannot rewind history and replay events after making small controlled changes, causation can only be inferred, never exactly known. This is referred to as the Fundamental Problem of Causal Inference – it is impossible to directly observe causal effects.[13]
A major goal of scientific experiments and statistical methods is to approximate as best as possible the counterfactual state of the world.[14] For example, one could run an experiment on identical twins who were known to consistently get the same grades on their tests. One twin is sent to study for six hours while the other is sent to the amusement park. If their test scores suddenly diverged by a large degree, this would be strong evidence that studying (or going to the amusement park) had a causal effect on test scores. In this case, correlation between studying and test scores would almost certainly imply causation.
Well-designed experimental studies replace equality of individuals as in the previous example by equality of groups. This is achieved by randomization of the subjects to two or more groups. Although not a perfect system, the likeliness of being equal in all aspects rises with the number of subjects placed randomly in the treatment/placebo groups. From the significance of the difference of the effect of the treatment vs. the placebo, one can conclude the likeliness of the treatment having a causal effect on the disease. This likeliness can be quantified in statistical terms by the P-value.
When experimental studies are impossible and only pre-existing data are available, as is usually the case for example in economics, regression analysis can be used. Factors other than the potential causative variable of interest are controlled for by including them as regressors in addition to the regressor representing the variable of interest. False inferences of causation due to reverse causation (or wrong estimates of the magnitude of causation due the presence of bidirectional causation) can be avoided by using explanators (regressors) that are necessarily exogenous, such as physical explanators like rainfall amount (as a determinant of, say, futures prices), lagged variables whose values were determined before the dependent variable's value was determined, instrumental variables for the explanators (chosen based on their known exogeneity), etc. See Causality#Economics. Spurious correlation due to mutual influence from a third, common causative, variable, is harder to avoid: the model must be specified such that there is a theoretical reason to believe that no such underlying causative variable has been omitted from the model; in particular, underlying time trends of both the dependent variable and the independent (potentially causative) variable must be controlled for by including time as another independent variable.

http://en.wikipedia.org/wiki/Correlation_does_not_imply_causation

Discriminant analysis and the statistics involved……By Vaibhav Jaiswal


Discriminant analysis is a technique for analysis data when the criterion or dependent variable is categorical in nature and predictor or independent variables are interval in nature.
The objectives of discriminant analysis are as follows

  1. Developments of discriminant function, which will best discriminate between the categories of the criterion or dependent variable (groups).
  2. Examination of whether significant difference exists among the groups, in terms of predictor variables
  3. Determination of whether the predicator variables contribute  to most of the inter-group differences
  4. Classification of cases to one of the groups based on the values of the predictor variables
  5. Evaluation of accuracy of classification

Discriminant analysis techniques are described by the number of categories possessed by the criterion variable. When the criterion variable has two categories, the technique is called two group discriminant analysis. When three or more categories are involved, the technique is referred to as multiple discriminant analysis.

Discriminant analysis model:

The discriminant analysis model involves linear combination of the following form
L = b1x1 + b2x2 + . + bnxn + c , where the b's are discriminant coefficients, the x's are the input variables or predictors, L is the discriminant score and c is a constant



The coefficients or weights (b) are estimated so that the groups differ as much as possible on the values of discriminant function. This occurs when the ration of between group sum of square to with-in group sum of squares for the discriminant function is at the maximum. Any other linear combination will result in smaller ratio.
I am giving a brief geometrical exposition of two group discriminant analysis. In the figure 1 we have two groups, A and B, and each member of these groups is measured on two variables X and Y which are the two axes. Members of A are denoted in ellipse A and members of B are denoted by ellipse B. the resultant ellipse encompasses some specified percentage of members in each group. A straight line is drawn through the two points where the ellipse intersects and then projected to new axes, I. The overlap between the univariate distributions A and B is smaller than would be obtained by any other line drawn through the ellipse representing the members of the groups. Thus the group differs maximum on I axes.

Some important statistics associated with discriminant analysis
  1. Canonical Correlation: it is the extent of association between the discriminant score and the groups.
  2. Classification matrix: it contains the number of correctly classified and misclassified cases.
  3. Discriminant function coefficient: these are the multipliers of the variables, when the variables are in the original units of measurement
  4. Eigen value: it is the ratio of between groups to with in group sum of squares. Larger eigen value means superior function

Use of discriminant analysis

Discriminant analysis is used in marketing at a very great extent in distinguishing the consumer behavior, store characteristics….etc.
But apart from being intensively in marketing discriminant analysis is also used in areas like finance and operations. I will give one example from operations where discriminant analysis is used effectively.

By using the discriminant analysis function within statistical software packages, managers can classify future production volume and work to separate defective products on the production line before they are delivered to customers. This study was conducted by David Lengacher.

Discriminant Analysis in Action

 

The following example outlines the creation of a data set and the use of discriminant analysis. A director of operations has been plagued with the same problem for months. The process is highly capable and internal scrap is low, but roughly 3 percent of the shipped products fail when the customer installs them into an assembly. Although the parts are within customer specifications, they are still failing. It may take months for a detailed chemical analysis to provide the director with a true root cause, but even then, there are no guarantees. However, upon graphing the last batch of customer returns against a random sample of parts ready for shipment, a clear pattern emerges (Figure 1).



It is much cheaper to scrap production inside the facility than it is to be charged for the customer’s complete assembly that failed because of a defective product. Therefore, the director’s goal is to be able to separate those units that are likely to fail as soon as they come off the production line.
At first glance, however, it appears that the situation is hopeless. Customer returns and normal parts are overlapping in the scatter plot. In situations like this, statistical analysis software can be of great help. By using software with a discriminant analysis feature, practitioners can separate production with a high degree of accuracy and minimal cost. 
The software will analyze the full data set and produce two group membership functions. These equations are constructed in a way that will minimize the misclassification of the sample data set. As shown in Figure 2, the functions are able to achieve a 91 percent accuracy level. Because each data point goes through both membership functions, practitioners can assign the data point to the function that returns the higher value. This solution can mitigate risk until the director of operations can identify the true root cause of the failures.
 Thanks to the findings of the discriminant analysis, the director has a statistically valid procedure for classifying future production volume. The next step is to monitor the performance of this new procedure. Although the next batch of returns should be smaller because of the implementation of the new decision support tool, practitioners will perform another round of discriminant analysis when they receive the next batch of parts returned by the customer. The accuracy of this tool will be directly proportional to the sample size used to construct it. In every sense, creating this decision support  tool is a continuous improvement project in itself.
Thus the example stated above shows how discriminant analysis can be used in making better decisions.
References
1.      Malhotra, N.K. (2007), “Marketing research an applied orientation.” 5th edition.
2.      David Langacher, “Discriminant Analysis Can Minimize Returned Products”, Available online at http://www.isixsigma.com/index.php?option=com_k2&view=item&id=709:discriminant-analysis-can-minimize-returned-products&Itemid=204#
3.      Wikipedia




By VAIBHAV JAISWAL


Tv show Numb3rs using statistics

Getting students engaged on a highly quantitative subject such as statistics or analytics is not an easy task.
Here I would like to share a TV series named Numb3rs which has been very popular and has attracted masses. I myself have been a huge fan of this show.
It’s about two brothers, the older one an FBI agent who constantly consults his younger brother, a math professor, in solving crimes. Each crime is solved using different mathematical analysis tools be it on profiling criminals, predicting their movements, identifying locations of victims, or some IT stuff such as image enhancement of surveillance videos or hacking computer systems.

Numb3rs opening line for every episode:
We all use math every day; to predict weather, to tell time, to handle money. Math is more than formulas or equations; it's logic, it's rationality, it's using your mind to solve the biggest mysteries we know.
One of the episodes in season 2 makes use of Linear discriminant analysis

Linear discriminant analysis (LDA), is sometimes referred to as Fisher's linear discriminant (although Fisher's original article The Use of Multiple Measures in Taxonomic Problems (1936) actually describes a slightly different discriminant, which does not make some of the assumptions of LDA such as normally distributed classes or equal class covariances). LDA is typically used as a feature extraction step before classification.

The episode vividly portrays the intricacies of face recognition using discriminant analysis. Despite the availability of commercial systems, face recognition continues to be an active topic in computer vision research. Current face recognition systems perform well under nearly ideal circumstances, but tend to suffer when variations in expression, illumination, decoration (i.e.,
glasses, facial hair), and/or pose are present. Most current face recognition research aims to improve recognition performance in the presence of such confounding factors. Face
recognition methods can be classified broadly into two : feature- or template-based.
a modular linear discriminant analysis (LDA) approach for face recognition. A set of observers is trained independently on different regions of frontal faces and each observer projects face images to a lower-dimensional subspace.
These lower-dimensional subspaces are computed using LDA methods, including a new algorithm that we refer to as direct,weighted LDA or DW-LDA. DW-LDA combines the advantages of two recent LDA enhancements, namely direct LDA (DLDA) and weighted pairwise Fisher criteria. Each observer performs recognition independently and the results are combined
using a simple sum-rule. Experiments compare the proposedapproach to other face recognition methods that employ linear dimensionality reduction. These experiments demonstrate that the modular LDA method performs significantly better than other linear subspace methods. The results also show that D-LDA does not necessarily perform better than the well-known principal component analysis followedby LDA approach.

Statistics plus videos plus numb3rs not a bad option !

Posted by :
Gayathri
12137
Marketing batch

Perceptual mapping - Fashion industry


We have seen in detail what perceptual map is all about and its various applications.


Having chosen the Fashion Industry segment, the sector as it is an evergreen and still booming sector. Fashion is a segment where perceptions and opinions play a major role in determining the sales figures . Understanding the complexity associated with the different attributes and brands can be made easier by developing a visual representation of each market. These are known as perceptual maps and they are used to determine how various brands are perceived according to the key attributes that customers value. In addition, it is possible to determine and map how customers see an ideal brand, based on the key attributes, and from this see how far away a brand is from occupying the ideal position.

Understanding the complexity associated with the different attributes and brands can be made easier by developing a visual representation of each market. These are known as perceptual maps and they are used to determine how various brands are perceived according to the key attributes that customers value. In addition, it is possible to determine and map how customers see an ideal brand, based on the key attributes, and from this see how far away a brand is from occupying the ideal position.

Perceptual mapping represents a geometric comparison of how competing products are perceived (Sinclair and Stalling, 1990). One thing to note is that the closer products/brands are clustered together on a perceptual map, the greater the competition. The further apart the positions, the greater the opportunity for new brands to enter the market, simply because the competition is less intense. For example, in fashion retailing there are numerous brands in the marketplace all competing with each other across differing core attributes, brand reputation, store presence, price, and clothing quality or trendy and stylish. To show how the differing brands might be positioned relative to each other using the attribute scores for each brand of fashion retailer we can measure and map the brand positioning for the respective brands. Figure below shows the positioning of a number of fashion retailers using the dimensions of price, store presence, and trendy and stylish.




Given the distance between the differing brands on the perceptual map,

Figure below shows us that there is a relatively high level of differentiation in brand positioning between the retailing brands Invogue and Fast Fashion. However, in contrast there is a low level of differentiation between Budget Fashion and Just Fashion. Given the distance of the differing brands from each attribute, Figure below also shows us that Invogue and Ladies Manor are more closely associated with the attribute of trendy and stylish; Budget Fashion, Just Fashion, and the Clothing Superstore with the attributes of price and value for money; and Stephanie’s and Fast Fashion with the attributes of in-store service, opening hours, and store convenience.

Determining attribute importance and mapping the brands across these attributes, we can discover how our brand and competing brands are perceived in the marketplace.

It is very rare that using just two attributes adequately reflects the diversity of opinion and preferences of the target market. Using multidimensional scaling techniques it is possible to add further attributes and create a composite picture of the main segments that constitute a market.

Perceptual mapping can provide significant insight into how a market operates.

For example, it provides marketers with an insight into how their brands are perceived and it also provides a view about how their competitors’ brands are perceived. In addition to this substitute products can be uncovered, based on their closeness to each other (Day et al., 1979). All of the data reveal strengths and weaknesses that in turn can assist strategic decisions about how to differentiate on the attributes that matter to customers and how to compete more effectively in the target market.

Source: http://www.oup.com/uk/orc/bin/9780199290437/baines_ch06.pdf


Himanshu Arora

Roll - 12024

Marketing Batch 2009 - 11