Search This Blog

24.1.11

Prediction Analytics

Robert Nozick once said, “There is no justifiable prediction about how the hypothesis will hold up in the future; its degree of corroboration simply is a historical statement describing how severely the hypothesis has been tested in the past.”

A finance manager has to predict the future EPS just like a marketing manager has to predict the future revenues. An operations manager has to predict the future inventory requirements just like a hr manager has to predict the future attrition rate and finally a GM has to predict the overall growth of the firm and all of them are responsible to someone or the other. But what if the predictions go wrong?

Prediction modeling is the process by which data is modeled and diagnosed to try to best predict the probability of an outcome. But what is the data that is being referred here? The companies realized the importance of this ‘data’ sometime in 1960s. Since then, the trend of data mining and warehousing has picked up in a big way. Almost all the companies across industries started storing and accumulating the data in the form of past performances, demographics of the society, research papers etc. This data is today being used to predict the future trends. Take the example of the banking industry. They use past information related to credit history, job information, loan history and other parameters about a person to determine his credit worthiness and therefore predict whether he or she will default in the future. This is a classic example of prediction analytics, a topic that has been the talk of the day since the past two days at college during the SPSS workshop.

Apart from predicting future behavior, predictive models also anticipate the consequences of change. According to Victor Holman (Performance Management Expert at Lifecycle Performance Professionals), there are three main types of models associated with Predictive analytics. First being the prediction model, second being the descriptive model and third being the decision model.

Predictive analytics' central building block is the predictor, a single value measured for each customer. For example, 'most recent', which is based on the number of weeks since the customer's last purchase, has higher values for more recent customers. This predictor is usually a reliable campaign response predictor: you will receive more responses from those customers more highly ranked by 'most recent'. That means that if you contact your customers in order of 'most recent' - first, call the most-recent customer; next, call the next-most-recent customer; and so on - you will improve your response rate. For each prediction goal, there are an abundance of predictors that will help rank your customer database. For example, consider a customer's online behavior: Customers who spend less time logged on may be less likely to renew their annual subscription. In this case, retention campaigns can be cost-effectively targeted to customers with a low monthly usage predictor value.

Descriptive models quantify the relationships between data in order to classify customers into groups. While predictive models focus on predicting one customer's behavior, descriptive models identify relationships between several customers or products. Descriptive models do not predict a target value, but focus more on the intrinsic structure, relations, interconnectedness, etc. The example that I mentioned earlier about the banks using past demographic data to predict the creditworthiness of the person, follows a descriptive model.

On the first day of the workshop, we were given a small training on Cluster analysis. It is a descriptive modeling technique that identifies clusters embedded in the data. A cluster is a collection of data objects that are similar in some sense to one another. Another descriptive modeling technique’s mention that I found on Victor Holmon’s blog is the k-means algorithm. K-means algorithm is a distance-based clustering algorithm that partitions the data into a predetermined number of clusters (provided there are enough distinct cases). The k-means algorithm works only with numerical attributes. Distance-based algorithms rely on a distance metric (function) to measure the similarity between data points.

Finally, Decision models describe the relationship between all decision elements and predict the results of decisions, allowing you to try different scenarios, and optimize results. Decision Support Systems and Management Information Systems at hospitals use predictive analysis in the health care industry to determine at risk patients and sometimes to determine which course of action would be best given a multiple array of variables. Rational decision models are based around a cognitive judgment of the pros and cons of various options. It is organized around selecting the most logical and sensible alternative that will have the desired effect. The decisions are normally organized through a detailed analysis of alternatives and a comparative assessment of the advantages of each. Weighted criteria scoring are an example of rational decision models.

To sum up, organizations these days have to spend considerable amount of resources on essential tools of survival like data mining and prediction analytics because the future is highly uncertain.

Submitted by:
Kunal Shah
Roll no. 12029
Finance Batch
SIBM Bangalore

(Certain text has been taken from Victor Holman’s knoll on the same topic)
http://knol.google.com/k/victor-holman/three-basic-predictive-analysis-models/2650srcv52e8t/62#

No comments:

Post a Comment