A. Decide how to treat missing data
1. List wise deletion may shrink the sample size too much.
2. Pairwise deletion may make the analysis impossible depending on missing data patterns.
3. Either list wise or pairwise deletion has the potential to bias results; consider imputation.
4. Be prepared to run multiple analyses employing different missing data strategies.
B. Assess suitability of data for factor analysis
1. Distributions should not be excessively far from normal; may need to transform.
2. Relationships need to be essentially linear; may need to transform.
3. Sampling adequacy (often optional). Desirable to have high RS relative to partial RS.
b. Anti-image correlations: for individual variables, use the same standard as for KMO, and this can help decide inclusion/exclusion.
C. Extraction method
1. Use principal components only if looking to use all information in pure data reduction (i.e., if it is reasonable to assume ~0% of the info is measurement error).
2. Can use maximum likelihood extraction, and its associated goodness of fit test for number of factors, if
a. variable distributions are fairly close to normal, i.e., skewness less than 2 and kurtosis less than 7, or if willing to transform or drop non-normal variables.
b. sample size is moderate: goodness of fit test is sensitive to small or large n.
3. In most social science cases, principal axis factoring works well.
D. Number of factors in a principal axis factor solution
1. Plan on needing to try multiple solutions.
2. Begin by extracting factors from all initial eigenvalues of at least 0.7. (Much research has shown that including only initial eigenvalues greater than equal 1 can under- or over-estimate the number of factors.) Once you see how solutions turn out after extraction and rotation, you can specify a minimum-eigenvalue or number-of-factors criterion based on something more informed than just the Kaiser Guttman rule of "eigenvalues greater than equal 1."
3. Scree test, visual or numerical, is best performed on extracted eigenvalues, rather than initial (although initial is what is displayed in SPSS).
4. Factor correlations should not be too high. If they are, and if largest initial eigenvalue is more than 10 times as large as the next, consider a one-factor solution (Harman; Kriebeck; Tracey).
b. Consider Terence (T. J. G.) Tracey's method of partialing out the first un-rotated factor as a "general factor" attributable to bias, e.g., to social desirability or common method bias.
5. Each factor should have at the very least 2 and probably 3 variables with high loadings.
6. Only retain those factors that are interpretable
E. Rotation method
1. Use oblique rotation as the default. There is no nothing to lose: if correlations among factors are very low, you can always switch to varimax (in which case results should hardly change anyway). But by starting with varimax you risk ignoring meaningful correlations.
2. To display two factors' relationship after finalizing a solution, can use varimax to plot the two clusters of variables and show the factors' correlation (which will be represented by the angle between them, with r equal to the cosine).
Source: An article on factor analysis written by Roland B. Stark in 19th February, 2007
Website link: www.integrativestatistics.com
Submitted By
Dipayan Kabiraj
12133
Group 5
SIBM Bangalore
No comments:
Post a Comment