← Back to list
AI & Data
#다중공선성#회귀분석#VIF#PCA#Ridge#132회
Last updated · 2026-09-30

Multicollinearity

1. Overview

A. Definition

A phenomenon in multiple regression where strong linear correlation exists among the independent (explanatory) variables, making the estimation of individual regression coefficients unstable and distorting their interpretation.

The essence of multiple regression analysis is to isolate and estimate, as a coefficient, the "pure, unique effect" of a single variable on the dependent variable while holding all other variables constant. This is precisely why a coefficient βⱼ is interpreted as "the change in Y per one-unit change in Xⱼ when the other variables are held fixed." For this separation of effects to hold statistically, however, each independent variable must carry sufficiently independent information. If two variables move together in almost the same direction, the regression equation loses any basis for deciding "whether A or B produced the change in the dependent variable."

Mathematically, the problem reduces to the columns of the design matrix X approaching linear dependence. The least-squares estimator is given by β̂ = (XᵀX)⁻¹Xᵀy; when the columns approach linear dependence, XᵀX approaches a singular (non-invertible) matrix, its determinant converges toward zero, and the elements of the inverse explode. Because the variance of the coefficient estimates is proportional to the diagonal elements of this inverse, the variance of β̂ inflates and the estimates swing wildly with even slight changes in the sample. In other words, the key is to understand that multicollinearity is not a problem caused by errors in the data but an identification problem arising from the structure of the variables themselves.

B. Problems and the need to address it

Severe multicollinearity produces three symptoms. First, the standard errors of the coefficients grow (variance inflation), so t-statistics shrink and a genuinely important variable may be wrongly judged "statistically insignificant." Second, coefficients may be estimated with a sign reversed against theory or with an unrealistically large value. For example, entering income and consumption as explanatory variables simultaneously causes their coefficients to offset and compensate for each other, so the signs waver. Third, even a slight change in the sample or adding/removing a single variable causes coefficients to change dramatically, collapsing the stability of the model.

What is interesting is that even in such situations, the overall predictive power of the model (R², the F-test) can remain intact. It merely fails to apportion the contribution of individual variables; the prediction produced jointly by the correlated variables is still valid. This observation is the starting point of any response strategy. That is, if the goal is to "interpret or make inferences about the influence of variables," it must be addressed, but if the goal is "predictive accuracy only," it can often be left alone or managed with regularization. Defining the analytical purpose first is thus the first button to fasten in dealing with multicollinearity.

2. Causes and Mechanism

To understand the causes of multicollinearity, we must trace back through the data-generating process to ask "why the columns of the design matrix approach linear dependence." Below is a conceptual diagram showing the overall structure by which causes lead to coefficient instability.

flowchart LR
  A["Variable design stage"] --> B["Redundant/derived variables entered"]
  A --> C["Naturally correlated variables"]
  A --> G["Small sample/narrow distribution"]
  B --> D["Columns of design matrix near linear dependence"]
  C --> D
  G --> D
  D --> E["XᵀX near non-invertible"]
  E --> F["Inverse elements explode → coefficient variance inflates"]
  F --> H["Standard errors↑ · sign reversal · instability"]

First, redundant variables—entering the same information twice with only the unit or scale changed—are the most serious. If height is entered in both cm and inches, the two columns are exactly proportional, producing perfect collinearity; in this case XᵀX is completely non-invertible and coefficient estimation itself becomes impossible. The dummy variable trap, where every category of a categorical variable is included and becomes perfectly dependent with the intercept, is a case of the same nature.

Second, derived or composite variables built arithmetically from existing variables become structurally entangled with the originals. For example, entering "total score" together with "Korean, English, and math scores" makes the total perfectly dependent because it is the sum of the three subjects, and entering revenue together with unit price and quantity creates a similar problem. Such derived variables add no new information while inducing collinearity.

Third, natural correlation in the real world where the data arise is the most common and the trickiest to handle in practice. Variables that inherently move together—income and consumption, housing floor area and number of rooms, age and career length—are hard to remove and require judgment about which to keep. On top of this, when the sample is small or the distribution of the explanatory variables is narrow (e.g., observing only a particular income band), the correlation among variables is coincidentally exaggerated and the collinearity worsens.

Cause Example Nature Severity
Redundant variable Height (cm) and height (inch) entered together Perfect collinearity Very high (fix by removal)
Dummy variable trap All category dummies + intercept Perfect dependence High (drop one category)
Derived/composite variable Total = sum of subjects, revenue = price×quantity Structural dependence High
Natural correlation Income-consumption, area-rooms Approximate collinearity Judgment needed
Small sample/narrow distribution Small sample, narrow observation range Coincidental·structural mix Medium

3. Diagnostic Methods

The core diagnostic question is "how well is a given variable explained by a linear combination of the remaining variables." The representative metric that quantifies this is the VIF (Variance Inflation Factor). Letting Rᵢ² be the coefficient of determination from regressing the i-th variable on the remaining independent variables, VIFᵢ = 1 / (1 − Rᵢ²). The meaning of this formula is intuitive. When Rᵢ² is near 0 (not explained by other variables), VIF is near 1, and when Rᵢ² is near 1 (almost fully explained by other variables), VIF diverges to infinity. For example, if Rᵢ² = 0.9 then VIF = 1/(1−0.9) = 10, which means the coefficient variance of that variable has inflated tenfold compared with no collinearity at all. Accordingly, the standard error grows by √10 ≈ 3.16 times.

Below is a detailed process diagram showing the diagnostic procedure step by step.

flowchart TD
  S["Fix candidate explanatory variables"] --> C["① Correlation matrix · check pairwise correlation"]
  C --> C1{"±0.8 or more?"}
  C1 -->|Suspicious| V["② Compute VIF · capture multivariate relations"]
  C1 -->|None| V
  V --> V1{"VIF ≥ 10?"}
  V1 -->|Yes| K["③ Condition index · diagnose severity"]
  V1 -->|No| OK["Collinearity mild · proceed"]
  K --> K1{"Condition index 30 or more?"}
  K1 -->|Yes| ACT["Apply remedies"]
  K1 -->|No| OK

Because the correlation matrix looks only at pairs of two variables, it can miss multicollinearity where three or more variables are entangled together. For example, when three variables are entangled as X₃ ≈ X₁ + X₂, the pairwise correlations X₁-X₃ and X₂-X₃ may each be modest, yet entering all three together produces severe collinearity. Because of this limitation, VIF and the condition index provide supplements. The condition index is defined as the square root of the ratio of the largest to the smallest eigenvalue of XᵀX (or the correlation matrix); the larger the value, the closer the matrix is to singular. Conventionally, a condition index of 10–30 is moderate and above 30 indicates severe collinearity.

Diagnosis is safer when the three methods are applied in stages rather than concluding from a single metric. First screen out obvious pairwise problems with the correlation matrix, then quantify multivariate relations with VIF, and for borderline cases check even the condition index and eigenvalue structure. For instance, a variable whose VIF sits around the threshold of 10 at 8–9 should be handled conservatively for a small sample or an analysis where causal interpretation matters, while for large-scale prediction-oriented data it can be absorbed through regularization—varying the judgment according to context.

Method Criterion Principle Limitation
Correlation coefficient |r| ≥ 0.8 between independent variables Captures pairwise linear relations Misses relations among 3+ variables
VIF/tolerance VIF ≥ 10 (strict 5), tolerance = 1/VIF ≤ 0.1 Degree to which one variable is explained by the rest Threshold is conventional
Condition index Severe at 30 or above Square root of max/min eigenvalue ratio Interpretation requires experience

4. Remedies

The most important point is that the remedy depends on the previously defined analytical purpose (interpretation vs. prediction) and on the cause (redundancy vs. natural correlation). There is no single correct answer; it is a matter of judging the trade-offs case by case.

When interpretation is the goal, the most direct method is variable removal. From a highly correlated pair of variables, one removes the side that is less important in the domain or less accurately measured. However, removing a theoretically necessary variable can introduce omitted variable bias that distorts the remaining coefficients, so one must weigh the balance between "resolving collinearity" and "inducing bias" carefully. If information loss is a concern, principal component analysis (PCA) or factor analysis is used to bundle the correlated variables and reconstruct them into mutually uncorrelated new axes. PCA removes collinearity at the source, but at the cost that the new axes (principal components) lose the meaning of the original variables, so interpretability drops.

When prediction is the goal, regularized (penalized) regression that shrinks coefficients toward zero is practical. Ridge regression (L2 penalty) penalizes the sum of squared coefficients, greatly lowering coefficient variance and stabilizing estimates even under collinearity, while keeping all variables rather than removing them. Lasso regression (L1 penalty) penalizes the sum of absolute values of coefficients, driving some coefficients of correlated variables exactly to zero and thus producing an automatic variable selection effect. Elastic Net, combining the two, is strong at jointly selecting or excluding groups of correlated variables, so it is often used for high-dimensional problems with many, highly correlated variables such as genomic data.

In addition, enlarging the sample to increase the amount of information mitigates coincidental correlation, and collinearity caused by interaction or polynomial terms can be substantially reduced by centering, which shifts variables to their mean.

Remedy Content Suitable situation Cost·caution
Variable removal Remove one of the highly correlated variables Redundancy·perfect collinearity, interpretation goal Risk of omitted variable bias
Dimensionality reduction (PCA) Reconstruct into uncorrelated components Information preservation·many correlations Interpretability drops
Regularized regression Ridge(L2)·Lasso(L1)·Elastic Net Prediction goal·high dimension Accept slight bias
Data augmentation·centering Enlarge sample, centering Small sample·interaction terms May not be a fundamental fix

5. Case Comparison

Let us see the difference with concrete numbers. Suppose a house-price prediction model enters "net floor area (㎡)" together with "number of rooms," and the VIF of area comes out to 8.5. The two variables are strongly naturally correlated, yet each gives independent information about price (for the same area, more rooms implies different preferences), so it is wasteful to remove one outright. Here, if prediction is the goal, stabilizing the two coefficients with Ridge is the reasonable choice. Conversely, if one must causally interpret "the net effect of adding one room in a policy report," one must interpret the room-count coefficient cautiously after controlling for area, or reconsider the design. For the same data, the prescription diverges by purpose.

As another example, in marketing data, entering "TV advertising spend" together with "total advertising spend" (total = TV + online + offline) often produces near-perfect collinearity that flips the sign of the TV coefficient to negative. This is a textbook case of derived-variable redundancy and is resolved by dropping total spend or constructing only per-channel variables. Tree-based models (random forests, XGBoost) work by branching rules and are relatively insensitive to collinearity, so when pure prediction is the goal, the model choice itself can become the remedy.

6. In Depth: Recent Trends from the Machine Learning and Regularization Perspective

In traditional statistics, multicollinearity was treated as a "defect to be removed," but in modern machine learning the perspective has shifted to something to be managed with regularization and feature engineering. Standard tools such as scikit-learn and statsmodels provide VIF computation and Ridge/Lasso/Elastic Net out of the box, and choosing the regularization strength (λ) in a data-driven way via cross-validation has become standard practice. In particular, in high-dimensional (p ≫ n) situations where the number of variables p exceeds the sample size n (genomics, text, sensor data), collinearity is inevitable, so L1/L2 regularization and dimensionality reduction—rather than variable removal—are effectively the only practical solutions.

From the MLOps and feature-store perspective, prevention is emphasized over after-the-fact remedy. At the stage of registering features in a feature store, highly correlated derived features are standardized and documented, and from a data-governance standpoint the reproduction of redundant features is prevented. Moreover, in the field of causal inference, collinearity has been reinterpreted as a fundamental limit of causal identification, strengthening a move toward identification strategies such as experimental design and instrumental variables (IV) rather than simple correlation control.

7. Considerations and Implications

  • Purpose-first judgment: Multicollinearity is not a statistical defect but a choice problem depending on the analytical purpose. For services that need only predictive accuracy (recommendation, demand forecasting), allowing collinearity and managing it with regularization is reasonable, but for analyses that require policy or causal interpretation (identifying factor effects), it must be removed or reconstructed.
  • Link with model selection: Linear and logistic regression are sensitive to collinearity, whereas tree-based and regularized models are relatively insensitive. Therefore, choosing an algorithm suited to the nature of the problem can itself be a first line of defense, and ensembles and regularization are natural response mechanisms.
  • Prevention is best: Rather than after-the-fact diagnosis and prescription, designing variables based on domain knowledge so that redundant and derived variables are not entered in the first place is the most effective. This connects directly to feature-store standardization, data governance, and reproducible pipeline management.
  • Relativity of thresholds: Criteria such as VIF ≥ 10 and condition index ≥ 30 are not absolute laws but conventions. Apply them flexibly according to sample size, model purpose, and industry practice, while documenting the basis for judgment and the processing history to secure reproducibility and auditability.
  • Boundary with causal inference: Controlling collinearity only reduces correlation; it does not guarantee causation. When causal effects are required, a mature professional-engineer approach goes beyond regression techniques to combine identification strategies such as experimental design, quasi-experiments (natural experiments), and instrumental variables.

References


In one line: Multicollinearity is an identification problem in which strong linear correlation among independent variables brings the design matrix XᵀX close to non-invertible and destabilizes regression coefficient estimation; it is diagnosed via VIF (≥10) and condition index (≥30) and addressed—according to the analytical purpose (interpretation vs. prediction)—by variable removal, PCA, and Ridge/Lasso/Elastic Net, with the modern view shifting from after-the-fact removal to management through regularization and feature design.