How does stepwise selection work in feature selection?
It uses L1 or L2 regularization to shrink irrelevant feature coefficients to zero.
It transforms the original features into a lower-dimensional space while preserving important information.
It ranks features based on their correlation with the target variable and selects the top-k features.
It iteratively adds or removes features based on a statistical criterion, aiming to find the best subset.
How do GLMs handle heteroscedasticity, a situation where the variance of residuals is not constant across the range of predictor values?
They require data transformations to stabilize variance before analysis.
They use non-parametric techniques to adjust for heteroscedasticity.
They ignore heteroscedasticity as it doesn't impact GLM estimations.
They implicitly account for it by allowing the variance to be a function of the mean.
What distinguishes a random slope model from a random intercept model in HLM?
Random slope models allow intercepts to vary, while random intercept models don't.
Random slope models handle categorical variables, while random intercept models handle continuous variables.
Random slope models are used for smaller datasets, while random intercept models are used for larger datasets.
Random slope models allow slopes to vary, while random intercept models don't.
When using Principal Component Analysis (PCA) as a remedy for multicollinearity, what is the primary aim?
To increase the sample size of the dataset
To create new, uncorrelated variables from the original correlated ones
To remove all independent variables from the model
To introduce non-linearity into the model
What does a Variance Inflation Factor (VIF) value greater than 10 generally suggest?
Severe multicollinearity
Heteroscedasticity
Perfect multicollinearity
No multicollinearity
How do Generalized Linear Models (GLMs) extend the capabilities of linear regression?
By assuming a strictly linear relationship between the response and predictor variables.
By limiting the analysis to datasets with a small number of observations.
By enabling the response variable to follow different distributions beyond just normal distribution.
By allowing only categorical predictor variables.
If a predictor has a p-value of 0.02 in a multiple linear regression model, what can you conclude?
The predictor is not statistically significant.
The predictor explains 2% of the variance in the outcome.
The predictor has a practically significant effect on the outcome.
The predictor is statistically significant at the 0.05 level.
How do polynomial features help in capturing non-linear relationships in data?
They introduce non-linear terms, allowing the model to fit curved relationships.
They make the model less complex and easier to interpret.
They convert categorical variables into numerical variables.
They reduce the impact of outliers on the regression line.
What is a key limitation of relying solely on Adjusted R-squared for model evaluation in linear regression?
It doesn't provide information about the magnitude of prediction errors.
It can be misleading when comparing models with different numbers of predictors.
It is difficult to interpret.
It is highly sensitive to outliers.
What is a primary risk of using high-degree polynomials in Polynomial Regression?
It can lead to overfitting, where the model learns the training data too well but fails to generalize to new data.
It always improves the model's performance on unseen data.
It reduces the computational cost of training the model.
It makes the model too simple and reduces its ability to capture complex patterns.