lm(score ~ hours + motivation, data = test_study2)
Data Analysis for Psychology in R 2
Department of Psychology
University of Edinburgh
2026–2027
|
Introduction to linear models (with Dr. Patrick Sturt) |
Introduction to linear regression |
| Interpreting linear models | |
| Testing predictors and evaluating linear models | |
| F-tests and model comparison | |
| Practice analysis: Linear models | |
|
Extending linear models (with Dr. Elizabeth Pankratz) |
Categorical predictors and treatment coding |
| Sum coding | |
| Assumptions and diagnostics | |
| Bootstrapping and confidence intervals | |
| Practice analysis: Categorical predictors (led by Patrick) |
|
Interactions (with Dr. Elizabeth Pankratz) |
Mean-centering and numeric/categorical interactions |
| Numeric/numeric interactions | |
| Categorical/categorical interactions | |
| Testing simple effects and correcting for multiple comparisons | |
| Practice analysis: Interactions (led by Patrick) | |
|
Logistic regression (with Dr. Elizabeth Pankratz) |
Probabilities and log-odds |
| Modelling binary outcomes with logistic regression | |
| Interactions, assumptions, diagnostics, comparisons | |
| Practice analysis: Logistic regression (led by Patrick) | |
| Report feedback, exam prep, full-course Q&A |
Understand how to interpret significance tests for \(\beta\) coefficients
Understand how to calculate and interpret \(R^2\) and adjusted- \(R^2\) as a measure of model quality
Be able to locate information on the significance of individual predictors and overall model fit in R lm model output
\[ y_i = \beta_0 + (\beta_1 \cdot x_{1i}) + (\beta_2 \cdot x_{2i}) + (\beta_j \cdot x_{ji}) + \epsilon_i \]
\[ \text{Score}_i = \beta_0 + (\beta_1 \cdot \text{Hours}_{i}) + (\beta_2 \cdot \text{Motivation}_{i}) + \epsilon_i \]
R:So far we have estimated values for the key parameters of our model ( \(\beta\)s )
Evaluating a model will consist of:
Evaluating the individual coefficients
Evaluating the overall model quality
Evaluating the model assumptions
Important: Before accepting a set of results, all three of these aspects of evaluation must be considered
Is our model informative about the association between X and Y?
Is study time a useful predictor of test score?
This is kind of vague, what does it mean to be a “useful predictor”?
We need to turn this into a testable statistical hypothesis
Steps in hypothesis testing:
Research questions are statements of what we intend to study
A good question defines:
For example:
Does increased study time improve test scores in school-age children?
Statistical hypotheses are testable mathematical statements
In typical testing in Psychology, we define a null hypothesis (\(H_0\)) and an alternative hypothesis (\(H_1\)).
Remember, we can only ever test the null hypothesis
We select a significance level, \(\alpha\) (typically .05)
Then we calculate the \(p\)-value associated with our test statistic
If the associated \(p\) is smaller than \(\alpha\), then we reject the null
If it is larger, then we fail to reject the null
In this example data, there’s no pattern in how values of \(x\) are associated with values of \(y\).
We can’t make any guesses about \(y\) based on the value of \(x\).
In this scenario, the slope \(\beta_1\) of the line associating \(x\) and \(y\) is 0.
Why is the slope 0?
Our null hypothesis is that there is no association between \(x\) and \(y\). Formally:
\[H_0: \beta_1 = 0\] \[H_1: \beta_1 \neq 0\]
\[t = \frac{\beta}{SE(\beta)} = \frac{\text{estimated coefficient}}{\text{standard error estimated for that coefficient}}\]
Call:
lm(formula = score ~ hours + motivation, data = test_study2)
Residuals:
Min 1Q Median 3Q Max
-12.9548 -2.8042 -0.2847 2.9344 13.8240
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 6.86679 0.65473 10.488 <2e-16 ***
hours 1.37570 0.07989 17.220 <2e-16 ***
motivation 0.91634 0.38376 2.388 0.0182 *
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Residual standard error: 4.386 on 147 degrees of freedom
Multiple R-squared: 0.6696, Adjusted R-squared: 0.6651
F-statistic: 148.9 on 2 and 147 DF, p-value: < 2.2e-16
For motivation, which we’ll represent as \(\beta_2\):
\[ t = \frac{\beta_2}{SE(\beta_2)} = \frac{0.9163}{0.3838} = 2.388 \]
We have an observed \(t\)-statistic, 2.388. Now we want to find out how likely it is to observe a \(t\)-value this extreme or more extreme, assuming that \(H_0\) is true.
We know the distribution that \(t\)-values follow when \(H_0\) is true: a \(t\)-distribution.
\(t\)-distributions can vary, depending on their degrees of freedom (df):
hours and motivation)So: \(150 - 2 - 1 = 147\) degrees of freedom.
As with all tests we need to set our \(\alpha\): our probability of incorrectly rejecting the null hypothesis, even when it’s true
We’ll use this \(\alpha\) value, together with the null distribution from the last slide, to find the critical values that we’ll compare our observed \(t\)-value to
The shaded red area, times 2 because it’s a two-tailed test, gives us a \(p\)-value of .018.
\(p = .018\) is less than \(\alpha = 0.05\), so we can reject the null hypothesis that the association between motivation and score is equal to zero.
When we measure an outcome variable \(y\), there will be variability in that outcome.
We quantify how good our model is based on what proportion of the variability in our data it can explain.
To work out this proportion, we need to quantify three different kinds of variability:
| Type of variability in \(y\) | How we quantify it |
|---|---|
| Variability explained by model | \(SS_{Model}\) “Model sum of squares” |
| Variability not explained by model | \(SS_{Residual}\) “Residual sum of squares” |
| Total variability | \(SS_{Model} + SS_{Residual} = SS_{Total}\) “Total sum of squares” |
The proportion of variability that the model explains is then computed as
\[\frac{SS_{Model}}{SS_{Total}} \quad \text{or} \quad 1 - \frac{SS_{Residual}}{SS_{Total}}\]
We call this proportion the coefficient of determination and represent it as \(R^2\).
\[ \begin{align} R^2 &= \frac{SS_{Model}}{SS_{Total}}\\[1em] &= 1 - \frac{SS_{Residual}}{SS_{Total}} \\ \end{align} \]
Let’s see how to calculate each of these three sums of squares.
\[SS_{Total} = \sum_{i=1}^{n}(y_i - \bar{y})^2\]
In words: For every observation, find the difference between the observed value and the mean of all observed \(y\) values, square that difference, and add up all the squares.
\[SS_{Residual} = \sum_{i=1}^{n}(y_i - \hat{y}_i)^2\]
In words: For every observation, find the difference between that observation and the value predicted by the model, square that difference, and add up all the squares.
\[SS_{Model} = \sum_{i=1}^{n}(\hat{y}_i - \bar{y})^2\]
In words: For every observation, find the difference between the predicted value for that observation and the mean of all observed \(y\) values, square that difference, and add them all up.
In the current example, these values are:
\[ R^2 = \frac{SS_{Model}}{SS_{Total}} \]
\[ \begin{align} R^2 &= \frac{5729.23}{8556.06} \\[.5em] &= 0.6695 \\ \end{align} \]
\(R^2\) = 0.6695 means that 66.95% of the variation in test scores is accounted for by hours of revision and student motivation
lm()’s Multiple R-squared!
Call:
lm(formula = score ~ hours + motivation, data = test_study2)
Residuals:
Min 1Q Median 3Q Max
-12.9548 -2.8042 -0.2847 2.9344 13.8240
Coefficients:
Estimate Std. Error t value Pr(>|t|)
(Intercept) 6.86679 0.65473 10.488 <2e-16 ***
hours 1.37570 0.07989 17.220 <2e-16 ***
motivation 0.91634 0.38376 2.388 0.0182 *
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Residual standard error: 4.386 on 147 degrees of freedom
Multiple R-squared: 0.6696, Adjusted R-squared: 0.6651
F-statistic: 148.9 on 2 and 147 DF, p-value: < 2.2e-16
When there are two or more predictors, \(R^2\) tends to be an inflated estimate of the corresponding population value
Due to random sampling fluctuation, even when \(R^2 = 0\) in the population, it’s value in the sample may \(\neq 0\)
In smaller samples, the fluctuations from zero will be larger on average
With more predictors, there are more opportunities to add to the positive fluctuation
We therefore compute an adjusted \(R^2\)
\[ \hat R^2 = 1 - (1 - R^2)\frac{n-1}{n-k-1} \]
Residual standard error: 4.386 on 147 degrees of freedom
Multiple R-squared: 0.6696, Adjusted R-squared: 0.6651
F-statistic: 148.9 on 2 and 147 DF, p-value: < 2.2e-16
Based on adjusted R-squared, hours studying and student motivation explain 66.5% of the variance in test scores
As the sample size is large and the number of predictors small, unadjusted (0.67) and adjusted R-squared (0.665) are similar
If there were even more predictors, then adjusted R-squared would be a much more accurate representation of proportion variance explained
We have an inferential test, based on a \(t\)-distribution, for the significance of \(\beta\)
We are more likely to find a statistically significant effect when residuals are small and we have a large sample
We can assess the degree to which our model explains variance in the outcome based on \(R^2\)
When we have multiple predictors, we should use the adjusted \(R^2\) to get a more conservative estimate
Tasks:
Work on exercises in labs
Complete the weekly quiz
Get support:
Consult the flash cards
Ask questions anonymously on Piazza
We really like seeing you in office hours!
\[ SE(\beta_j) = \sqrt{\frac{ SS_{Residual}/(n-k-1)}{\sum(x_{ij} - \bar{x_{j}})^2(1-R_{xj}^2)}} \]
\[ SE(\beta_j) = \sqrt{\frac{ SS_{Residual}/(n-k-1)}{\sum(x_{ij} - \bar{x_{j}})^2(1-R_{xj}^2)}} \]
We want our \(SE\) to be smaller - this means our estimate is precise
Examining the above formula we can see that: