DAPR2 Lab Exercises
  • Block 1: Intro LM
    • 01: Simple linear regression
    • 02: Multiple regression
    • 03: Significance tests
  • Block 2: Extending LM
  • Block 3: Interactions
  • Block 4: Logistic regression

On this page

  • Recap from last week
  • Test significance of \(\beta_1\)
  • Variance explained
  • Report regression models
  • Bonus conceptual question
  • Optional: Knit your report to PDF

03: Significance tests

This week, you’ll revisit the same model you fit last week, focusing now on manually recreating the significance tests reported in summary(). You’ll also practice writing up the model results using APA-style reporting.

NoteGet set up
  1. Open RStudio.
  2. Create a new R Markdown (.Rmd) file for this week’s exercises.
  3. Save it somewhere you can find it again.
  4. Give it a clear name (for example, dapr2_lab03.Rmd).
  5. In the first code chunk, load the packages you’ll need this week (and install them if you don’t have them already):
    • tidyverse

Research question (RQ): Is there an association between wellbeing and time spent outdoors, when controlling for the effect that social interaction has on wellbeing?

Data dictionary:

variable description
age Age in years of respondent
outdoor_time Self report estimated number of hours per week spent outdoors
social_int Self report estimated number of social interactions per week (both online and in-person)
routine Binary 1=Yes/0=No response to the question 'Do you follow a daily routine throughout the week?'
wellbeing Warwick-Edinburgh Mental Wellbeing Scale (WEMWBS), a self-report measure of mental health and wellbeing. The scale is scored by summing responses to each item, with items answered on a 1 to 5 Likert scale. The minimum scale score is 14 and the maximum is 70
location Location of primary residence (City, Suburb, Rural)
steps_k Average weekly number of steps in thousands (as given by activity tracker if available)
NoteMore detail about this dataset

From the Edinburgh & Lothians, 100 city/suburb residences and 100 rural residences were chosen at random and contacted to participate in the study. The Warwick-Edinburgh Mental Wellbeing Scale (WEMWBS) was used to measure mental health and wellbeing.

Participants filled out a questionnaire including items concerning: estimated average number of hours spent outdoors each week, estimated average number of social interactions each week (whether on-line or in-person), whether a daily routine is followed (yes/no). For those respondents who had an activity tracker app or smart watch, they were asked to provide their average weekly number of steps.

Recap from last week

Question 1

Read in the data from https://uoepsy.github.io/data/wellbeing_rural.csv and store it in a variable named mwdata (“mw” stands for “mental wellbeing”).

Using the function lm(), fit the linear model represented by the mathematical model formulation below, and name the result mdl.

\[ \text{wellbeing} = \beta_0 + (\beta_1 \cdot \text{outdoor\_time}) + (\beta_2 \cdot \text{social\_int}) + \epsilon \]

mwdata <- read_csv("https://uoepsy.github.io/data/wellbeing_rural.csv")
mdl    <- lm(wellbeing ~ outdoor_time + social_int, data = mwdata)

Test significance of \(\beta_1\)

Question 2

Our RQ is interested in the association between hours spent outdoors and wellbeing, holding social interaction constant. Thus the coefficient we’re interested in testing for this RQ is \(\beta_1\). \(\beta_1\) represents the slope over the predictor variable outdoor_time, holding the other predictor variable social_int constant.

We’ll start by defining our null hypothesis H0 and our alternative hypothesis H1 for the significance test for \(\beta_1\).

  • Write a sentence in plain English that states the null hypothesis that we’ll test for \(\beta_1\).
  • Then write another sentence that states the alternative hypothesis for \(\beta_1\).
  • Finally, write the null and alternative hypotheses again, this time using mathematical notation.

🗂️ See Specifying hypotheses flash card.

In plain English:

  • The null hypothesis is that \(\beta_1\) is equal to zero (in other words, that the slope over outdoor_time is zero, or that the line that associates outdoor_time with wellbeing, while holding social_int constant, is flat).
  • The alternative hypothesis is that \(\beta_1\) is not equal to zero.


\[ \begin{align} H_0 &: \beta_1 = 0 \\ H_1 &: \beta_1 \neq 0\\ \end{align} \]

Write:

$$
\begin{align}
H_0 &: \beta_1 = 0 \\
H_1 &: \beta_1 \neq 0\\
\end{align}
$$

Question 3

Run the code summary(mdl) to produce the model summary.

Write down the observed t-statistic for \(\beta_1\).

(Which coefficient maps to \(\beta_1\)? Recall that \(\beta_1\) is the slope over outdoor_time, which is represented in the model summary as the value in the Estimate column for outdoor_time.)

summary(mdl)

Call:
lm(formula = wellbeing ~ outdoor_time + social_int, data = mwdata)

Residuals:
     Min       1Q   Median       3Q      Max 
-15.7611  -3.1308  -0.4213   3.3126  18.8406 

Coefficients:
             Estimate Std. Error t value Pr(>|t|)    
(Intercept)  28.62018    1.48786  19.236  < 2e-16 ***
outdoor_time  0.19909    0.05060   3.935 0.000115 ***
social_int    0.33488    0.08929   3.751 0.000232 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 5.065 on 197 degrees of freedom
Multiple R-squared:  0.1265,    Adjusted R-squared:  0.1176 
F-statistic: 14.26 on 2 and 197 DF,  p-value: 1.644e-06


\(\beta_1\) is the coefficient for outdoor_time, so the observed t-statistic is 3.935.

Question 4

To find out whether this t-statistic is significantly different from zero, we must compare it to a t-distribution with the appropriate number of degrees of freedom. The “appropriate” degrees of freedom are determined based on the data and the model. In this question, you’ll work out what that quantity should be.

The t-distribution will have degrees of freedom equal to \(n - k - 1\), where

  • \(n\) = sample size, i.e., number of observations
  • \(k\) = number of predictors, i.e., number of \(\beta\) coefficients not including the intercept

Figure out the values for \(n\) and \(k\). Use them to calculate how many degrees of freedom this coefficient’s t-distribution should have.

In mwdata, we have one observation per row. So to work out \(n\), we can look at the number of observations in the RStudio environment, or alternatively:

nrow(mwdata)
[1] 200


The model has two predictors, so \(k = 2\).

Thus:

\[ \begin{align} df &= n - k - 1 \\ &= 200 - 2 - 1 \\ &= 197 \end{align} \]

Question 5

Assume \(\alpha = .05\) (the standard in Psychology) and a two-tailed test.

Copy the code below. It uses qt() to find the critical t-values (that is, the t-values which establish the boundaries for statistical significance) on the t-distribution with the appropriate degrees of freedom.

Replace ... with the degrees of freedom you determined in Q4.

Then run the code.

c(
  lower = round(qt(0.025, ...), 3),
  upper = round(qt(0.975, ...), 3)
)

Write down your answer: Does the observed t-value for \(\beta_1\) fall beyond these critical values? Can we reject the H0?

🗂️ See Reconstructing model estimates > t value flash card.

If our observed t-value for \(\beta_1\) is smaller than –1.972 or greater than 1.972, then we conclude that \(\beta_1\) is significantly different from zero.

Our observed t-value is 3.935, which is greater than 1.972. Thus \(\beta_1\) is significantly different from zero, and we can reject the H0.

Variance explained

Question 6

All model summaries present two \(R^2\) values, one unadjusted and one adjusted.

Write down your answer: For this model, which one should we pay attention to? Why?

Then write down your answer: How much variance in wellbeing scores does our model explain?

🗂️ See Assessing model fit > R-squared and adjusted R-squared flash card.

The model summary presents both “Multiple R-squared” and “Adjusted R-squared”.

Multiple \(R^2\) is appropriate for linear models with one single predictor. Adjusted \(R^2\) is appropriate for linear models with more than one predictor. We should therefore pay attention to the adjusted \(R^2\).

Why is \(R^2\) appropriate for LMs with more than one predictor? Because we will always be able to explain the data better if we add in more predictors. This means that multiple \(R^2\) will always increase, the more predictors we add. To take this into account, adjusted \(R^2\) modifies this value depending on how many predictors the model contains.

summary(mdl)

Call:
lm(formula = wellbeing ~ outdoor_time + social_int, data = mwdata)

Residuals:
     Min       1Q   Median       3Q      Max 
-15.7611  -3.1308  -0.4213   3.3126  18.8406 

Coefficients:
             Estimate Std. Error t value Pr(>|t|)    
(Intercept)  28.62018    1.48786  19.236  < 2e-16 ***
outdoor_time  0.19909    0.05060   3.935 0.000115 ***
social_int    0.33488    0.08929   3.751 0.000232 ***
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 5.065 on 197 degrees of freedom
Multiple R-squared:  0.1265,    Adjusted R-squared:  0.1176 
F-statistic: 14.26 on 2 and 197 DF,  p-value: 1.644e-06

Our model expalins 11.76% of the variance in wellbeing scores (based on the adjusted \(R^2\)).

Report regression models

Question 7

Imagine you’re writing a report to describe this analysis.

Write a sentence or two to describe:

  • what kind of analysis you ran and
  • the structure of your model (what’s the outcome? what are the predictors?).

In your description, make sure to mention the units/measurement scale of each variable.

We ran a simple linear model to predict people’s WEMWBS wellbeing score as a function of how many hours they spend outdoors per week as well as how many social interactions they have per week.

Question 8

Write a sentence to explain to yourself what the outdoor_time coefficient in the model summary means.

Write another sentence to state whether you can or cannot reject the H0 you defined above.

🗂️ See Multiple regression > Interpreting results flash card.

Holding the number of social interactions constant, increasing one hour of outdoor time is associated with an increase of 0.20 points on the WEMWBS scale.

The result is statistically significant, so yes, we can reject H0.

(Note: the coefficient interpretation is only complete if it mentions holding the other predictor(s) constant.)

Question 9

The APA has official guidelines on how to report regression coefficient estimates. Use this template structure:

\(b\) = (coefficient estimate), 95% CI [(lower bound), (upper bound)], \(p\) = (p-value)

(Tip: use confint(mdl) to get each coefficient’s 95% CI.)

Here’s an example of how to include this information in a sentence:

The slope was significantly different from zero (\(b\) = 0.54, 95% CI [0.32, 0.76], \(p\) = .041). I therefore reject the H0 that this estimate is equal to zero.

When reporting the p-value, you should report exact p-values to three decimal places, with no zero in front of the decimal, for example: \(p\) = .041 or \(p\) = .222. However, if \(p\) is smaller than .001, just write \(p\) < .001.

Based on these guidelines, take your sentences from Q8 and rewrite them using APA notation. This means:

  • Write down the meaning of the coefficient estimate, without referring to R variable names.
  • Include APA-formatted estimates of the relevant statistical quantities.
  • State whether the coefficient is significantly different from zero and what this means for the H0.

🗂️ See the APA Numbers and Statistics Guide.

confint(mdl) |> round(2)
             2.5 % 97.5 %
(Intercept)  25.69  31.55
outdoor_time  0.10   0.30
social_int    0.16   0.51


According to the model, increasing one hour of outdoor time is associated with an increase of 0.20 points on the WEMWBS scale, while holding the number of social interactions constant. This estimate is significantly different from zero (\(b\) = 0.20, 95% CI [0.10, 0.30], \(p\) < .001), and we can therefore reject the H0.

Bonus conceptual question

Question 10

You’ll notice that in the model summary, all \(\beta\) coefficients (intercept AND slopes) are associated with a t-value and a p-value. The H0 being tested for all of these parameters is that their estimates are equal to zero.

This H0 makes a lot of sense for the slope coefficients: if a slope is equal to zero, then the line is flat. When the line is flat, there’s no association between \(x\) and \(y\).

In this question, we want you to ask yourself, and write down your thinking: To what extent does this H0 also make sense for the intercept? If our intercept estimate is significantly different from zero, does that tell us anything interesting? Are values of zero even possible for the intercept of this model?

If you like thinking visually, you might find it helpful to recreate the model-fitted plot from last week using the following code:

sjPlot::plot_model(
    mdl, 
    type = 'eff', 
    terms = c(
        'outdoor_time', 
        'social_int'
    ),
    show.data = TRUE
)

The H0 for the intercept is that its estimate is equal to zero. In other words, that the average estimated outcome value, when all predictors are equal to zero, is zero.

This is usually not really a plausible or interesting H0, and it’s definitely not plausible for our wellbeing example! We’d never actually expect wellbeing scores to be equal to zero. All the scores we observe are between 22 and 59, and the range of all possible values is 14 to 70. Zero isn’t even a possibility!

So when the significance test for the intercept tells us that our intercept estimate is significantly different from zero, it’s like … thanks I guess? We didn’t really think it was going to be equal to zero anyway.

In short: Comparing the intercept coefficient to zero is not really that informative or interesting when we have a continuous outcome variable that we’d expect to be different from zero anyway. The intercept will almost always be significantly different from zero, and we don’t usually care.

Optional: Knit your report to PDF

Note: During the labs, please spend your time on the exercises above, not on rendering your document. This guidance is here in case you have extra time at the end.

Rmarkdown is a useful format because it can be used as a foundation for generating different kinds of documents. For the group report, you’ll be asked to render your Rmd file to PDF. Why not practice now?

First, make sure you have tinytex installed in R. This is what lets you convert things into PDF format.

Run the following code to install tinytex (you’ll only need to run it once).

install.packages("tinytex")
tinytex::install_tinytex()

Then ensure that the “yaml” (rhymes with “mammal”) header—the bit at the top of your Rmd document—looks something like this:

---
title: "this is my report title"
author: "B1234506"
date: "07/09/2025"
output: bookdown::pdf_document2
---

Then follow the instructions from the Rmd bootcamp to knit your Rmd document to PDF.

If you encounter errors, refer to the DAPR1 formatting resources / knitting checklist.

TipWhat to do if you cannot knit to PDF

If you are having issues knitting directly to PDF, try the following:

  • Knit to HTML file
  • Open your HTML in a web-browser (e.g. Chrome, Firefox)
  • Print to PDF (Ctrl+P, then choose to save to PDF)
  • Open file to check formatting
TipHiding code and/or output from code chunks

Review Lesson 5 of the rmd bootcamp for a detailed description/worked examples.

Hiding R Code

To not show the code of an R code chunk, and only show the output, write:

```{r, echo=FALSE}
# code goes here
```

Hiding R Output:

To show the code of an R code chunk, but hide the output, write:

```{r, results='hide'}
# code goes here
```

Hiding R code AND output

To hide both code and output of an R code chunk, write:

```{r, include=FALSE}
# code goes here
```