| variable | description |
|---|---|
| age | Age in years of respondent |
| outdoor_time | Self report estimated number of hours per week spent outdoors |
| social_int | Self report estimated number of social interactions per week (both online and in-person) |
| routine | Binary 1=Yes/0=No response to the question 'Do you follow a daily routine throughout the week?' |
| wellbeing | Warwick-Edinburgh Mental Wellbeing Scale (WEMWBS), a self-report measure of mental health and wellbeing. The scale is scored by summing responses to each item, with items answered on a 1 to 5 Likert scale. The minimum scale score is 14 and the maximum is 70 |
| location | Location of primary residence (City, Suburb, Rural) |
| steps_k | Average weekly number of steps in thousands (as given by activity tracker if available) |
03: Significance tests
This week, you’ll revisit the same model you fit last week, focusing now on manually recreating the significance tests reported in summary(). You’ll also practice writing up the model results using APA-style reporting.
- Open RStudio.
- Create a new R Markdown (.Rmd) file for this week’s exercises.
- Save it somewhere you can find it again.
- Give it a clear name (for example,
dapr2_lab03.Rmd). - In the first code chunk, load the packages you’ll need this week (and install them if you don’t have them already):
tidyverse
Research question (RQ): Is there an association between wellbeing and time spent outdoors, when controlling for the effect that social interaction has on wellbeing?
Data dictionary:
From the Edinburgh & Lothians, 100 city/suburb residences and 100 rural residences were chosen at random and contacted to participate in the study. The Warwick-Edinburgh Mental Wellbeing Scale (WEMWBS) was used to measure mental health and wellbeing.
Participants filled out a questionnaire including items concerning: estimated average number of hours spent outdoors each week, estimated average number of social interactions each week (whether on-line or in-person), whether a daily routine is followed (yes/no). For those respondents who had an activity tracker app or smart watch, they were asked to provide their average weekly number of steps.
Recap from last week
Read in the data from https://uoepsy.github.io/data/wellbeing_rural.csv and store it in a variable named mwdata (“mw” stands for “mental wellbeing”).
Using the function lm(), fit the linear model represented by the mathematical model formulation below, and name the result mdl.
\[ \text{wellbeing} = \beta_0 + (\beta_1 \cdot \text{outdoor\_time}) + (\beta_2 \cdot \text{social\_int}) + \epsilon \]
Test significance of \(\beta_1\)
Our RQ is interested in the association between hours spent outdoors and wellbeing, holding social interaction constant. Thus the coefficient we’re interested in testing for this RQ is \(\beta_1\). \(\beta_1\) represents the slope over the predictor variable outdoor_time, holding the other predictor variable social_int constant.
We’ll start by defining our null hypothesis H0 and our alternative hypothesis H1 for the significance test for \(\beta_1\).
- Write a sentence in plain English that states the null hypothesis that we’ll test for \(\beta_1\).
- Then write another sentence that states the alternative hypothesis for \(\beta_1\).
- Finally, write the null and alternative hypotheses again, this time using mathematical notation.
🗂️ See Specifying hypotheses flash card.
Run the code summary(mdl) to produce the model summary.
Write down the observed t-statistic for \(\beta_1\).
(Which coefficient maps to \(\beta_1\)? Recall that \(\beta_1\) is the slope over outdoor_time, which is represented in the model summary as the value in the Estimate column for outdoor_time.)
To find out whether this t-statistic is significantly different from zero, we must compare it to a t-distribution with the appropriate number of degrees of freedom. The “appropriate” degrees of freedom are determined based on the data and the model. In this question, you’ll work out what that quantity should be.
The t-distribution will have degrees of freedom equal to \(n - k - 1\), where
- \(n\) = sample size, i.e., number of observations
- \(k\) = number of predictors, i.e., number of \(\beta\) coefficients not including the intercept
Figure out the values for \(n\) and \(k\). Use them to calculate how many degrees of freedom this coefficient’s t-distribution should have.
Assume \(\alpha = .05\) (the standard in Psychology) and a two-tailed test.
Copy the code below. It uses qt() to find the critical t-values (that is, the t-values which establish the boundaries for statistical significance) on the t-distribution with the appropriate degrees of freedom.
Replace ... with the degrees of freedom you determined in Q4.
Then run the code.
c(
lower = round(qt(0.025, ...), 3),
upper = round(qt(0.975, ...), 3)
)Write down your answer: Does the observed t-value for \(\beta_1\) fall beyond these critical values? Can we reject the H0?
🗂️ See Reconstructing model estimates > t value flash card.
Variance explained
All model summaries present two \(R^2\) values, one unadjusted and one adjusted.
Write down your answer: For this model, which one should we pay attention to? Why?
Then write down your answer: How much variance in wellbeing scores does our model explain?
🗂️ See Assessing model fit > R-squared and adjusted R-squared flash card.
Report regression models
Imagine you’re writing a report to describe this analysis.
Write a sentence or two to describe:
- what kind of analysis you ran and
- the structure of your model (what’s the outcome? what are the predictors?).
In your description, make sure to mention the units/measurement scale of each variable.
Write a sentence to explain to yourself what the outdoor_time coefficient in the model summary means.
Write another sentence to state whether you can or cannot reject the H0 you defined above.
🗂️ See Multiple regression > Interpreting results flash card.
The APA has official guidelines on how to report regression coefficient estimates. Use this template structure:
\(b\) = (coefficient estimate), 95% CI [(lower bound), (upper bound)], \(p\) = (p-value)
(Tip: use confint(mdl) to get each coefficient’s 95% CI.)
Here’s an example of how to include this information in a sentence:
The slope was significantly different from zero (\(b\) = 0.54, 95% CI [0.32, 0.76], \(p\) = .041). I therefore reject the H0 that this estimate is equal to zero.
When reporting the p-value, you should report exact p-values to three decimal places, with no zero in front of the decimal, for example: \(p\) = .041 or \(p\) = .222. However, if \(p\) is smaller than .001, just write \(p\) < .001.
Based on these guidelines, take your sentences from Q8 and rewrite them using APA notation. This means:
- Write down the meaning of the coefficient estimate, without referring to R variable names.
- Include APA-formatted estimates of the relevant statistical quantities.
- State whether the coefficient is significantly different from zero and what this means for the H0.
🗂️ See the APA Numbers and Statistics Guide.
Bonus conceptual question
You’ll notice that in the model summary, all \(\beta\) coefficients (intercept AND slopes) are associated with a t-value and a p-value. The H0 being tested for all of these parameters is that their estimates are equal to zero.
This H0 makes a lot of sense for the slope coefficients: if a slope is equal to zero, then the line is flat. When the line is flat, there’s no association between \(x\) and \(y\).
In this question, we want you to ask yourself, and write down your thinking: To what extent does this H0 also make sense for the intercept? If our intercept estimate is significantly different from zero, does that tell us anything interesting? Are values of zero even possible for the intercept of this model?
If you like thinking visually, you might find it helpful to recreate the model-fitted plot from last week using the following code:
sjPlot::plot_model(
mdl,
type = 'eff',
terms = c(
'outdoor_time',
'social_int'
),
show.data = TRUE
)
Optional: Knit your report to PDF
Note: During the labs, please spend your time on the exercises above, not on rendering your document. This guidance is here in case you have extra time at the end.
Rmarkdown is a useful format because it can be used as a foundation for generating different kinds of documents. For the group report, you’ll be asked to render your Rmd file to PDF. Why not practice now?
First, make sure you have tinytex installed in R. This is what lets you convert things into PDF format.
Run the following code to install tinytex (you’ll only need to run it once).
install.packages("tinytex")
tinytex::install_tinytex()Then ensure that the “yaml” (rhymes with “mammal”) header—the bit at the top of your Rmd document—looks something like this:
---
title: "this is my report title"
author: "B1234506"
date: "07/09/2025"
output: bookdown::pdf_document2
---
Then follow the instructions from the Rmd bootcamp to knit your Rmd document to PDF.
If you encounter errors, refer to the DAPR1 formatting resources / knitting checklist.
If you are having issues knitting directly to PDF, try the following:
- Knit to HTML file
- Open your HTML in a web-browser (e.g. Chrome, Firefox)
- Print to PDF (Ctrl+P, then choose to save to PDF)
- Open file to check formatting
Review Lesson 5 of the rmd bootcamp for a detailed description/worked examples.
Hiding R Code
To not show the code of an R code chunk, and only show the output, write:
```{r, echo=FALSE}
# code goes here
```Hiding R Output:
To show the code of an R code chunk, but hide the output, write:
```{r, results='hide'}
# code goes here
```Hiding R code AND output
To hide both code and output of an R code chunk, write:
```{r, include=FALSE}
# code goes here
```