library(tidyverse)
library(broom)01: Simple linear regression
This week, you’ll use R to fit your first linear model, discover a few ways to extract information from it, and use it to generate predictions.
- Open RStudio.
- Create a new R Markdown (.Rmd) file for this week’s exercises.
- Save it somewhere you can find it again.
- Give it a clear name (for example,
dapr2_lab01.Rmd). - In the first code chunk, use the
library()function to load the packages you’ll need this week:tidyversebroom
Here’s how to do that:
If you don’t have these packages installed, then in the RStudio console (bottom left of the RStudio window), run the following code:
install.packages(c('tidyverse', 'broom'))Research question (RQ): Does the number of weekly social interactions influence people’s wellbeing?
Data dictionary:
| variable | description |
|---|---|
| age | Age in years of respondent |
| outdoor_time | Self report estimated number of hours per week spent outdoors |
| social_int | Self report estimated number of social interactions per week (both online and in-person) |
| routine | Binary 1=Yes/0=No response to the question 'Do you follow a daily routine throughout the week?' |
| wellbeing | Warwick-Edinburgh Mental Wellbeing Scale (WEMWBS), a self-report measure of mental health and wellbeing. The scale is scored by summing responses to each item, with items answered on a 1 to 5 Likert scale. The minimum scale score is 14 and the maximum is 70 |
| location | Location of primary residence (City, Suburb, Rural) |
| steps_k | Average weekly number of steps in thousands (as given by activity tracker if available) |
From the Edinburgh & Lothians, 100 city/suburb residences and 100 rural residences were chosen at random and contacted to participate in the study. The Warwick-Edinburgh Mental Wellbeing Scale (WEMWBS) was used to measure mental health and wellbeing.
Participants filled out a questionnaire including items concerning: estimated average number of hours spent outdoors each week, estimated average number of social interactions each week (whether on-line or in-person), whether a daily routine is followed (yes/no). For those respondents who had an activity tracker app or smart watch, they were asked to provide their average weekly number of steps.
Read in and explore data
Use the tidyverse function read_csv() to read in the data from https://uoepsy.github.io/data/wellbeing_rural.csv.
Store the data in a variable named mwdata (“mw” stands for “mental wellbeing”).
Here’s the basic usage of read_csv() and how you can assign its output to a variable named mwdata.
Replace the ... with the actual location of the data.
mwdata <- read_csv("...")
Our research question (RQ) is about how the number of social interactions per week might be associated with people’s wellbeing.
The relevant variables for this question are social_int and wellbeing.
Use ggplot to make a scatterplot that displays these two variables in relation to one another. Show social_int on the x axis (because it’s the independent variable) and wellbeing on the y axis (because it’s the dependent variable).
Replace the ... with the appropriate variable or column names.
mwdata |>
ggplot(aes(x = ..., y = ...)) +
geom_point()
Set up and fit linear model
Write the mathematical model formulation for a simple linear model with wellbeing as the outcome variable (i.e., dependent variable) and social_int as the one predictor variable (i.e., independent variable).
Here’s how to write a model formulation using the mathematical notation in R markdown documents:
- In the markdown part of your document (that is, not in a code chunk), create the section that will use math mode: Start by writing two dollar signs, a space, and two more dollar signs:
$$ $$ - In the middle of those four dollar signs, you can write the mathematical notation.
Here’s an example that creates the basic formulation for a simple linear model with outcome y and predictor x:
$$ y = \beta_0 + (\beta_1 \cdot x) + \epsilon $$
- To write variable names as regular text inside the
$$math mode notation, you can write\text{wellbeing}and\text{social\_int}. - What’s up with that backslash in
social_int? Sometimes the underscore can cause problems in math mode (you can also see it creating a subscript after \(\beta\)), but not if we writesocial\_intinstead, placing a backslash in front of the underscore.
🗂️ See Simple regression > Model specification flash card.
Use the function lm() to fit a linear model which predicts wellbeing as a function of social_int. Name the result mdl.
🗂️ See Simple regression > Model building flash card.
R offers many different ways to obtain the coefficient estimates from a fitted model. Run each of the following lines of code and see what each one does. Which one gives you the most information?
mdl
mdl$coefficients
coef(mdl)
coefficients(mdl)
summary(mdl)
Take the intercept and slope coefficients estimated in mdl, round them both to two decimal places, and substitute them appropriately into the mathematical model expression you wrote in Q3.
Bonus R challenge: use coef() to give you the coefficients, and use round() to round them to two decimal places.
🗂️ See Simple regression > Model building flash card.
In the context of this data and RQ, write one sentence interpreting what the intercept means and one sentence interpreting what the slope over social_int means.
🗂️ See Simple regression > Interpreting results flash card.
Use linear model to make predictions
The third row of mwdata represents somebody with 11 social interactions per week:
mwdata[3,]# A tibble: 1 x 7
age outdoor_time social_int routine wellbeing location steps_k
<dbl> <dbl> <dbl> <dbl> <dbl> <chr> <dbl>
1 25 19 11 1 35 rural 49.8
Calculate the wellbeing score that the model would estimate for this person with social_int = 11.
To do this:
- Take the mathematical model expression you wrote for Q6,
- remove the \(+ \epsilon\) bit (why? because model predictions correspond to the exact line the model has estimated, so we don’t want to include the error/residual term),
- substitute the value 11 for the variable \(\text{social\_int}\),
- and conduct the multiplication and addition as defined by the mathematical expression. (Remember that you can use the console in RStudio as a calculator.)
🗂️ See Compute model-predicted values > Example flash card.
Now you’ll try out a new function called augment(). It comes from the package broom, and it is very useful for computing predictions from linear models.
Run the following code:
augment(mdl)You should see a data frame with a bunch of columns. For now, we are only interested in the first four:
wellbeing: the observedwellbeingvalues we gave to the modelsocial_int: the observedsocial_intvalues we gave to the model.fitted: the predicted outcome values from the model (“fitted” or model-fitted” is another way to say “predicted”).resid: the residuals, that is, the difference between each predicted value and each actual observed values
Look at the third row of this data frame. Does the value in the .fitted column approximately match the value you calculated above in Q8?
There may be differences in the hundredths place because we rounded our numbers to two decimal places, but augment() is working with all the decimal places. Don’t worry about that—does the rest of the number match?
🗂️ See Compute model-predicted values > For sample data flash card.
The Piazza forum
Finally: we want you to get familiar with the course’s Piazza page.
Access Piazza from Learn > Quick links > Piazza Q&A.
Piazza is a discussion forum where you can anonymously post questions that your coursemates and instructors can see and respond to. Please use Piazza to ask us your questions!
Asking on Piazza is better than asking by email because on Piazza, everybody can benefit from your questions. And if you have a question, you certainly won’t be the only one.
Your final task this week is to get to know Piazza by navigating to Elizabeth’s post called “Lab 01: Something nice from this summer” and replying (anonymously or not, as you prefer) with something nice that YOU experienced this summer.
(Note: we may switch Q&A systems partway through the year, because the uni might be changing from Piazza to something new. We’ll keep you posted!)
Optional: Knit your report to PDF
Note: During the labs, please spend your time on the exercises above, not on rendering your document. This guidance is here in case you have extra time at the end.
Rmarkdown is a useful format because it can be used as a foundation for generating different kinds of documents. For the group report, you’ll be asked to render your Rmd file to PDF. Why not practice now?
First, make sure you have tinytex installed in R. This is what lets you convert things into PDF format.
Run the following code to install tinytex (you’ll only need to run it once).
install.packages("tinytex")
tinytex::install_tinytex()Then ensure that the “yaml” (rhymes with “mammal”) header—the bit at the top of your Rmd document—looks something like this:
---
title: "this is my report title"
author: "B1234506"
date: "07/09/2025"
output: bookdown::pdf_document2
---
Then follow the instructions from the Rmd bootcamp to knit your Rmd document to PDF.
If you encounter errors, refer to the DAPR1 formatting resources / knitting checklist.
If you are having issues knitting directly to PDF, try the following:
- Knit to HTML file
- Open your HTML in a web-browser (e.g. Chrome, Firefox)
- Print to PDF (Ctrl+P, then choose to save to PDF)
- Open file to check formatting
Review Lesson 5 of the rmd bootcamp for a detailed description/worked examples.
Hiding R Code
To not show the code of an R code chunk, and only show the output, write:
```{r, echo=FALSE}
# code goes here
```Hiding R Output:
To show the code of an R code chunk, but hide the output, write:
```{r, results='hide'}
# code goes here
```Hiding R code AND output
To hide both code and output of an R code chunk, write:
```{r, include=FALSE}
# code goes here
```