Building maximal models and interpreting LMM estimates
In this lab, you’ll apply the tools you saw in the lectures this week to the same three datasets you got to know in Week 1. You’ll work out the maximal model structure for each one.
And in Question 10 (which is more substantial than the first nine), you’ll fit one of those maximal models and interpret its estimates.
Get set up
Create a new .Rmd file for this week’s exercises.
Save it somewhere you can find it again.
Give it a clear name (for example, dapr3_lab03.Rmd).
In the first code chunk, load the packages you’ll need this week:
RQ: Are people more likely to purchase clothing when they see it displayed on a model, and is this association dependent on item price?
variable
description
purch_rating
Purchase rating (sliding scale 0 to 100, with higher ratings indicating greater perceived likelihood of purchase)
price
Price presented for item (range £5 to £100)
ppt
Participant identifier
condition
Whether items are seen on a model or on a white background
More detail about this dataset
Thirty participants were presented with a set of pictures of items of clothing, and rated each item how likely they were to buy it. Each participant saw 20 items, ranging in price from £5 to £100. 15 participants saw these items worn by a model, while the other 15 saw the items hanging against a white background.
From the Week 1 lab, here’s the key information about clothing:
Outcome: purch_rating
Predictors: price, condition (and their interaction)
Randomly-varying grouping variable: ppt
Question 1
Write the R formula for the fixed effects part of this model (that is, the y ~ x + z part).
(Note: We know this has to be an interaction model because the RQ is asking about how an association between predictor and outcome depends on the value of another predictor.)
Question 2
In the Week 1 lab, you identified the randomly-varying grouping variable(s) in this dataset. For each randomly-varying grouping variable, we know that a maximal model requires a random intercept—an adjustment to the fixed intercept for each level of the grouping variable. Additionally, for each of those variables, a maximal model may contain a random slope over each predictor—an adjustment to the fixed slope for each level of the grouping variable.
Identify the random slopes that this model can contain. You can figure this out by asking: for a given randomly-varying grouping variable, do at least some of its levels appear with more than one distinct value of a given predictor? If yes, then it’s possible to include a random slope over the given predictor by the given grouping variable.
Specifically, for clothing:
Do at least some levels of ppt appear with more than one distinct value of price? Can we include a random slope over price by ppt?
Do at least some levels of ppt appear with more than one distinct value of condition? Can we include a random slope over condition by ppt?
Do at least some levels of ppt appear with more than one distinct value of price?Yes, every ppt sees every price once. In other words, price varies within ppt.
Each value of ppt appears with more than one value of price (that is, each participant sees more than one price), so we can include a random slope over price by ppt.
Do at least some levels of ppt appear with more than one distinct value of condition?No, every ppt sees only one condition or the other, not both. In other words, condition varies between ppts.
Each value of ppt appears with only one value of condition (that is, each participant sees only one condition, not both), so we cannot include a random slope over condition by ppt.
Question 3
Write out the complete R formula for this maximal model (expanding on the fixed-effects-only model you wrote above by adding on the appropriate random effects).
RQ: How is the social status of monkeys associated with their ability to solve problems, while controlling for the difficulty of the problem?
variable
description
status
Social status of monkey (adolescent, subordinate adult, or dominant adult)
difficulty
Problem difficulty ('easy' vs 'difficult')
monkeyID
Monkey name
solved
Whether or not the problem was successfully solved by the monkey
More detail about this dataset
Researchers have given a sample of Rhesus Macaques various problems to solve in order to receive treats. Troops of Macaques have a complex social structure, but adult monkeys tend can be loosely categorised as having either a “dominant” or “subordinate” status. The monkeys in our sample are either adolescent monkeys, subordinate adults, or dominant adults. Each monkey attempted various problems before they got bored/distracted/full of treats. Each problems were classed as either “easy” or “difficult”, and the researchers recorded whether or not the monkey solved each problem.
From the Week 1 lab, here’s the key information about monkeystatus:
Outcome: solved
Predictors: status, difficulty
Randomly-varying grouping variable: monkeyID
Question 4
Write the R formula for the fixed effects part of this model.
(Note: We know this is not an interaction model because the RQ is asking for us to control for a specific variable, not to estimate how one effect depends on a different variable. To control for a specific covariate, we just need to add it on.)
Question 5
We already know what random intercepts the model must have (one per randomly-varying grouping variable).
Because each monkey only has one status, no random slope over status is possible.
Do at least some levels of monkeyID appear with more than one distinct value of difficulty?Yes, nearly every monkeyID sees both levels of difficulty.
xtabs(~ monkeyID + difficulty, data = monkey)
difficulty
monkeyID difficult easy
Aliyya 3 4
Ashley 3 3
Billy 3 4
Brianna 3 2
Catherine 4 5
Celestina 3 2
Cheyenne 2 4
Cinoi 4 6
Courtney 2 5
Daniel 5 3
David 5 6
Delray 5 1
Devonte 2 3
Efren 4 3
Eric 3 5
Erik 2 4
Fuaad 5 1
Gienry 5 6
Iffat 1 7
Jaarallah 5 6
Jashjaun 2 4
Jason 6 4
Jayson 5 5
Jennifer 7 3
Jessica 4 4
Jonathan 2 5
Kaamil 7 4
Maya 2 5
Micah 5 5
Nadheera 0 7
Najiyya 5 6
Nathan 3 2
Nichol 4 6
Peter 1 6
Rachel 5 6
Rebecca 4 5
Richard 3 0
Riley 3 5
Robert 4 3
Saydie 4 4
Sean 3 6
Seunghoo 3 6
Shajee'a 7 3
Shannon 6 4
Sierra 5 1
Sydney 4 4
Taylor 3 3
Van 3 7
Vance 7 1
Young Joo 3 4
Because at least some monkeys have seen both levels of difficulty, we can include a random slope by monkey over difficulty.
(It’s fine that Nadheera and Richard only see one difficulty level each—we have enough data from the other monkeys to still be able to estimate the random slopes.)
Question 6
Write out the complete R formula for this maximal model.
RQ: How is the delivery format of jokes (audio-only vs. audio AND video) associated with differences in humour ratings?
variable
description
ppt
Participant identification number
joke_label
Joke presented
joke_id
Joke identification number
delivery
Experimental manipulation: whether joke was presented in audio-only ('audio') or in audiovideo ('video')
rating
Humour rating chosen on a slider from 0 to 100
More detail about this dataset
These data are simulated to imitate an experiment that investigates the effect of visual non-verbal communication (i.e., gestures, facial expressions) on joke appreciation. Ninety participants took part in the experiment, in which they each rated how funny they found a set of 30 jokes. For each participant, the order of these 30 jokes was randomised for each run of the experiment. For each participant, the set of jokes was randomly split into two halves, with the first half being presented in audio-only, and the second half being presented in audio and video. This meant that each participant saw 15 jokes with video and 15 without, and each joke would be presented with video roughly half of the time.
From the Week 1 lab, here’s the key information about laughs:
Outcome: rating
Predictors: delivery
Randomly-varying grouping variables: ppt and joke_label/joke_id (two forms of the same information; for simplicity we can just use joke_id)
Question 7
Write the R formula for the fixed effects part of this model.
Fit the LMM that you have developed for the laughs dataset, using the maximal model formula you wrote for Q9. (You can just use R’s default treatment coding for delivery, so that audio is the reference level and video is the non-reference level.)
Use the model summary to respond to the following questions.
Fixed effects:
What does the model’s intercept mean?
What does the coefficient deliveryvideo mean?
By-participant random effects:
Within what range will approximately 95% of participant-level intercepts fall?
Within what range will approximately 95% of participant-level slopes over delivery fall?
Could some participants show an effect of delivery that goes in the opposite direction than the fixed effect?
By-joke random effects:
Within what range will approximately 95% of joke-level intercepts fall?
Within what range will approximately 95% of joke-level slopes over delivery fall?
Could some jokes elicit an effect of delivery that goes in the opposite direction than the fixed effect?
The R code yields slightly different values because it takes into account a lot of decimal places, while we only used two, so the difference is due to rounding error. This is not a big deal (the values are approximations anyway), so either version is acceptable.
Could some participants show an effect of delivery that goes in the opposite direction than the fixed effect?
The fixed effect is positive, but some participants are estimated to show a negative slope.
We know this because the lower bound of the range where 95% of participant-level slopes fall is negative (the opposite direction from the fixed effect).
So: yes! It’s estimated that some people find audio-only presentation to be funnier than audio-video presentation.
Within what range will approximately 95% of joke-level intercepts fall?
The fixed intercept is 39.02 points, and the SD of joke-level intercept adjustments is 2.15 points.