Research Methods: From Theory to Practice · 1st Edition

Variations on Experimental Designs

Chapter 9 · Audio study guide with word-level transcript

Thank you for studying with us

The website closes on August 31st and the chapter audio moves to YouTube, free. Everything here is unlocked until then.

If you've supported us already — thank you, genuinely. If this helped you and you'd like to put something toward the last of the running costs, it means a lot.

Support LML
Variations on Experimental Designs
0:00 / 0:00
Up NextChapter 10 · Observation, Case Studies, Archival Research, and Meta-analysis
Report an issue

ⓘ This audio and summary are simplified educational interpretations and are not a substitute for the original text.

Key Takeaways

  • Single-case designs use subjects as their own controls through systematic measurement and intervention manipulation to establish causality.
  • Reversal and ABAB designs verify causality by introducing and withdrawing interventions, though raising ethical concerns about removing beneficial treatments.
  • Multiple baseline designs stagger intervention timing across behaviors or settings to confirm treatment effects rather than time passage.
  • Quasi-experimental designs lack random assignment but allow evaluation in natural settings where random assignment is infeasible or unethical.
  • Factorial designs examine multiple independent variables simultaneously, revealing both main effects and interactions between factors.
  • Higher-order factorial designs with three or more factors improve real-world representation but increase interpretive complexity.
Chapter SummaryWhat this audio overview covers
Experimental designs extend beyond the standard randomized controlled trial format to accommodate practical constraints, ethical considerations, and research questions requiring nuanced approaches. Single-case experimental designs apply rigorous experimental control to individual subjects rather than groups, using the subject as their own control through systematic measurement and intervention manipulation. The reversal design establishes a baseline, introduces an intervention, and withdraws it to verify causality, while the ABAB design repeats this cycle for increased certainty. Multiple baseline designs stagger intervention timing across behaviors or settings to confirm that observed changes result from treatment rather than time passage. Although single-case designs offer convenience and preliminary intervention evaluation, they struggle with generalizability and raise ethical concerns about intentionally removing beneficial treatments. Quasi-experimental designs occupy a middle ground between observational and true experimental research, incorporating some experimental elements while lacking random assignment or complete variable manipulation. These designs excel in natural settings where random assignment is infeasible or unethical, such as evaluating educational interventions in actual classrooms, but sacrifice the ability to draw definitive causal conclusions. Factorial designs address research questions involving multiple independent variables simultaneously, examining both main effects and interactions between factors. A main effect represents the isolated impact of one variable, while interactions reveal how one variable's influence depends on another variable's level, often visualized as nonparallel lines on interaction charts. Factorial designs vary in structure: between-subjects designs assign different participants to factor combinations, within-subjects designs expose the same participants to all combinations, and mixed designs combine both approaches. Higher-order factorial designs investigating three or more factors provide comprehensive real-world representation but introduce interpretive complexity that increases with each additional factor.

Chapter Transcript

Read a transcript excerpt below, or use Study Mode for synchronized audio follow-along.

0:17Imagine you secretly switch your roommate's evening coffee to decaf to cure her insomnia. A classic maneuver. Right. And it works. She is finally sleeping through the night. But to actually prove to science that the coffee was the culprit, you now have to secretly switch her back to caffeinated coffee, watch her toss and turn and just, you know, document her suffering. Which sounds terrible. Exactly. Is that ethical or are you just being a terrible roommate in the name of data?

0:47Today we are diving into the messy, sometimes morally gray and wildly fascinating world of real life experimental design. It is a perfect dilemma to start with. Because when we picture a scientific experiment, we usually imagine a pristine white laboratory. You know, tightly controlled variables, pipettes, test tubes. Yeah, like out of a movie. Right. It feels absolute. It gives us a comforting, clear cause and

1:11effect. But the real world refuses to behave like a sterile lab. Right. You step outside and suddenly you are dealing with unpredictable humans in chaotic environments. The diagnostic landscape just turns into mud. Oh, absolutely. So how do we actually know what we know when we can't control everything? We're taking a deep dive into the variations on experimental designs. Basically, the toolkit researchers use when things get complicated.

1:37It gets very complicated. It really does. And we're going to build this from the ground up, starting with the absolute smallest scale possible. Like, what happens when you don't even have a group of participants? How do you run a valid scientific experiment on just one person? N equals one. Well, we call that a single case experimental design. Now, we have to distinguish this from a simple case study because a case study is where you observe someone, maybe someone with a rare neurological condition, and you get incredibly rich descriptive information.

2:10But it's idiosyncratic. You aren't manipulating anything. You're just watching. Right. A single case experimental design uses strict experimental control. You manipulate an independent variable and you measure a dependent variable. Okay. Let's unpack this. Because if you only have one person, what are you comparing them against? Don't you need a control group who gets a placebo or something? That's the genius of it, actually. The individual subject serves as their own control.

2:33Oh, interesting. Yeah. They participate in all of the conditions across time. So let's look at your roommate's caffeine habit. You notice she drinks coffee every day at 6 p .m. and she can't sleep. To test this scientifically on just her, you use the most basic single case option. It's called the reversal design, also known as the ABA design. ABA. Okay. So we're talking about distinct phases of the experiment.

3:00Exactly. A is your baseline. For five days, you don't change anything. You just measure how long it takes her to fall asleep after drinking her regular caffeinated coffee. Just to get a baseline. Right. You need to establish what normal looks like for her. Then you introduce B, the intervention. You secretly swap her stash for decaf. And you keep it a secret to avoid the placebo effect, right?

3:22Where just knowing she's drinking decaf might make her relax and fall asleep faster. Yes, exactly. So we measure her sleep time during phase B. And if the time it takes her to fall asleep drops dramatically, we've proved it right. Case closed. Oh, not quite. We have some evidence, sure. But what if she just happened to finish a stressful project at work that same week? Oh, I see.

3:41Or what if the weather cooled down and her room is more comfortable? Knowing the baseline first was crucial, but returning to the baseline is just as critical to rule out those coincidences. That brings us to the final A in the ABA design. You withdraw the intervention. Oh, man. You switch her back to the caffeinated coffee. You do. And if we mentally visualize this graph, like the ones in figures 9 .1 through 9 .4 in the text, picture a line tracking how many hours she's awake in bed.

4:11During the first A phase, that line is high. She's staring at the ceiling. Then we hit B, the decaf phase, and the line plunges down. She's sleeping great. But then we hit that second A phase and the line shoots right back up. Precisely. If the insomnia returns when the caffeine returns, you have established a highly probable cause and effect relationship. Some researchers even push it to an ABA design, cycling through it twice.

4:35Just to be extra sure. Exactly. Because if you can turn the insomnia on and off like a light switch multiple times, it is incredibly unlikely to be a coincidence. But this raises an immediate red flag for me. We're purposefully giving someone their insomnia back. We are removing a helpful treatment. That feels problematic. It does. And it raises an important question. It's actually the primary disadvantage of the single case design, the ethics of withdrawal.

5:02In the coffee scenario, your friend returns to a sleepless state. Methodologically, the justification is that the certainty of the answer, knowing definitively what causes the problem, outweighs the temporary discomfort. But scale that up. Imagine the intervention isn't decaf coffee, but, I don't know, an antidepressant medication for severe depression or a behavioral therapy that stops a child from self -harming. Oh, wow. Yeah, you can't just withdraw that to see if they start self -harming again.

5:29That's totally unacceptable. Right. So researchers use an alternative called a multiple baseline design. Instead of withdrawing the treatment, you stagger when the treatment starts across different situations, times, or even different behaviors. Okay, how does that work? Imagine measuring your roommate's sleep, but also measuring her midday anxiety and her morning jitteriness. You introduce decaf for the jitters first, wait a few days, then introduce a wind -down routine for the sleep.

5:55Look at figure 9 .5 for this, you'll see a staggered graph. Right? It looks like a staircase almost. Exactly. If each specific problem improves only after its specific intervention is applied, you prove the treatment works without ever having to take it away. That makes so much sense. It proves the intervention caused the change, not just the passage of time. Exactly. But even with that solved, curing my friend's insomnia doesn't prove caffeine causes insomnia for everyone in the world.

6:23We need groups for that. Which leads us to a new problem. What if we have groups, but we can't randomly assign people to them? Then we enter the realm of the quasi -experimental design. It's partially experimental, it has some strict laboratory controls, but it lacks a full set of controls, usually because you cannot randomly assign participants to the groups or you can't fully manipulate the independent variable.

6:46There is a brilliant study that captures this perfectly. It was done by Deloach, Uddle, and Rosengren in 2004, looking at toddlers and something called scale errors. Oh, I love this one. It's wild. They brought kids, aged 18 to 30 months, into a playroom and they documented these hilarious bizarre behaviors. If you look at the photos in Figure 9 .6, you have these giant toddlers trying to physically slide down a miniature toy slide that's only a few inches long or trying to squeeze their bodies into a tiny dollhouse chair.

7:19The psychology behind it is fascinating. The toddlers' brains haven't yet learned to integrate visual information about the size of an object with the motor planning required to interact with it. So funny. They see a chair, their brain says sit, and they just ignore the fact that the chair is the size of an apple. But how they designed the study is the real star here. They had standardized laboratory rooms, standardized toys, exact playtime limits.

7:42The experimenter sat in the corner and followed a strict, minimal script. Those are intense experimental controls. Yes, but they didn't randomly assign kids to a miniature toy group and a regular toy group. It was a within -subjects design where all the kids just played in the room naturally. And they didn't manipulate an independent variable in the traditional sense. They just exposed the children to the stimuli loosely to observe what happened.

8:07To me, doing a quasi -experiment is kind of like cooking in someone else's kitchen. Oh, I like that. You can control your recipe and you can control your own chopping technique. That is your experimental control. But you absolutely cannot control what kind of stove they have or how their pantry is organized. You have to trade some of your ultimate causal certainty for the flexibility to just work in the environment you're given.

8:30That is a phenomenal analogy. Flexibility is the entire point. Sometimes the real world demands it. Look at Stephen Asher's work from the Inside Research Box in the text. Back in 1967, he wanted to replicate the famous 1940s doll test by Kenneth and Mamie Clark. Right. The original study looked at children's racial preferences by having them choose between a black doll and a white doll. Exactly. Asher wanted to see if those preferences had shifted by 1967 following the civil rights movement.

9:00He also went on to do incredible peer -peering studies to help socially isolated kids. So how did he set up the doll test replication? He used puppets instead of dolls, but the key was that he didn't manipulate the environment. He didn't lock kids in a sterile room and force an artificial variable on them. He went into their natural environment and observed their choices as they naturally occurred.

9:21He went to them. He traded the pristine lab for the messy real world context because that's where the truth of the behavior actually lived. Okay, so quasi -experiments handle the constraints of the environment. We can observe toddlers playing or kids picking puppets. But what about the complexity of the variables themselves? In real life, things aren't just caused by one single factor. Not usually, no. If I have a headache, it might be because I drank too much coffee, but it also depends on how much water I drank and whether I slept.

9:50How do researchers capture that layered reality? They use a factorial design. This is any experimental design that has more than one independent variable. We call these variables factors. And we use this because studying variables in total isolation often gives us a false picture of the world. Factorial designs dramatically increase external validity because they mimic the interconnected nuances of reality. Let's break down the mechanics of this because the terminology can sound a bit intimidating.

10:18You'll hear researchers talk about a two -by -two design. The numbers just tell you how many factors you have and how many levels are inside each factor. So a two -by -two design means you have two factors and each factor has two levels. Let's apply that to a therapy setting like in Table 9 .1. Factor A could be the therapist style that has two levels. Directive therapy where the therapist leads the session or non -directive therapy where the client leads.

10:41And then factor B is the client's openness to therapy. That also has two levels. Low openness or high openness? If we connect this to the bigger picture of study design, researchers have to calculate their cells. A cell is a unique experimental condition. To find out how many groups you need, you multiply the levels. A two -by -two design means two times two, which equals four distinct conditions.

11:05This is where the logistics get crazy. To run this, you need a group of low -openness clients getting directive therapy, low -openness clients getting non -directive therapy, high openness getting directive, and high openness getting non -directive. You have to fill four separate buckets with participants. And notice the difference in those variables. The therapist style is an experimental variable. The researcher can actively manipulate it. But the client's openness is a participant variable.

11:32It's a pre -existing characteristic. You can't randomly assign someone to suddenly be a highly open person. Factorial designs let us mix manipulated variables with the reality of who the participants actually are. Here's where it gets really interesting. There's a study from 2014 by Yun and Vargas called The Almighty Avatars. And it uses a three -by -two factorial design to look at how video games change human behavior.

11:55A three -by -two design. So three times two means six distinct cells. Exactly. Factor one was the avatar the participant played as in a game for just five minutes. It had three levels, Superman, Voldemort, or a neutral circle. Factor two was the type of food the participant was then asked to taste and portion out for an unknown anonymous volunteer to eat. It had two levels, delicious chocolate or scorching hot chili sauce.

12:20What's fascinating here is the psychological mechanism they were testing, priming. Does slipping into the digital skin of a hero or a villain temporarily overwrite your own moral compass? Yes. And the results were wild. The heroes who played as Superman for five minutes were twice as likely to act prosocially and pour out generous portions of chocolate for the stranger. The villains who played as Voldemort were significantly more likely to pour out the hot chili sauce, essentially choosing to inflict pain on a stranger.

12:47By using a factorial design, the researchers didn't just ask does playing Voldemort make you mean. They asked, does playing Voldemort change how you interact with specific types of stimuli, reward versus punishment? The design captured the nuance of the behavior, but once you have this complex data, you have to know how to read it. When you look at the results of a factorial design, you are hunting for two distinct things.

13:12The first is the main effect. This is the overall sweeping effect of a single factor acting alone, totally ignoring the other factor. Like overall, does directive therapy work better than non -directive therapy? You just average the scores and look at the headline. But the headline is really the whole story. The second thing you look for, and arguably the entire reason you run a factorial design in the first place, is the interaction.

13:35An interaction occurs when the effect of one factor is contingent on the levels of another factor. It is the scientific proof of, it depends. Let's visualize this. Look at Figures 9 .7 and 9 .8. If you take your data and map it on a graph, and you draw lines connecting the data points. If those lines are perfectly parallel to each other, like railroad tracks, there is no interaction.

13:56The variables are just doing their own independent things. But if those lines cross, or even if they just aren't parallel, you have an interaction. The effect of one variable is fundamentally shifting based on the presence of the other. Mickalene Shi did a brilliant Between Subjects memory study in 1978 that illustrates this perfectly. She wanted to know what drives our ability to remember things. Is it raw cognitive maturity, which comes with age, or is it domain -specific expertise?

14:23So she built a two -by -two design. Right. Factor one was the memory task, recalling a list of numbers, or recalling the positions of chess pieces on a board. Factor two was the participant's age and expertise profile. She used child chess experts and adult novices. If you visualize the graph of her results, like in Figure 9 .9, it forms a giant X. The lines completely cross. On one side of the graph, looking at the numberless task, the adult novices massively outperformed the children.

14:51Their adult brains had better general working memory for rote tasks. But follow the lines across the X to the chess task. The child experts completely obliterated the adults. The children's deep, specialized knowledge of chess patterns allowed them to chunk the information visually, bypassing their overall age disadvantage. This raises an important question about how we consume information. When you find a massive interaction like that crossing X, it forces you to completely rethink the main effect.

15:20Right. You can't just publish a paper with the main effect headline, Adults Have Better Memories Than Children, because another scientist will point to the chess data and say, well no, it depends on the task and their expertise. The interaction makes this simple explanation obsolete. Exactly. The interaction tells the true story. Now, Chai's study was a between -subjects design. The kids were in one group, the adults were in another.

15:41But what if we want to track the same people across different conditions? We can do it within subjects factorial design. Imagine studying how caffeine dosage, high or low, affects performance on different memory tasks, visual or auditory. Instead of finding four different groups of people, every single participant does all four combinations. I come into the lab, I get high caffeine and do the visual task. Then I get high caffeine and do the auditory task.

16:06Then low caffeine visual, low caffeine auditory. I am repeatedly measured across every cell. And we can actually combine those two philosophies into an intentional hybrid, called a mixed factorial design. We use this when we need the specific benefits of both methods. Take the treatment of coulrophobia, the intense fear of clowns. A very rational fear, if you ask me. Fair enough. Suppose you want to test whether exposure therapy is better than supportive therapy for curing this fear.

16:33You measure the participant's anxiety before they start therapy, the pre -test. And then you measure it again after therapy, the post -test. Okay, so time, pre -test versus post -test is your within -subjects factor. Every single participant is measured at both times. You are tracking their personal evolution. Correct. But the type of therapy, exposure versus supportive, is a between -subjects factor. You randomly assign a person to get only exposure therapy or only supportive therapy.

17:02Why not have them do both therapies? Because of carryover effects. If I give you exposure therapy on Monday and supportive therapy on Tuesday and your fear goes away, which therapy cured you? Or did the combination cure you? By making the therapy a between -subjects factor, we keep the treatments strictly separated. We get the best of both worlds. We track individual personal growth over time, but we avoid contaminating the therapies.

17:26That makes total sense. But researchers don't have to stop at two factors, right? You can keep stacking them into higher -order factorial designs. You can. You can consider three, four, or more factors simultaneously. A researcher studying memory might set up a 2x2x3 design. Okay, doing the math, factor one is steady strategy. Rote rehearsal versus imagery. That's two. Factor two is gender. Male versus female. That's two. Factor three is age.

17:52High school, college, or middle age. That's three. So 2x2x3 equals 12 distinct experimental cells. And that allows you to look for a three -way interaction. You might discover that the imagery strategy is highly effective, but only for high school and college -age women, providing absolutely zero benefit for middle -age women or for men of any age. The nuance is incredible. Okay, I have to push back here.

18:14Visualizing a two -way interaction was just an X on a graph. Visualizing a three -way interaction in my head sounds like a dimensional nightmare. And running an experiment with 12 separate participant buckets sounds like a logistical nightmare. At what point does adding all these factors become so complicated that the data is just useless noise? You've hit on the exact practical limit of these designs. At a level of interpretation, evaluating a four -way or five -way interaction is mathematically possible.

18:39But cognitively, it is almost impossible for a human to interpret meaningfully. The textbook gives a very clear warning about this. Just because you can run a complex four -way factorial design doesn't mean you should. Precisely. The rule of thumb is that your research design must always be driven by your research question, not by a desire to use fancy, overly complex methodology. You start simple. You use a basic design.

19:03You only introduce variations and higher -order factors if the specific complexity of the real -world problem absolutely demands it. So what does this all mean? Look at the journey we've taken today. We started with the microscopic focus of n equals one, the single -case designs where a single person serves as their own control to prove a treatment works. Then we zoomed out to the real -world compromises of quasi -experiments, treating random assignment for the flexibility to study toddlers in playrooms and children's natural choices.

19:31And finally, we explored the multi -layered reality of factorial designs, where we learned that the magic and the truth really happens in the interactions between variables. And if we connect this to the bigger picture of your daily life, there is a crucial takeaway here. We live in a society that is constantly bombarded by information. The next time you're scrolling through your news feed and you see a bold headline confidently stating that X causes Y, a simple standalone main effect, I want you to pause.

20:00Put on your researcher hat. Exactly. Ask yourself, what hidden variable are they ignoring? What is the invisible participant variable or the specific environmental context that creates an interaction? What is the condition that proves it actually depends? Because reality is rarely a simple main effect. That is exactly the kind of critical, protective thinking we need to be doing. Thank you so much for joining us for this deep dive into the messy, complicated, and brilliant reality of experimental variations.

20:28Keep questioning the variables. From the Last Minute Lecture Team, thank you for listening, and we'll see you in the next deep dive. Remember, keep your lab coat handy, but don't be afraid to step out into the mud.