Developing Your Research Protocol
Thank you for studying with us
The website closes on August 31st and the chapter audio moves to YouTube, free. Everything here is unlocked until then.
If you've supported us already — thank you, genuinely. If this helped you and you'd like to put something toward the last of the running costs, it means a lot.
ⓘ This audio and summary are simplified educational interpretations and are not a substitute for the original text.
Key Takeaways
- Random sampling is ideal but practical constraints often require alternatives like stratified or convenience sampling methods.
- Measurement instruments introduce error; researchers should prioritize established, validated scales over creating new measures.
- Prospective power analysis before data collection determines necessary sample size to detect meaningful effect sizes.
- Comprehensive pre-registered analysis plans prevent p-hacking and spurious findings from unplanned statistical exploration.
- WEIRD population oversampling limits generalizability; researchers must ensure diverse, representative samples across studies.
- Research design requires balancing methodological ideals against real constraints of time, budget, and participant access.
Chapter Transcript
Read a transcript excerpt below, or use Study Mode for synchronized audio follow-along.
0:17So in 1962, a researcher named Jacob Cohen looked at the most prestigious abnormal psychology journal in the world. And well, he found something just completely terrifying. Oh, the statistical power survey. Right. Yeah. That one is it's a huge wake up call. It really is. Because when he ran the numbers, he realized that like half of the published studies in that journal were essentially just flipping a coin.
0:41Wow. Yeah. Researchers were failing to detect these real potentially life -saving psychological phenomena almost 50 % of the time. And it was all just because they were skipping this one mathematical step before they even started collecting data. Which is such a tragedy, honestly. It is. So welcome back to another deep dive, everyone. Think of this as, you know, your personal one -on -one tutoring session. Today, we're looking at chapter six of research methods from theory to
1:08practice, and it's titled Developing Your Research Protocol. Our mission here is to figure out how to keep you from making that exact same mistake Cohen found. Because it's a trap, right? Like people assume the hardest part of science is coming up with brilliant world -changing theory. Right. The eureka moment. Exactly. But a theory is really just a sketch. The actual heavy lifting is translating that abstract sketch into a rigid, functional blueprint.
1:34And chapter six provides this brilliant flow chart for it. Okay. Let's unpack this blueprint because it shows these four highly interrelated steps for protocol design. You basically have to juggle all at the same time. Yeah. The four pillars. You've got obtaining a sample, choosing your measures, conducting a power analysis, and formulating an analysis plan. And the textbook introduces us to Dr. Amy Bonert to kind of show what this juggling act looks like in the real world.
2:01Like she started her career as an undergrad at the University of Michigan, literally driving an hour across the state going into people's living rooms to coordinate a study on the friendships of kids with traumatic brain injuries. Right. Which is incredibly hands -on. Totally. And today she evaluates Girls in the Game, which is this massive Chicago nonprofit helping at -risk youth. But the point is whether she's driving to a single participant's house or analyzing data for thousands of kids, she is relying on those exact same four pillars.
2:35Absolutely. And the first of those pillars is really the who of your study. So obtaining your sample. And we have to clearly separate the concept of a population from a sample. The textbook uses a really great candy analogy for this. Oh, M &Ms. I love that one. Yeah. So if you want to know about M &Ms, the population is literally every single M &M ever manufactured in the world.
2:55Which you obviously can't test. You can't eat all of them. Right. As much as we might want to. So instead you buy one single snack pack. That snack pack is your sample. And the entire validity of your research relies on the assumption that this tiny little snack pack represents the millions of candies still sitting in the factory. But I mean, that raises an immediate massive question for anyone designing a study.
3:17Does your sample always have to be a perfect, perfectly diverse miniature replica of the broader population? Like if I'm studying human psychology, do I need a snapshot of literally all humanity? Well, it entirely depends on the underlying mechanics of what you're actually trying to measure. OK, how so? So if you're setting a low level universal neurological process, representativeness might not actually be the priority. Take the Stroop effect, for instance.
3:45That is right. The color word thing. Exactly. That's the phenomenon where your brain's reaction time slows down. If you're asked to name the color of the ink, a word is printed in, but the word itself says a different color. So like the word RAD printed in blue ink. Oh man, that one always messes with my head. Same. But because human processing pathways are hardwired in very, very similar ways across the species, a sample of 20 random people off the street might give you perfectly valid data for that specific test.
4:15But assuming human experience is universally hardwired is like incredibly dangerous, right? Oh, absolutely. Because the textbook uses figure 6 .1 to totally shatter that assumption using the Miller -Lyer illusion. Yes, the arrow lines. Right. So imagine you're looking at two horizontal lines. One line has arrows on the ends pointing inward, like a double -sided dart. The other has arrows pointing outward, kind of like the corners of a room.
4:42And to most of us, the line with the inward arrows looks substantially longer. Exactly. Even though if you take a ruler to them, both lines are exactly the same length. And for a long time, psychologists just assumed that this optical illusion fooled every human brain on earth. But cross -cultural studies revealed something completely different. When researchers presented this illusion to like 15 different communities worldwide, European and American samples fell for it.
5:07But many non -Western populations just, they didn't. They saw the lines as exactly equal. I remember reading about the mechanics of this. It's called the Carpentered World Hypothesis. The theory is that people raised in environments with lots of straight edges, you know, rectangular buildings, rigid corners. They literally trained their brains to interpret two -dimensional angles as three -dimensional depth cues. Right. So an outward pointing arrow is processed as the inside corner of a room.
5:35Exactly. Which seems further away. So the brain just sort of scales it up automatically. But if you grow up in a natural environment without those rigid architectural lines, your brain never develops that specific shortcut. And that completely upends what we call the college sophomore problem that plagues modern psychology. The weird problem. Yes, the weird samples. For decades, researchers have relied on the easiest available subjects, which usually means 18 to 24 -year -olds taking intro psych classes just to fulfill a credit requirement.
6:04Right. This creates what the scientific community calls weird samples. That stands for Western, educated, industrialized, rich, and democratic. And so by using these weird samples, researchers publish papers claiming to reveal like fundamental truths about human nature when functionally they're just describing the neuroses of an incredibly specific outlier subset of humanity. Precisely. As a researcher, you have to relentlessly interrogate how your specific sample might bottleneck your ability to generalize your findings to the rest of the world.
6:39And part of that interrogation is also how you label the people you do manage to recruit, especially with nonstandard populations. What's fascinating here is how much the labels themselves have evolved. The text details the history of how we talk about Down syndrome. So in 1866, an English doctor named John Langdon Down first formally categorized the genetic disorder. With the extra 21st chromosome. Right. But in his published medical paper, he used terms like idiots and imbeciles.
7:07And over the decades, those labels shifted to terms like mongoloid or mentally retarded. Which are completely jarring to read today. It feels so offensive. It does. Today, we use people -first language. The scientific standard is to emphasize the individual's humanity before the condition. So people with Down syndrome. That makes sense. But this isn't just about evolving societal politeness. It has a direct methodological impact. If you design a study and use outdated or offensive labels in your recruitment materials, your participants will immediately recognize that you don't understand them.
7:41Right. They'll be like, who is this person? Exactly. They won't trust you. They might drop out of the study or they might mask their true behavior, which effectively destroys the validity of your data. So how do we physically gather these people? I know the gold standard is random sampling, where every single person in your target population has an identical mathematical probability of being selected. Right. Like drawing cards from a perfectly shuffled deck.
8:07Yeah. But to keep the odds truly equal, if you draw the king of hearts, you have to put it back in the deck and shuffle again before your next draw. The textbook calls that sampling with replacement. But true random sampling is often a mathematical fiction. Because you can't put the entire population of Chicago in a giant hat. Exactly. So we use alternative mechanisms like stratified random sampling, which is where you pre -divide the population into demographic groups like age or income brackets, and then sample within those groups to guarantee proportional representation.
8:39Or if you need to ensure a minority group isn't statistically drowned out, you might use oversampling. Right. Deliberately recruiting a disproportionately high number of people from that group, just to ensure you have enough mathematical weight to draw a valid conclusion about them. Yep. And when mathematical probability isn't an option at all, researchers turn to non -probability samples. The most common one is convenient sampling, which is exactly what it sounds like just asking whoever is nearby to take your survey.
9:09Which happens a lot. And then there's snowball sampling, which I thought was a brilliant mechanism for reaching hidden populations. Oh, it really is. Because if you want to study, say, undocumented immigrants, they aren't going to respond to a flyer on a bulletin board. You get to find one participant you can build trust with, and you ask them to recruit two friends, and those friends recruit two friends.
9:28The sample just grows like a snowball rolling down a hill. But the mechanical flaw there is self -selection bias. You kind of have to assume the kind of person who volunteers for a scientific study is fundamentally different from someone who just ignores your email. Right. And today, a lot of this recruitment happens on platforms like Amazon's Mechanical Turk or MTurk, where researchers just pay people small fees to complete tasks online.
9:53Which introduces behavioral economics into your methodology. Yes. You might assume that paying people more money guarantees better data. But behavioral research actually shows that paying participants to do a task they already find inherently interesting can destroy their intrinsic motivation. It's such a delicate balance. Offer too little, and they rush through it. Offer too much, and you actually alter their psychological state. Okay, but let's say you pull off that perfect sample.
10:18You navigate the weird bias. You use first language, and you gather the exact right group. That sample is completely worthless if the tool you use to extract data from them is broken. Exactly. Which brings us to step two. Choosing your measures. This is the how of your protocol. And we have to confront the reality of measurement error here. Your instrument is never going to perfectly capture reality.
10:44There's always going to be a gap between the true phenomenon and the tool spits out. Right. Like the textbook's example on the dual task effect in older adults. Yes. Neurologically, older adults will unconsciously slow their walking speed when they're asked to hold a conversation at the same time. But imagine trying to measure that walking speed using a cheap handheld stopwatch. The delay between your eyes seeing them cross the finish line and your thumb physically clicking the button introduces massive measurement error.
11:12Right. You might completely fail to detect them slowing down. Not because the effect isn't real, but just because your thumbs are too slow. To minimize that error, you really have to understand the mathematical language of your tools. Table 6 .1 breaks down the four scales of measurement in ascending order of mathematical complexity. Okay, lay them on me. The foundation is the nominal scale. These are strictly unordered categories.
11:35Think of political affiliation, Democrat, Republican, Libertarian. You might code them in your spreadsheet as one, two, and three, but those numbers have zero mathematical value. A two is not twice as much as a one. Got it. And the next step up is ordinal, which introduces rank order. So if a preschool teacher ranks Johnny as the most aggressive kid in class and Samantha as the second most aggressive, we know the hierarchy, right?
11:59We know who's worse, yeah. But we have no idea about the psychological distance between them. Johnny might be throwing desks through windows while Samantha occasionally refuses to share a crayon. The rank tells us the order, but the magnitude of the gap remains totally invisible. Which is exactly why researchers push for the third level, interval scales. Suddenly, the mathematical distances between the points are exactly equal. The gap between an IQ of 100 and 110 represents the exact same amount of cognitive difference as a gap between 90 and 100.
12:31But interval scales still lack a true zero, right? Like, a temperature of zero degrees Fahrenheit doesn't mean the physical absence of temperature. It's just another arbitrary point on the line. Exactly. To get that true zero, you need the top tier. Ratio scales. Things like reaction time, physical length, or annual income. Zero dollars earned literally means an absolute absence of dollars. Which is structurally different. Very much so.
12:56Think about it mechanically. You cannot calculate a true mathematical average of Republican and Democrat. You can only count frequencies. But if you have an interval of a ratio scale, suddenly you can calculate averages, variances, standard deviations. And that opens the door to parametric statistical tests. Yes. Parametric tests use those averages to cut through the noise and detect incredibly subtle patterns in your data. If you are stuck with nominal or ordinal data, you have to use non -parametric tests, which just rely on counting and ranking, and they're inherently less sensitive.
13:28Okay, this totally explains the massive ongoing war over the Likert scale. Yeah, those surveys rate your pain from one to five. Strongly disagree to strongly agree. Oh yes, the great debate. Because psychologists constantly fight over whether the psychological distance between strongly disagree and disagree is the exact same mathematical distance as the gap between agree and strongly agree. If the gaps aren't equal, it's just an ordinal scale.
13:56But researchers so want to use those sensitive parametric tests that they often just treat Likert data as an interval scale anyway and just, you know, hope for the best. Here's where it gets really interesting though. Even if your math checks out, you're constantly trading off between reliability and validity. Yes, consistency versus accuracy. Right. Reliability is consistency. If I step on a scale 10 times, does it give me the exact same weight every time?
14:22Validity is accuracy. Is the scale actually measuring my weight or is it broken and measuring my height and just giving me a useless number? The textbook's example of measuring human creativity captures this trade -off perfectly. If you want to measure someone's creativity, you could just count the sheer number of ideas they generate in five minutes. Which is highly reliable. Any two researchers can agree on the exact count.
14:44But is it valid? Does yelling out 50 terrible useless ideas actually make someone creative? Probably not. A much more valid measure would be evaluating the novelty and usefulness of those ideas. But novelty is subjective. It's incredibly hard to measure reliably across different researchers. So you are almost always sacrificing a little bit of reliability to get closer to true validity. Okay. So assume your sample is pristine and your measures are both reliable and valid.
15:11That brings us to step three and the tragic 1962 Cohen survey we show with. You have to conduct a power analysis. Yes, the juice. You need to know if your study has enough mathematical magnification to actually see the result. Let's use a visual analogy for statistical power here. Imagine you're steering at a drop of pond water. Statistical power is the physical strength of your microscope lens. If you're looking for a giant amoeba, which represents a massive effect size, a weak, cheap lens is perfectly fine.
15:41You'll spot it easily. But if you're looking for a tiny, subtle virus, a very small effect size, you need an incredibly powerful lens. And in research methodology, thickening that lens usually means recruiting a massive sample of participants. Statistical power is the probability that your study will detect an effect, assuming that the effect actually exists in reality. If your lens is weak, the virus is still swimming around in the water.
16:04You just physically cannot see it. And psychologists generally aim for a power of 0 .8, meaning an 80 % chance of detecting a true effect. To achieve this, you should ideally conduct a perspective power analysis before you begin. You estimate how big the phenomenon is likely to be the effect size, and the math tells you exactly how many participants you need to recruit to build a strong enough lens.
16:27But if researchers know how to build the right lens, why was Cohen finding that half the studies in the 1960s were effectively blind? Because building a massive lens is expensive and time consuming. Recruiting 500 participants takes years. So researchers often evaded the problem by building a weak lens, testing a small sample, but throwing dozens of different hypotheses at the wall. Oh, wow. Yeah. Imagine throwing a handful of darts at a dart board while blindfolded.
16:53Most will miss, but by pure mathematical probability, one dart will eventually hit a bull's eye. They would find a statistically significant result on one random variable out of 30, publish the paper, and just ignore the fact that the entire study was drastically underpowered. So what does this all mean for the real world? The text asks you to imagine testing a brand new, highly effective psychotherapeutic intervention for panic disorder.
17:19You run the study, but because you skipped the perspective power analysis, your sample size is far too small. You don't have enough power. You look through the weak microscope and you just don't see the cure. The math comes back negative. You mistakenly conclude the therapy failed. The paper is shelved, clinical therapists never learn about the technique, and thousands of patients are deprived of a treatment that genuinely works.
17:42And it's all because a researcher tried to save time and skipped a math equation. It's chilling. And the temptation to avoid that kind of failure leads exactly into the absolute necessity of step four, formulate an analysis plan. Right, because the moment you look at your collected data, the psychological pressure to find a successful result is overwhelming. You must decide exactly how you will analyze your data before you ever start collecting it.
18:07Because if you don't bolt yourself to a pre -plan, you leave the door wide open for p -hacking. Let's look at the mechanics of that term. In psychology, the standard threshold for getting a paper published is achieving a p -value of less than 0 .05. Which means there is less than a 5 % probability that your results are just random statistical noise. Exactly. So p -hacking is this shady practice of deliberately manipulating your data set after the fact to force your results under that magical 0 .05 barrier.
18:39You might quietly drop a few outlier participants who didn't react the way you wanted or test unexpected variables until you find something, anything that looks significant. It creates a false positive. You are essentially painting the bullseye around the dart after you already threw it. Oh wait, I want to push back on this a little bit because the textbook explores this exact tension. If I'm exploring my data and I accidentally stumble onto a cure for cancer that wasn't in my original blueprint, am I supposed to just pretend I didn't see it?
19:05Isn't serendipitous data exploration how some of the greatest scientific discoveries are actually made? It absolutely is. And as Simmons, Nelson, and Simonson laid out in a landmark 2011 paper, the core issue isn't exploration. Okay. If we connect this to the bigger picture, the dividing line between noble data exploration and shady data mining comes down to one absolute mechanism. Transparency. Transparency. You are allowed to explore, but if you deviate from your pre -plan, you must fully disclose every single alternative analysis you ran.
19:40You have to show the scientific community the discarded darts. Right. So if you transparently report your entire process, other scientists can look at your math and judge for themselves whether you've stumbled onto a genuine discovery or if your results are just the illusion of arbitrary analytic decisions. Exactly. Which brings us to the final reality check of the chapter. Step five. The art of juggling choices. We have outlined the perfect blueprint here, but a blueprint rarely survives contact with the construction site.
20:06Designing a protocol is fundamentally an exercise in compromise. You're constantly hemmed in by practical constraints, like participant constraints. If you can't get enough older adults to commute to your university lab, you might have to compromise your controlled environment and literally move your equipment to a noisy shopping mall to intercept them. Or you face time constraints. A college senior writing a thesis has one semester. They physically cannot conduct a 10 -year longitudinal study.
20:35Or if your phenomenon is seasonal affective disorder, your entire data collection window slams shut the moment spring arrives. Then there are massive financial constraints. Do you rely on affordable MTurk workers? Or do you have the budget to fly in highly specialized participants? Even the questionnaires themselves can be a barrier. Many established highly reliable psychological measures are strictly copyrighted. Which is wild, right? Requiring researchers to pay massive fees just to print the forms.
21:02It's crazy. And finally, equipment constraints. Every cognitive psychologist would love to use an fMRI machine to scan brain activity in real time. But at $400 an hour, that's simply not obtainable for most projects. You have to pit the optimal against the obtainable. Yeah, a clever low tech behavioral observation might ultimately yield more valid data than a poorly executed fMRI scan anyway. There is no such thing as a flawless study.
21:29The mastery of research methods is learning how to balance these constraints while fearfully protecting the data. So let's look at the finished blueprint. First, obtain your sample identifying who you are studying and whether they truly represent the target population. Second, choose your measures selecting the proper scale to minimize error and balance reliability with validity. Third, conduct a power analysis ensuring your statistical microscope is strong enough to actually see the effect.
21:55And fourth, formulate an analysis plan setting the mathematical rules in stone to keep But I'll leave you with this mechanism to ponder. This raises an important question. We just spent a lot of time emphasizing strict adherence to pre -registered analysis plans to prevent same hacking. But if we force researchers to put on those blinders, what happens to the anomalies? Could the ultimate safeguard against bad science also be the exact mechanism that prevents us from noticing completely unexpected paradigm shifting discoveries hiding right there in the margins of our own spreadsheets?
22:31That is a phenomenal puzzle to wrestle with. A huge thank you from the last minute lecture team for letting us guide you through this material today. The sketch of your research idea is beautiful, but now you have the tools to actually build it. Keep questioning everything and we'll see you on the next deep dive.