0% of the question bank attempted

Types of Variables

Every analysis starts by asking what kind of variable you have, because that decides which graphs and summaries make sense.
  • Categorical (qualitative) variables place individuals into groups: eye color, favorite sport, yes/no answers
  • Quantitative variables are numbers where arithmetic makes sense: height, test score, minutes of sleep
  • A number can still be categorical: zip codes and jersey numbers are labels, so averaging them is meaningless
  • Categorical data → bar charts and pie charts; quantitative data → dotplots, stemplots, histograms and boxplots
  • Always identify the individuals (who or what is measured) and the variables (what is recorded about each)

Describing a Distribution: SOCS

To describe a quantitative distribution, cover Shape, Outliers, Center and Spread, always in context.
  • Shape: symmetric, skewed right (long tail toward large values), skewed left (long tail toward small values); unimodal or bimodal
  • Outliers: values that fall far from the rest; the 1.5 × IQR rule flags points below Q1 − 1.5·IQR or above Q3 + 1.5·IQR
  • Center: the mean (balance point) or median (middle value)
  • Spread: range, interquartile range (IQR = Q3 − Q1) or standard deviation
  • Always name the variable and units, e.g. 'the median commute is 22 minutes', not just 'the median is 22'

Resistant vs. Non-resistant Measures

Some summaries barely move when an extreme value appears; others get dragged toward it.
  • The median and IQR are resistant: an outlier has little effect on them
  • The mean, standard deviation and range are not resistant: one extreme value can move them a lot
  • In a right-skewed distribution the mean is usually greater than the median; in a left-skewed one it is usually less
  • For skewed data or data with outliers, report median and IQR; for roughly symmetric data, mean and standard deviation
  • The five-number summary (min, Q1, median, Q3, max) is what a boxplot draws
A histogram of quantitative data has touching bars; a bar chart of categories has gaps. Don't mix them up.
'Skewed right' means the tail points right (toward larger values), not that most data is on the right.
A large standard deviation doesn't mean the data is 'wrong'. It just means the values are spread out.

Density Curves and the Normal Model

A normal curve is a smooth, symmetric, bell-shaped model described completely by its mean μ and standard deviation σ.
  • The total area under any density curve is 1, and area over an interval is the proportion of values in it
  • A normal curve is symmetric about μ; the mean, median and mode are all equal
  • σ controls the width: a larger σ gives a flatter, wider curve
  • The inflection points of the curve sit one standard deviation from the mean
  • Many measurements are approximately normal (heights, measurement error), but not all. Check with a graph first

The Empirical Rule (68–95–99.7)

In a normal distribution, fixed percentages of the data fall within 1, 2 and 3 standard deviations of the mean.
  • About 68% of values fall within μ ± 1σ
  • About 95% fall within μ ± 2σ
  • About 99.7% fall within μ ± 3σ
  • By symmetry, about 16% lie above μ + 1σ and about 2.5% lie above μ + 2σ
  • The rule only applies to distributions that are approximately normal

z-Scores and Percentiles

A z-score says how many standard deviations a value is from the mean, which lets you compare values from different distributions.
  • z = (x − μ) / σ; positive z is above the mean, negative z is below
  • A z-score of 2 means the value is two standard deviations above average
  • The pth percentile is the value with p% of the data at or below it
  • For normal data, a z-table or calculator (normalcdf) converts z-scores to proportions
  • Comparing z-scores shows which result is more unusual relative to its own group
Don't apply the empirical rule to skewed data. It only works for approximately normal distributions.
A negative z-score isn't 'bad'. It only means the value is below the mean.
Percentile is not percent correct: scoring at the 90th percentile means beating about 90% of others.

Scatterplots and Association

A scatterplot shows the relationship between two quantitative variables measured on the same individuals.
  • Put the explanatory variable on the x-axis and the response variable on the y-axis
  • Describe direction (positive/negative), form (linear/curved), strength and unusual points
  • Positive association: as x increases, y tends to increase; negative: y tends to decrease
  • An outlier in a scatterplot falls outside the overall pattern
  • An influential point is one whose removal would noticeably change the line or correlation

The Correlation Coefficient r

r measures the direction and strength of a linear relationship on a scale from −1 to 1.
  • r near +1 or −1 means a strong linear relationship; r near 0 means a weak linear relationship
  • r has no units and doesn't change if you swap x and y or change units of measurement
  • r only measures linear association: a perfect curve can have r close to 0
  • r is not resistant: a single outlier can greatly change it
  • r² (coefficient of determination) is the fraction of variation in y explained by the linear model with x

Why Correlation Isn't Causation

Two variables can move together without one causing the other.
  • A confounding (lurking) variable can drive both: ice cream sales and drownings both rise with summer heat
  • The direction could be reversed: does exercise improve mood, or do happier people exercise more?
  • Coincidence: with enough variables, some will correlate by chance
  • Only a well-designed randomized experiment can establish cause and effect
  • Observational studies can show association and motivate experiments, but can't prove causation
r = 0 doesn't mean 'no relationship'. It means no linear relationship; there could be a strong curve.
A strong correlation between two variables doesn't mean one causes the other.
Don't extrapolate: predictions far outside the range of x-values in the data are unreliable.

Populations, Samples and Bias

We study a sample to learn about a population, which only works if the sample represents it.
  • Population: the entire group we want information about; sample: the part we actually examine
  • A parameter describes a population (μ, p); a statistic describes a sample (x̄, p̂)
  • Bias is a systematic tendency to over- or under-estimate the truth, and bigger samples don't fix it
  • Voluntary response samples (online polls, call-ins) overrepresent people with strong opinions
  • Convenience samples (whoever is easy to reach) usually aren't representative

Random Sampling Methods

Chance-based sampling methods avoid favoritism and let us quantify uncertainty.
  • Simple random sample (SRS): every group of n individuals has the same chance of being chosen
  • Stratified random sample: split the population into similar groups (strata) and take an SRS from each
  • Cluster sample: split into groups (clusters), randomly select whole clusters, and survey everyone in them
  • Systematic sample: pick a random start, then every kth individual
  • Undercoverage, nonresponse and wording of questions can still bias results even with random selection

Observational Studies vs. Experiments

In an experiment the researcher imposes a treatment; in an observational study they only watch.
  • Experiments use random assignment to create comparable groups, which allows cause-and-effect conclusions
  • Principles of experimental design: comparison, random assignment, control of other variables, replication
  • A placebo is a fake treatment; blinding keeps subjects (single-blind) or also evaluators (double-blind) unaware of groups
  • A block design groups similar subjects first, then randomizes treatments within each block
  • Random sampling lets you generalize to the population; random assignment lets you infer causation
Stratified and cluster sampling are different: stratified samples from every group, cluster samples a few whole groups.
An observational study with a huge sample still can't prove causation.
Random assignment and random sampling aren't the same thing. They answer different questions.

What Probability Means

Probability is the long-run relative frequency of an outcome over many repetitions of a chance process.
  • Probabilities are between 0 (impossible) and 1 (certain)
  • The sample space is the set of all possible outcomes; all their probabilities add to 1
  • Law of large numbers: as trials increase, the proportion of an outcome approaches its true probability
  • Short runs can look streaky; the 'gambler's fallacy' wrongly expects chance to self-correct in the short term
  • Simulation estimates probabilities by imitating a chance process many times

Probability Rules

A few rules combine the probabilities of events.
  • Complement rule: P(not A) = 1 − P(A)
  • Addition rule: P(A or B) = P(A) + P(B) − P(A and B)
  • Mutually exclusive events can't both happen, so P(A and B) = 0
  • Multiplication rule for independent events: P(A and B) = P(A) · P(B)
  • 'At least one' problems are easiest with the complement: 1 − P(none)

Conditional Probability and Independence

Conditional probability updates the chance of an event once we know another event occurred.
  • P(A | B) = P(A and B) / P(B), read 'probability of A given B'
  • Events are independent if knowing one occurred doesn't change the probability of the other: P(A | B) = P(A)
  • Two-way tables make conditional probabilities concrete: restrict to the row or column you're given
  • Tree diagrams organize multi-stage processes; multiply along branches, add across branches
  • Mutually exclusive events with nonzero probabilities are never independent
Don't add probabilities of events that can happen together without subtracting the overlap.
Mutually exclusive isn't the same as independent. Disjoint events are dependent.
P(A | B) is usually not equal to P(B | A).

Sampling Distributions

A statistic varies from sample to sample; its sampling distribution describes that variation.
  • The sampling distribution of x̄ is centered at μ: the sample mean is an unbiased estimator
  • Its standard deviation is σ/√n, so quadrupling n cuts the variability in half
  • Central Limit Theorem: for large n (often n ≥ 30), the distribution of x̄ is approximately normal regardless of the population's shape
  • For proportions, p̂ is centered at p with standard deviation √(p(1−p)/n)
  • Variability of a statistic depends on sample size, not on population size (as long as the population is much larger)

Building a Confidence Interval

A confidence interval gives a range of plausible values for a parameter: estimate ± margin of error.
  • General form: statistic ± (critical value)(standard error)
  • For a proportion: p̂ ± z*·√(p̂(1−p̂)/n); z* = 1.645 for 90%, 1.96 for 95%, 2.576 for 99%
  • For a mean with unknown σ: x̄ ± t*·s/√n, using the t distribution with n − 1 degrees of freedom
  • Conditions: random sample, independence (n ≤ 10% of population), and normality/large counts
  • Large counts for proportions: n·p̂ ≥ 10 and n(1 − p̂) ≥ 10

Interpreting Confidence

The confidence level describes the method, not any single interval.
  • Interpret an interval: 'We are 95% confident the interval from __ to __ captures the true [parameter in context]'
  • Confidence level: if we repeated the sampling many times, about 95% of the intervals would capture the parameter
  • Higher confidence → wider interval; larger sample size → narrower interval
  • The margin of error covers only random sampling error, not bias from poor sampling or wording
  • It's wrong to say there is a 95% probability the parameter is in this specific interval: the parameter is fixed
A 95% confidence interval doesn't contain 95% of the data values. It estimates a parameter.
The margin of error doesn't account for bias.
Doubling the sample size doesn't halve the margin of error. You need four times the sample.

Hypotheses and the Logic of a Test

A significance test asks whether sample evidence is strong enough to reject a claim about a parameter.
  • Null hypothesis H₀: the 'no effect' or status-quo claim, e.g. p = 0.5
  • Alternative hypothesis Hₐ: what we suspect is true instead, e.g. p > 0.5 (one-sided) or p ≠ 0.5 (two-sided)
  • Hypotheses are always about parameters (p, μ), never about sample statistics
  • We assume H₀ is true and ask how surprising our sample result would be
  • Test statistic: (statistic − null value) / standard error

P-values and Decisions

The P-value measures how surprising the data are if the null hypothesis is true.
  • P-value: the probability, assuming H₀ is true, of getting a result at least as extreme as the one observed
  • Small P-value → the data would be unusual if H₀ were true → evidence against H₀
  • Compare with the significance level α (often 0.05): if P ≤ α, reject H₀; otherwise fail to reject H₀
  • Failing to reject H₀ doesn't prove H₀ is true; there just isn't convincing evidence against it
  • Always state the conclusion in context: 'There is convincing evidence that the true proportion of...'

Errors and Power

Every decision can be wrong, and the two kinds of mistakes have different consequences.
  • Type I error: rejecting H₀ when it's actually true (a false alarm); P(Type I) = α
  • Type II error: failing to reject H₀ when Hₐ is true (a missed effect)
  • Power = P(rejecting H₀ when Hₐ is true) = 1 − P(Type II)
  • Power increases with larger sample size, larger α, and a larger true effect
  • Statistical significance isn't practical importance: a tiny effect can be 'significant' with a huge sample
The P-value is not the probability that H₀ is true.
Never 'accept H₀'. You either reject it or fail to reject it.
Hypotheses use parameters (p, μ), not statistics (p̂, x̄).

Misleading Graphs

A graph can be accurate in its numbers and still leave a false impression.
  • Truncated y-axis: starting the axis above 0 makes small differences look huge in bar charts
  • Inconsistent scales or uneven time intervals distort trends
  • Pictographs that scale an image in two dimensions exaggerate differences (double the height → four times the area)
  • 3-D and pie charts with many slices make comparisons hard to judge
  • Cherry-picked time windows can make a trend appear or disappear

Misleading Numbers

The same data can be summarized in ways that inflate or hide an effect.
  • Relative vs. absolute risk: a risk rising from 1 in 10,000 to 2 in 10,000 is a '100% increase' but only 1 extra case per 10,000
  • Averages can mislead with skewed data: the mean income of a room changes a lot when a billionaire walks in
  • Percent vs. percentage points: going from 10% to 15% is a 5-percentage-point rise, but a 50% increase
  • Survivorship bias: studying only the successes (e.g. companies still in business) hides the failures
  • Base rates matter: a 99%-accurate test for a rare condition still produces many false positives

Simpson's Paradox and Reading Claims Critically

A trend in every group can reverse when the groups are combined.
  • Simpson's paradox happens when a lurking variable is unevenly distributed across groups
  • Classic example: a treatment can have a higher success rate in each hospital separately but a lower rate overall
  • Ask who was studied, how they were chosen, and what was compared
  • Check whether the study was an experiment or observational before accepting a causal claim
  • Look for sample size, margin of error and whether the source has a stake in the result
Don't judge a bar chart by bar heights alone. Check where the y-axis starts.
'Doubles your risk' can still mean a very small absolute risk.
An impressive statistic from a biased sample is still biased.
Term
Press Enter or Space to flip the card. Left and right arrows move between cards. 1 marks it known, 2 marks it still learning.
Click or press Enter to flip · Rate yourself to track weak cards
Browse all 64 flashcards as a list

Unit 1: Describing Data & Graphs

Categorical variable
Places each individual into a group or category (e.g. blood type). Summarize with counts or percents.
Quantitative variable
A numerical value where arithmetic makes sense (e.g. weight in kg).
SOCS
Shape, Outliers, Center, Spread: the four things to describe about a quantitative distribution, in context.
Skewed right
Long tail toward larger values; mean is usually pulled above the median.
Median
The middle value of ordered data (average of the two middle values if n is even). Resistant to outliers.
IQR
Q3 − Q1; the spread of the middle half of the data. Resistant to outliers.
1.5 × IQR rule
A value is an outlier if it is below Q1 − 1.5·IQR or above Q3 + 1.5·IQR.
Standard deviation
Roughly the typical distance of data values from the mean. Not resistant.

Unit 2: Distributions & Normal Curve

Normal distribution
Symmetric, bell-shaped density curve fully described by mean μ and standard deviation σ.
Empirical rule
In normal data: ~68% within 1σ, ~95% within 2σ, ~99.7% within 3σ of the mean.
z-score
z = (x − μ)/σ; the number of standard deviations a value is from the mean.
Percentile
The percent of observations at or below a given value.
Density curve
A curve on or above the axis with total area 1; area over an interval = proportion of observations.
Standard normal
The normal distribution with mean 0 and SD 1; the distribution of z-scores.
Inflection points
Where a normal curve changes concavity; located at μ ± σ.
normalcdf
Calculator function giving the area (proportion) under a normal curve between two bounds.

Unit 3: Correlation vs Causation

Scatterplot
Graph of two quantitative variables on the same individuals; explanatory on x, response on y.
Correlation r
Measures direction and strength of a linear relationship; between −1 and 1, no units.
r²
Coefficient of determination: fraction of the variation in y explained by the linear model.
Confounding variable
A lurking variable related to both explanatory and response, making cause unclear.
Residual
Observed y − predicted y; the vertical distance from a point to the regression line.
Extrapolation
Using a model to predict far outside the range of the data. Unreliable.
Least-squares line
The line minimizing the sum of squared residuals; passes through (x̄, ȳ).
Influential point
A point whose removal markedly changes the slope, intercept or correlation.

Unit 4: Sampling & Study Design

Population vs. sample
Population: everyone of interest. Sample: the subset actually studied.
Parameter vs. statistic
Parameter describes a population (μ, p); statistic describes a sample (x̄, p̂).
Simple random sample
Every set of n individuals has an equal chance of being chosen.
Stratified sample
Divide into similar strata; take an SRS within each stratum.
Cluster sample
Divide into clusters; randomly choose whole clusters and survey everyone in them.
Confounding
When the effects of two variables on the response can't be separated.
Random assignment
Using chance to place subjects into treatment groups; allows causal conclusions.
Placebo effect
Improvement caused by believing you received treatment, even a fake one.

Unit 5: Probability Basics

Probability
The long-run relative frequency of an outcome; a number from 0 to 1.
Sample space
The set of all possible outcomes of a chance process.
Complement rule
P(not A) = 1 − P(A).
Addition rule
P(A or B) = P(A) + P(B) − P(A and B).
Mutually exclusive
Events that can't happen at the same time; P(A and B) = 0.
Independent events
One occurring doesn't change the probability of the other; P(A and B) = P(A)P(B).
Conditional probability
P(A | B) = P(A and B)/P(B): the chance of A given that B occurred.
Law of large numbers
Over many trials, the observed proportion approaches the true probability.

Unit 6: Confidence Intervals

Sampling distribution
The distribution of a statistic's values over all possible samples of the same size.
Central Limit Theorem
For large n, the sampling distribution of x̄ is approximately normal, regardless of population shape.
Standard error
Estimated standard deviation of a statistic, e.g. s/√n or √(p̂(1−p̂)/n).
Confidence interval
statistic ± margin of error; a range of plausible values for a parameter.
Margin of error
critical value × standard error; accounts for random sampling variation only.
Confidence level
Long-run percent of intervals from the method that capture the true parameter.
z* values
90%: 1.645 95%: 1.96 99%: 2.576.
Large counts condition
n·p̂ ≥ 10 and n(1 − p̂) ≥ 10, so p̂ is approximately normal.

Unit 7: Hypothesis Testing Logic

Null hypothesis H₀
The claim of no effect or no difference, stated about a parameter.
Alternative hypothesis Hₐ
The claim we seek evidence for: >, < (one-sided) or ≠ (two-sided).
P-value
Probability, assuming H₀ is true, of a result at least as extreme as the one observed.
Significance level α
Cutoff for 'unusual'; reject H₀ if P-value ≤ α. Equals P(Type I error).
Type I error
Rejecting H₀ when it's actually true (false positive).
Type II error
Failing to reject H₀ when Hₐ is true (false negative).
Power
Probability of correctly rejecting a false H₀; 1 − P(Type II).
Test statistic
(statistic − null value)/standard error.

Unit 8: Misleading Statistics in the Wild

Truncated axis
A y-axis that doesn't start at 0, exaggerating differences in bar charts.
Relative vs. absolute risk
Relative: percent change in risk. Absolute: actual change in cases per population.
Percentage points
The arithmetic difference between two percentages (4% → 6% is +2 points).
Survivorship bias
Drawing conclusions only from cases that 'survived' a selection process.
Simpson's paradox
A trend within every group reverses when the groups are combined, due to a lurking variable.
Base rate fallacy
Ignoring how rare a condition is when interpreting a positive test result.
Cherry-picking
Showing only the data or time window that supports a conclusion.
Conflict of interest
When a study's source benefits from a particular result.
Press 1–4 to answer · Enter for next

Unit 1: Describing Data & Graphs

Types of Variables
Every analysis starts by asking what kind of variable you have, because that decides which graphs and summaries make sense.
Describing a Distribution: SOCS
To describe a quantitative distribution, cover Shape, Outliers, Center and Spread, always in context.
Resistant vs. Non-resistant Measures
Some summaries barely move when an extreme value appears; others get dragged toward it.
Key fact
Mean > median usually signals right skew; mean < median usually signals left skew.
Key fact
IQR = Q3 − Q1 measures the spread of the middle 50% of the data.
Key fact
Outlier fences: Q1 − 1.5·IQR and Q3 + 1.5·IQR.
Key fact
Standard deviation is roughly the typical distance of values from the mean.

Unit 2: Distributions & Normal Curve

Density Curves and the Normal Model
A normal curve is a smooth, symmetric, bell-shaped model described completely by its mean μ and standard deviation σ.
The Empirical Rule (68–95–99.7)
In a normal distribution, fixed percentages of the data fall within 1, 2 and 3 standard deviations of the mean.
z-Scores and Percentiles
A z-score says how many standard deviations a value is from the mean, which lets you compare values from different distributions.
Key fact
68–95–99.7: the percent of normal data within 1, 2 and 3 standard deviations of the mean.
Key fact
z = (x − μ)/σ measures position in standard deviations.
Key fact
z = 0 is the mean; about 84% of normal data lies below z = 1.
Key fact
The area under a density curve is always exactly 1.

Unit 3: Correlation vs Causation

Scatterplots and Association
A scatterplot shows the relationship between two quantitative variables measured on the same individuals.
The Correlation Coefficient r
r measures the direction and strength of a linear relationship on a scale from −1 to 1.
Why Correlation Isn't Causation
Two variables can move together without one causing the other.
Key fact
−1 ≤ r ≤ 1; the sign gives direction, the magnitude gives strength of the linear relationship.
Key fact
r² = proportion of variation in y explained by the regression on x.
Key fact
Correlation doesn't imply causation; confounding variables are the usual culprit.
Key fact
Least-squares regression line: ŷ = a + bx, passing through (x̄, ȳ).

Unit 4: Sampling & Study Design

Populations, Samples and Bias
We study a sample to learn about a population, which only works if the sample represents it.
Random Sampling Methods
Chance-based sampling methods avoid favoritism and let us quantify uncertainty.
Observational Studies vs. Experiments
In an experiment the researcher imposes a treatment; in an observational study they only watch.
Key fact
Random sampling → generalize to the population. Random assignment → conclude cause and effect.
Key fact
Parameters describe populations (μ, p); statistics describe samples (x̄, p̂).
Key fact
Larger samples reduce variability but do not reduce bias.
Key fact
Double-blind: neither subjects nor those measuring outcomes know who got which treatment.

Unit 5: Probability Basics

What Probability Means
Probability is the long-run relative frequency of an outcome over many repetitions of a chance process.
Probability Rules
A few rules combine the probabilities of events.
Conditional Probability and Independence
Conditional probability updates the chance of an event once we know another event occurred.
Key fact
P(not A) = 1 − P(A).
Key fact
P(A or B) = P(A) + P(B) − P(A and B).
Key fact
Independent: P(A and B) = P(A)·P(B), equivalently P(A | B) = P(A).
Key fact
P(at least one) = 1 − P(none).

Unit 6: Confidence Intervals

Sampling Distributions
A statistic varies from sample to sample; its sampling distribution describes that variation.
Building a Confidence Interval
A confidence interval gives a range of plausible values for a parameter: estimate ± margin of error.
Interpreting Confidence
The confidence level describes the method, not any single interval.
Key fact
Standard deviation of x̄ = σ/√n; of p̂ = √(p(1−p)/n).
Key fact
95% confidence uses z* = 1.96.
Key fact
Margin of error shrinks as n grows (by a factor of 1/√n) and grows with confidence level.
Key fact
CLT: x̄ is approximately normal for large n, whatever the population's shape.

Unit 7: Hypothesis Testing Logic

Hypotheses and the Logic of a Test
A significance test asks whether sample evidence is strong enough to reject a claim about a parameter.
P-values and Decisions
The P-value measures how surprising the data are if the null hypothesis is true.
Errors and Power
Every decision can be wrong, and the two kinds of mistakes have different consequences.
Key fact
P-value = P(result at least this extreme | H₀ true).
Key fact
Reject H₀ when P-value ≤ α.
Key fact
Type I = false positive (probability α); Type II = false negative.
Key fact
Power = 1 − P(Type II error); bigger n means more power.

Unit 8: Misleading Statistics in the Wild

Misleading Graphs
A graph can be accurate in its numbers and still leave a false impression.
Misleading Numbers
The same data can be summarized in ways that inflate or hide an effect.
Simpson's Paradox and Reading Claims Critically
A trend in every group can reverse when the groups are combined.
Key fact
Truncated axes exaggerate differences in bar charts.
Key fact
Relative risk sounds bigger than absolute risk. Ask for both.
Key fact
A change from 10% to 15% is 5 percentage points but a 50% relative increase.
Key fact
Simpson's paradox: a lurking variable can reverse a trend when groups are combined.
Common mistakes for each unit — read the mistake, then make sure you know why it's wrong.

Unit 1: Describing Data & Graphs

Watch out
A histogram of quantitative data has touching bars; a bar chart of categories has gaps. Don't mix them up.
Watch out
'Skewed right' means the tail points right (toward larger values), not that most data is on the right.
Watch out
A large standard deviation doesn't mean the data is 'wrong'. It just means the values are spread out.

Unit 2: Distributions & Normal Curve

Watch out
Don't apply the empirical rule to skewed data. It only works for approximately normal distributions.
Watch out
A negative z-score isn't 'bad'. It only means the value is below the mean.
Watch out
Percentile is not percent correct: scoring at the 90th percentile means beating about 90% of others.

Unit 3: Correlation vs Causation

Watch out
r = 0 doesn't mean 'no relationship'. It means no linear relationship; there could be a strong curve.
Watch out
A strong correlation between two variables doesn't mean one causes the other.
Watch out
Don't extrapolate: predictions far outside the range of x-values in the data are unreliable.

Unit 4: Sampling & Study Design

Watch out
Stratified and cluster sampling are different: stratified samples from every group, cluster samples a few whole groups.
Watch out
An observational study with a huge sample still can't prove causation.
Watch out
Random assignment and random sampling aren't the same thing. They answer different questions.

Unit 5: Probability Basics

Watch out
Don't add probabilities of events that can happen together without subtracting the overlap.
Watch out
Mutually exclusive isn't the same as independent. Disjoint events are dependent.
Watch out
P(A | B) is usually not equal to P(B | A).

Unit 6: Confidence Intervals

Watch out
A 95% confidence interval doesn't contain 95% of the data values. It estimates a parameter.
Watch out
The margin of error doesn't account for bias.
Watch out
Doubling the sample size doesn't halve the margin of error. You need four times the sample.

Unit 7: Hypothesis Testing Logic

Watch out
The P-value is not the probability that H₀ is true.
Watch out
Never 'accept H₀'. You either reject it or fail to reject it.
Watch out
Hypotheses use parameters (p, μ), not statistics (p̂, x̄).

Unit 8: Misleading Statistics in the Wild

Watch out
Don't judge a bar chart by bar heights alone. Check where the y-axis starts.
Watch out
'Doubles your risk' can still mean a very small absolute risk.
Watch out
An impressive statistic from a biased sample is still biased.