0% of the question bank attempted

Shape, Center, and Spread

Every distribution of quantitative data is described by its shape (symmetric, skewed, unimodal/bimodal), center (mean or median), and spread (range, IQR, or standard deviation), along with any outliers.
  • Always describe distributions with SOCS: Shape, Outliers, Center, Spread, in context
  • Skewed left: tail points left, mean < median; skewed right: tail points right, mean > median
  • Symmetric distributions have mean ≈ median
  • Unimodal $=$ one peak, bimodal $=$ two peaks, uniform $=$ flat
  • Use the shape to decide which center/spread measures are appropriate

Measures of Center

The mean is the arithmetic average and is sensitive to outliers/skew, while the median is the middle value and resists extreme values.
  • Mean: $x̄ = (Σx)/n$
  • Median: middle value of ordered data (average of two middle values if n is even)
  • Mean is pulled toward the tail in a skewed distribution; median is resistant
  • For skewed data or data with outliers, report the median as the 'typical' value
  • Mode is the most frequent value; useful mainly for categorical or discrete data

Measures of Spread

Spread describes how variable the data are; range and IQR are resistant-friendly, while variance and standard deviation use every value.
  • Range $=$ max − min
  • IQR $=$ Q3 − Q1 (spread of the middle 50%), resistant to outliers
  • Variance: s² $= Σ(x − x̄)²/(n − 1)$; Standard deviation s $=$ √(variance)
  • SD measures the typical distance of a value from the mean
  • Standard deviation is not resistant — outliers inflate it
  • Use IQR with median (skewed data); use SD with mean (roughly symmetric data)

Five-Number Summary, Boxplots, and Outliers

The five-number summary (min, Q1, median, Q3, max) is displayed visually with a boxplot, and the 1.5×IQR rule flags potential outliers.
  • Five-number summary: Min, Q1, Median, Q3, Max
  • Outlier rule: any value below Q1 − 1.5(IQR) or above Q3 + 1.5(IQR)
  • Boxplots show the IQR as a box with whiskers to the most extreme non-outlier values
  • Modified boxplots plot outliers as separate points
  • Boxplots are good for comparing distributions across groups side by side

z-Scores and Standardization

A z-score converts a raw value into the number of standard deviations it lies from the mean, letting you compare values from different distributions.
  • z $=$ (x − $μ$) / $σ$ (or $(x − x̄)/s$ for sample data)
  • Positive z: value is above the mean; negative z: value is below the mean
  • Standardizing does not change the shape of a distribution
  • z-scores allow comparison of values measured on different scales (e.g., test scores from different exams)
  • Adding a constant to every value shifts center but not spread; multiplying by a constant changes both center and spread proportionally

Comparing Distributions and Density

When comparing two or more distributions, always compare shape, center, spread, and outliers in context rather than describing them separately.
  • Use parallel boxplots, back-to-back stemplots, or dotplots to compare groups
  • Comment on which group has the higher center and which has more/less variability
  • Relative frequency (percent) is used, not counts, when group sizes differ
  • Density curves show the proportion of data as area under the curve, total area $=$ 1
  • Always write comparative statements in context, using words like 'greater than' or 'less variable than'
‘Skewed left’ describes the direction of the longer tail, not where most data pile up
Standard deviation can never be negative; a computed negative SD means an arithmetic error
Adding a constant to all data shifts the mean but leaves the standard deviation unchanged — only multiplying changes spread
Comparing distributions requires comparative language in context, not two separate one-group descriptions

Density Curves

A density curve is a smooth curve that models the shape of a distribution, where the total area underneath equals 1 and area corresponds to proportion of data.
  • Total area under a density curve $=$ 1 (100% of data)
  • Area under the curve over an interval $=$ proportion of observations in that interval
  • The median divides the area in half; the mean is the balance point of the curve
  • In a skewed density curve, the mean is pulled toward the longer tail relative to the median
  • Density curves are idealized models, not the actual data

The Normal Distribution

The Normal distribution is a symmetric, bell-shaped density curve fully described by its mean $μ$ and standard deviation $σ$, following the Empirical (68-95-99.7) Rule.
  • Notation: $N(μ, σ)$
  • 68% of values fall within 1 SD of the mean ($μ$ ± $σ$)
  • 95% of values fall within 2 SD of the mean ($μ$ ± $2σ$)
  • 99.7% of values fall within 3 SD of the mean ($μ$ ± $3σ$)
  • Normal curves are symmetric with mean $=$ median $=$ mode at the center
  • Inflection points of the curve occur at $μ$ ± $σ$

Standardizing and z-Scores

Any Normal distribution can be converted to the standard Normal distribution N(0,1) by standardizing values with z $=$ $(x − μ)/σ$.
  • z $=$ $(x − μ)/σ$ converts x to a standard score
  • Standard Normal distribution: N(0, 1)
  • Use a z-table or calculator (normalcdf) to find proportions/probabilities
  • invNorm is used to find a value given a percentile/proportion
  • z-scores let you compare observations from different Normal distributions

Normal Probability Calculations

Calculator functions normalcdf and invNorm let you compute probabilities/proportions and find values from percentiles for any Normal distribution.
  • normalcdf(lower, upper, $μ, σ$) finds the proportion/probability between two values
  • invNorm(percentile, $μ, σ$) finds the value at a given percentile
  • For 'less than', use normalcdf(-1E99, x, $μ, σ$); for 'greater than', use normalcdf(x, 1E99, $μ, σ$)
  • Always sketch and shade the Normal curve before calculating
  • Check reasonableness: answers should be between 0 and 1 for probabilities

Assessing Normality

A Normal probability plot (or comparing the data to the 68-95-99.7 rule) helps determine whether a data set is approximately Normally distributed.
  • A Normal probability plot that is roughly linear suggests the data are approximately Normal
  • Curvature in the plot indicates skewness or non-Normality
  • Compare observed percentages within 1, 2, 3 SD of the mean to 68%, 95%, 99.7%
  • Histograms/boxplots showing strong skew or outliers suggest data are not Normal
  • Many statistical procedures assume Normality, so checking it matters for inference later
The Empirical Rule only applies to (approximately) Normal distributions, not all distributions
A Normal probability plot with a clear curve — not a straight line — signals the data are NOT Normal
Area under a density curve represents proportion, not probability of a single exact point (P(X $=$ x) $=$ 0 for continuous variables)
z-scores can be negative; a negative z simply means the value is below the mean, not an error

Scatterplots and Association

A scatterplot displays the relationship between two quantitative variables, revealing direction, form, strength, and unusual features (outliers).
  • Explanatory variable (x) is plotted on the horizontal axis; response variable (y) on the vertical axis
  • Direction: positive (as x increases, y increases) or negative association
  • Form: linear or nonlinear (curved) pattern
  • Strength: how closely points follow the pattern (tight vs. scattered)
  • Always note any outliers or unusual points separately

Correlation (r)

The correlation coefficient r measures the strength and direction of a linear relationship between two quantitative variables, ranging from -1 to 1.
  • r ranges from -1 to +1; sign matches the direction of association
  • r close to ±1 indicates a strong linear relationship; r close to 0 indicates a weak/no linear relationship
  • r measures only linear association — it can be misleading for curved relationships
  • r has no units and is not affected by which variable is x or y
  • r is strongly affected by outliers
  • r² (coefficient of determination) is discussed separately

Least-Squares Regression Line

The least-squares regression line ŷ $=$ a + bx minimizes the sum of squared residuals and is used to predict y from x.
  • Slope b $=$ r(sy/sx); the line passes through ($x̄$, ȳ)
  • Interpretation of slope: 'for each one-unit increase in x, ŷ is predicted to change by b units'
  • Interpretation of y-intercept a: predicted value of y when x $=$ 0 (only meaningful if $x=0$ is in scope)
  • Extrapolation (predicting far outside the observed x-range) is unreliable
  • Regression line is only appropriate for genuinely linear relationships

Residuals and Residual Plots

A residual is the difference between an observed and predicted value (residual $=$ y − ŷ), and residual plots help check whether a linear model is appropriate.
  • Residual $=$ observed y − predicted ŷ
  • A residual plot with no pattern (random scatter around 0) supports a linear model
  • A curved pattern in the residual plot indicates the relationship is not linear
  • Increasing spread (fan shape) in residuals indicates non-constant variability
  • The sum of residuals for a least-squares line is always 0

r² (Coefficient of Determination)

r² is the proportion of variation in y that is explained by the linear regression on x, and is always between 0 and 1.
  • r² $=$ (correlation r)²
  • Interpretation: 'r²×100% of the variation in y is explained by the linear relationship with x'
  • r² closer to 1 means the regression line explains most of the variability in y
  • r² does not indicate causation, only explanatory strength of the linear model
  • Standard deviation of residuals (s) measures typical prediction error in y's units

Outliers, Influential Points, and Transformations

Outliers and especially influential points (often high-leverage x-values) can dramatically change the regression line, and nonlinear data may require transformation.
  • An outlier in regression has a large residual (far from the pattern)
  • An influential point substantially changes the slope/intercept if removed, often due to extreme x-value
  • Log or power transformations can linearize exponential or power-model data before regression
  • Always check a residual plot before trusting a linear model's predictions
  • Correlation/regression describe association, never prove causation
Correlation does NOT imply causation — a strong r only shows association
A high r² does not fix a nonlinear relationship; always check the residual plot first
Extrapolation beyond the observed x-range is unreliable even with a high r²
Correlation r is not resistant — a single outlier can drastically change its value

Sampling Methods

Good sampling methods use chance to select a subset of a population so the sample represents the whole population without bias.
  • Simple Random Sample (SRS): every group of n individuals has an equal chance of being chosen
  • Stratified random sample: population divided into homogeneous strata, then SRS taken within each stratum
  • Cluster sample: population divided into heterogeneous clusters, then entire clusters are randomly selected
  • Systematic sample: select every kth individual from a list after a random start
  • Convenience and voluntary response samples are NOT random and typically produce biased results

Bias in Sampling

Bias occurs when a sampling method systematically favors certain outcomes, making the sample not representative of the population.
  • Undercoverage: some groups of the population are left out of the sampling frame
  • Nonresponse bias: individuals selected don't respond, and responders differ from non-responders
  • Response bias: participants respond inaccurately due to question wording, interviewer, or social pressure
  • Voluntary response bias: sample consists only of people who choose to respond (e.g., call-in polls), often overrepresenting strong opinions
  • Larger sample size does NOT fix a biased sampling method

Observational Studies vs. Experiments

An observational study measures variables without imposing treatments, while an experiment deliberately imposes treatments to study cause-and-effect.
  • Observational studies can show association but cannot establish causation due to potential confounding
  • Experiments impose a treatment on subjects and can establish causation if well-designed
  • Confounding variables are variables whose effects on the response cannot be separated from the explanatory variable's effect
  • Retrospective (case-control) and prospective (cohort) studies are types of observational studies
  • Random assignment (not random sampling) is what allows experiments to infer causation

Principles of Experimental Design

Well-designed experiments use control, randomization, and replication to reduce bias and allow cause-and-effect conclusions.
  • Control: control for lurking variables by comparing multiple treatments, e.g., including a control/placebo group
  • Randomize: use chance to assign subjects to treatment groups, balancing out lurking variables
  • Replicate: apply each treatment to many subjects to reduce the role of chance variation
  • Placebo effect: subjects respond simply to receiving any treatment; a placebo group controls for this
  • Double-blind: neither subjects nor those measuring the response know which treatment was given, reducing bias

Blocking and Matched Pairs

Blocking groups experimental units by a variable expected to affect the response before randomizing treatments within each block, increasing precision.
  • A block is a group of experimental units known to be similar in some way that affects the response (e.g., gender, age group)
  • Randomize treatments separately within each block ('block what you can, randomize what you cannot')
  • Matched pairs design: a special case of blocking with blocks of size two (e.g., before/after, or twins)
  • Blocking reduces variability from the blocking variable, making treatment effects easier to detect
  • Blocking is not the same as stratifying in sampling, though the logic is similar

Scope of Inference

The scope of conclusions you can draw depends on whether random sampling and/or random assignment were used in a study.
  • Random sampling → results can be generalized to the population
  • Random assignment → cause-and-effect conclusions can be drawn
  • Random sampling AND random assignment → can generalize causal conclusions to the population
  • No randomization at all → conclusions apply only to the subjects studied, and only as association, not causation
  • Always identify explanatory and response variables, and check for confounding before making claims
A large sample size cannot fix bias from a poor sampling method (e.g., a huge voluntary response sample is still biased)
Random sampling and random assignment are different: sampling allows generalization, assignment allows causal claims
Observational studies, no matter how large, cannot establish causation because of possible confounding variables
Blocking is done BEFORE random assignment; it's not the same as simply comparing groups after the fact

Basic Probability Rules

Probability quantifies the likelihood of an event, always between 0 and 1, and follows specific rules for combining events.
  • 0 ≤ P(A) ≤ 1 for any event A; P(sample space) $=$ 1
  • Complement Rule: P(not A) $=$ 1 − P(A)
  • Addition Rule (general): P(A or B) $=$ P(A) + P(B) − P(A and B)
  • Mutually exclusive (disjoint) events: P(A and B) $=$ 0, so P(A or B) $=$ P(A) + P(B)
  • Law of Large Numbers: as trials increase, the observed relative frequency approaches the true probability

Independence and Mutual Exclusivity

Two events are independent if the occurrence of one does not change the probability of the other; this is a different concept from mutually exclusive events.
  • Events A and B are independent if P(A|B) $=$ P(A) (and equivalently P(B|A) $=$ P(B))
  • Mutually exclusive events CANNOT both occur (P(A and B) $=$ 0); independent events CAN both occur
  • Mutually exclusive events with nonzero probabilities are NEVER independent
  • Check independence with the multiplication rule: P(A and B) $=$ P(A)×P(B) only if independent
  • Real-world context often needs the general multiplication rule if events aren't independent

Conditional Probability

Conditional probability P(A|B) gives the probability of event A occurring given that event B has already occurred, restricting the sample space to B.
  • P(A|B) $=$ P(A and B) / P(B), provided P(B) > 0
  • Two-way tables are useful for computing conditional probabilities directly from counts
  • Conditional probability restricts attention to only the outcomes where the given condition is true
  • Order matters: P(A|B) is generally not equal to P(B|A)
  • Use conditional probability to formally test for independence

The General Multiplication Rule

The multiplication rule finds the probability that two events both occur: P(A and B) $=$ P(A)×P(B|A), simplifying to P(A)×P(B) when independent.
  • General rule: P(A and B) $=$ P(A) × P(B|A)
  • If A and B are independent: P(A and B) $=$ P(A) × P(B)
  • Tree diagrams visually organize sequential/conditional probabilities and multiply along branches
  • Sampling without replacement typically creates dependent events (probabilities change after each draw)
  • Sampling with replacement typically creates independent events

Simulation

Simulation uses random number generators (or physical devices) to imitate chance behavior, estimating probabilities that are hard to calculate directly.
  • Steps: (1) State the assumptions/model, (2) Assign digits/outcomes to represent events, (3) Run many trials, (4) Compute the estimated probability from the results
  • More trials produce a more accurate simulation estimate (Law of Large Numbers)
  • Simulation gives an APPROXIMATION of the true probability, not an exact answer
  • Random digit tables or calculator random number generators (randInt) are common tools
  • Clearly define what constitutes 'success' in each trial before running the simulation

Two-Way Tables and Venn Diagrams

Two-way tables organize data by two categorical variables, and Venn diagrams visually represent unions, intersections, and complements of events.
  • Marginal probability: probability based on a row or column total divided by the grand total
  • Joint probability: probability of the intersection of two specific categories (a single cell / grand total)
  • Venn diagram overlap represents 'and' (intersection); combined regions represent 'or' (union)
  • Two-way tables make it easy to compute conditional probabilities: (cell count)/(row or column total)
  • Check independence in a two-way table by comparing conditional and marginal probabilities
Mutually exclusive events are NOT independent (if both have positive probability) — a very common mix-up
P(A|B) ≠ P(B|A) in general — order and conditioning direction matter
'At least one' problems are usually easiest using the complement rule: P(at least one) $=$ 1 − P(none)
Sampling without replacement changes probabilities from draw to draw, creating dependence, even if it 'feels' independent for large populations (approximately independent only when population is much larger than sample)

Discrete Random Variables & Probability Distributions

A discrete random variable X takes a countable set of numeric values, each with a probability; the probability distribution lists every value with P(X $=$ x).
  • Valid distribution requires every P(x) ≥ 0 and $ΣP(x) =$ 1
  • Probability histograms display the distribution visually; shape, center, and spread all matter
  • P(X $=$ a) is a single bar; P(X ≤ a) or P(X ≥ a) sums the relevant bars (cumulative probability)
  • Discrete RVs take isolated values (counts); continuous RVs take any value in an interval (measured with density curves/area)
  • Use two-way tables or context to build a distribution from raw counts before computing summary values

Expected Value (Mean) of a Random Variable

The expected value $μX$ is the long-run average outcome of a random variable if the chance process were repeated infinitely many times — a probability-weighted average, not a typical single outcome.
  • $μX =$ E(X) $= Σ$ x·P(x), summing over every possible value
  • E(X) need not be a value X can actually take (e.g., average number of children $=$ 2.3)
  • For a linear transformation: E(a + bX) $=$ a + b·E(X)
  • For the sum of two random variables: E(X + Y) $=$ E(X) + E(Y), always — no independence needed
  • E(X − Y) $=$ E(X) − E(Y), always true as well

Variance and Standard Deviation of a Random Variable

Variance measures the average squared deviation of a random variable from its mean, quantifying spread; standard deviation is its square root, back in the original units.
  • $σ²X =$ Var(X) $= Σ$ $(x − μX)²·P(x); σX =$ √ Var(X)
  • For a linear transformation: Var(a + bX) $=$ b²·Var(X); adding a constant a never changes spread
  • If X and Y are INDEPENDENT: Var(X + Y) $=$ Var(X) + Var(Y) and Var(X − Y) $=$ Var(X) + Var(Y) — variances always add for independent RVs, even when subtracting
  • Standard deviations do NOT add or subtract directly; you must add variances first, then take the square root
  • Combining random variables that are NOT independent requires a covariance term, which is beyond the standard AP formula but the AP exam assumes independence unless stated otherwise

Combining Random Variables

New random variables formed by adding, subtracting, or scaling other random variables follow predictable rules for their mean and variance.
  • Mean of a sum/difference: $μ(X±Y) = μX$ ± $μY$, regardless of independence
  • Variance of a sum/difference (independent X, Y only): $σ²(X±Y) = σ²X$ + $σ²Y$
  • Example: total time for two independent tasks — add means and add variances, then take the square root for the combined SD
  • Watch units and context carefully: 'X − Y' models a difference in outcomes, not a probability of X being less than Y
  • Check independence explicitly before adding variances; dependent variables need extra (non-AP-formula) covariance information

Binomial Distributions

A binomial random variable counts the number of successes in a fixed number of independent trials, each with the same two possible outcomes and the same probability of success.
  • BINS conditions: Binary outcomes, Independent trials, fixed Number of trials n, Same probability p of success on each trial
  • P(X $=$ k) $= C(n,k)·p^k·(1−p)^(n−k)$, where C(n,k) $=$ n!/(k!(n−k)!)
  • Mean: $μX =$ np; Standard deviation: $σX =$ √(np(1−p))
  • 10% condition: sampling without replacement from a finite population is approximately independent (binomial-like) only if the sample is less than 10% of the population
  • Large Counts condition (np ≥ 10 and n(1−p) ≥ 10) allows a Normal approximation to the binomial distribution

Geometric Distributions

A geometric random variable counts the number of trials needed to get the FIRST success in a sequence of independent trials with constant probability of success.
  • Conditions are BITS: Binary outcomes, Independent trials, constant probability p, trials continue until the first Success (no fixed n)
  • P(X $=$ k) $= (1−p)^(k−1)·p$, for k $=$ 1, 2, 3, … (X starts at 1, not 0)
  • Mean: $μX =$ 1/p; Standard deviation: $σX =$ √((1−p)/p²)
  • P(X > n) $= (1−p)^n$, the probability of no success in the first n trials
  • Distinguish from binomial: geometric has no fixed number of trials and asks 'how long until,' while binomial fixes n and counts successes
Variances add for independent RVs even when you SUBTRACT the variables: Var(X−Y) $=$ Var(X)+Var(Y), not Var(X)−Var(Y)
Standard deviations never add or subtract directly — always combine variances first, then take the square root at the end
A binomial variable requires a FIXED number of trials n; if the question asks 'how many trials until the first success,' that's geometric, not binomial
Geometric random variables start at X $=$ 1 (you need at least one trial), never at X $=$ 0

Sampling Distributions: Core Idea

A sampling distribution describes the distribution of a statistic (like a sample mean or sample proportion) computed from every possible sample of a given size — it tells us how much a statistic varies from sample to sample.
  • A statistic (from a sample) estimates a parameter (from the population); statistics vary, parameters are fixed
  • The sampling distribution's shape, center, and spread describe how the statistic behaves across repeated sampling
  • An estimator is unbiased if the mean of its sampling distribution equals the true parameter value
  • Variability of a sampling distribution decreases as sample size n increases
  • Bias relates to accuracy (centered on the parameter); variability relates to precision (how spread out the statistic's values are)

Sampling Distribution of a Sample Proportion

When repeatedly sampling from a population with true proportion p, the sample proportion p̂ has a predictable mean, standard deviation, and (under conditions) approximately Normal shape.
  • Mean: $μp̂ =$ p ($p̂$ is an unbiased estimator of p)
  • Standard deviation: $σp̂ =$ √(p(1−p)/n)
  • 10% condition: sample size n must be less than 10% of the population for independence (so $σp̂$ formula is accurate)
  • Large Counts condition: np ≥ 10 and n(1−p) ≥ 10 for the sampling distribution of p̂ to be approximately Normal
  • As n increases, $σp̂$ decreases (more precise estimates), but the mean stays at p

Sampling Distribution of a Sample Mean

When repeatedly sampling from a population with mean $μ$ and standard deviation $σ$, the sample mean $x̄$ has its own sampling distribution with predictable mean and standard deviation.
  • Mean: $μx̄ = μ$ (x̄ is an unbiased estimator of $μ$) — true for any sample size
  • Standard deviation (standard error): $σx̄ = σ/√n$, requires the 10% condition when sampling without replacement
  • If the population itself is Normally distributed, x̄ is exactly Normally distributed for any n
  • If the population is NOT Normal, x̄'s distribution becomes approximately Normal as n grows large (Central Limit Theorem)
  • Increasing n reduces $σx̄$ (more precision) but does not change $μx̄$

The Central Limit Theorem (CLT)

The Central Limit Theorem states that as sample size n increases, the sampling distribution of the sample mean x̄ approaches a Normal distribution, regardless of the shape of the population distribution.
  • Rule of thumb: CLT applies well once n ≥ 30 for most population shapes, though strongly skewed populations may need larger n
  • CLT justifies using Normal-based procedures (z-scores, confidence intervals) for means even when the population isn't Normal
  • The CLT applies to sums and means of independent, identically distributed random variables
  • The CLT does NOT change $μx̄ = μ$ or $σx̄ = σ/√n$ — it only affects the SHAPE of the sampling distribution
  • For proportions, the analogous Large Counts condition (np≥10, n(1−p)≥10) plays a similar Normal-approximation role

Standard Error and the Effect of Sample Size

Standard error is the estimated standard deviation of a sampling distribution when the population parameter is unknown; it shrinks as sample size grows, following an inverse square-root relationship.
  • Standard error of $p̂: SE(p̂) = √(p̂(1−p̂)/n)$, using the sample proportion since p is unknown
  • Standard error of $x̄: SE(x̄) =$ s/√n, using the sample standard deviation s since $σ$ is unknown
  • To cut the standard error in half, you must QUADRUPLE the sample size (since $σ/√n$ involves a square root)
  • Larger samples give more precise (less variable) estimates, but do not remove bias caused by poor sampling methods
  • Larger sample size never fixes a biased sampling method — that requires proper random sampling design, not more data
Increasing sample size reduces variability ($σ/√n$) but never removes bias from a flawed sampling method
$σx̄ = σ/√n$, NOT $σ/n$ — a common algebra slip that badly underestimates spread
The CLT describes the SHAPE of the sampling distribution of x̄, not the shape of the original population data
To halve the standard error, you must quadruple n, not just double it, because of the square root in $σ/√n$

Confidence Intervals: Logic and Structure

A confidence interval uses a sample statistic plus a margin of error to estimate an unknown population parameter, with a stated level of confidence in the method used to produce it.
  • General form: statistic ± (critical value)×(standard error)
  • Confidence level describes the long-run capture rate of the METHOD: e.g., 95% of all intervals built this way would capture the true parameter over many repetitions
  • Higher confidence level → wider interval (larger critical value), all else equal
  • Larger sample size → narrower interval (smaller standard error), all else equal
  • Conditions to check: Random (random sample/assignment), Normal/Large Counts (shape of sampling distribution), Independent (10% condition)

Confidence Interval for a Proportion and a Mean

The one-proportion z-interval and one-sample t-interval are the two core confidence interval procedures for estimating a population proportion or population mean.
  • One-proportion z-interval: p̂ ± z*·√(p̂(1−p̂)/n)
  • One-sample t-interval for a mean: $x̄$ ± t*·(s/√n), using t* from a t-distribution with n−1 degrees of freedom because $σ$ is unknown
  • The t-distribution is shorter and wider (fatter tails) than the Normal distribution, especially for small n, to account for extra uncertainty from estimating $σ$ with s
  • As degrees of freedom (n−1) increase, the t-distribution approaches the standard Normal distribution
  • Common critical values: $z*=1.645$ (90%), 1.96 (95%), 2.576 (99%); t* depends on both confidence level and df (found via table or technology)

Hypothesis Testing: Logic and Structure

A hypothesis test uses sample evidence to assess a claim about a population, weighing the null hypothesis (no effect/no difference) against an alternative hypothesis.
  • H0 (null): typically states 'no effect' or equality (e.g., p $=$ 0.5, $μ1 = μ2$); Ha (alternative): what we suspect is true (≠, <, or >)
  • The P-value is the probability of getting a test statistic at least as extreme as the one observed, ASSUMING H0 is true
  • Small P-value $=$ strong evidence against H0; if P-value < $α$ (significance level), reject H0; otherwise, fail to reject H0
  • 'Fail to reject H0' is NOT the same as 'accept H0' or 'prove H0 true' — it just means insufficient evidence against it
  • Same Random/Normal(or Large Counts)/Independent conditions from confidence intervals also apply before running a hypothesis test

Type I and Type II Errors, and Power

Every hypothesis test risks two types of mistakes: rejecting a true null hypothesis (Type I) or failing to reject a false null hypothesis (Type II); power measures the test's ability to correctly detect a real effect.
  • Type I error: rejecting H0 when H0 is actually TRUE; P(Type I error) $= α$, the significance level
  • Type II error: failing to reject H0 when H0 is actually FALSE (Ha is true); P(Type II error) $= β$
  • Power $=$ 1 − $β =$ the probability of correctly rejecting a false H0; higher power is better
  • Decreasing $α$ (e.g., from 0.05 to 0.01) reduces the risk of Type I error but INCREASES the risk of Type II error (lowers power)
  • Increasing sample size increases power without increasing $α$, because it reduces standard error

t-Tests, Chi-Square Tests, and Choosing a Procedure

Different inference procedures fit different data types: t-tests compare means, chi-square tests analyze categorical count data, and each has its own test statistic and distribution.
  • One-sample t-test statistic: t $= (x̄ − μ0)/(s/√n)$, with df $=$ n−1
  • Two-sample t-test statistic (independent samples): t $= (x̄1 − x̄2)/√(s1²/n1 + s2²/n2)$, for comparing two population means
  • Chi-square test statistic: $χ² = Σ$ (observed − expected)²/expected, used for goodness-of-fit, homogeneity, or independence tests with categorical data
  • Chi-square goodness-of-fit compares one categorical variable to a hypothesized distribution; chi-square test for independence/association examines two categorical variables in a two-way table
  • Choose a proportion procedure (z) for categorical data summarized as counts/percentages; choose a t-procedure for quantitative data summarized as means
'Fail to reject H0' does NOT mean H0 is proven true — it only means there's insufficient evidence against it
A P-value is NOT the probability that H0 is true; it's the probability of data this extreme (or more) GIVEN H0 is true
Decreasing $α$ to reduce Type I error risk INCREASES the risk of Type II error (lowers power) — you can't minimize both simultaneously without changing the sample size
Confidence level describes the long-run success rate of the METHOD across many samples, not the probability that one specific interval contains the true parameter
Term
Press Enter or Space to flip the card. Left and right arrows move between cards. 1 marks it known, 2 marks it still learning.
Click or press Enter to flip · Rate yourself to track weak cards
Browse all 80 flashcards as a list

Unit 1: Exploring Data

Mean
The arithmetic average, x̄ = (Σx)/n; sensitive to outliers and skew.
Median
The middle value of ordered data; resistant to outliers and skew.
Standard Deviation
A measure of typical distance from the mean; s = √[Σ(x−x̄)²/(n−1)].
IQR (Interquartile Range)
Q3 − Q1; the spread of the middle 50% of the data, resistant to outliers.
Outlier (1.5×IQR rule)
A value below Q1 − 1.5(IQR) or above Q3 + 1.5(IQR).
Five-Number Summary
Minimum, Q1, Median, Q3, Maximum — used to construct a boxplot.
z-score
z = (x − μ)/σ; the number of standard deviations a value is from the mean.
Resistant Measure
A statistic (median, IQR) that is not strongly affected by outliers or skew.
Skewed Right
Distribution with a long right tail; mean > median.
Skewed Left
Distribution with a long left tail; mean < median.

Unit 2: Modeling Distributions

Density Curve
A smooth curve modeling a distribution's shape; total area underneath = 1.
Normal Distribution
A symmetric, bell-shaped distribution denoted N(μ, σ), fully described by its mean and SD.
Empirical (68-95-99.7) Rule
In a Normal distribution, about 68% of data fall within 1 SD, 95% within 2 SD, and 99.7% within 3 SD of the mean.
Standard Normal Distribution
The Normal distribution N(0, 1), used after standardizing with z-scores.
z-score formula
z = (x − μ)/σ; converts a value to standard deviation units from the mean.
normalcdf
Calculator function that finds the proportion/probability of area between two values in a Normal distribution.
invNorm
Calculator function that finds the value corresponding to a given percentile/cumulative proportion.
Normal Probability Plot
A graph used to assess Normality; a roughly linear plot suggests the data are approximately Normal.
Inflection Point
The point(s) on a Normal curve, at μ ± σ, where curvature changes.
Percentile
The percentage of data at or below a given value in a distribution.

Unit 3: Describing Bivariate Data

Correlation (r)
A number from -1 to 1 measuring the strength and direction of a LINEAR relationship between two quantitative variables.
Least-Squares Regression Line
The line ŷ = a + bx that minimizes the sum of squared residuals; passes through (x̄, ȳ).
Slope (b)
b = r(sy/sx); the predicted change in y for each one-unit increase in x.
Residual
residual = y − ŷ, the difference between an observed and predicted value.
Coefficient of Determination (r²)
The proportion of variation in y explained by the linear relationship with x.
Residual Plot
A scatterplot of residuals vs. x used to check whether a linear model fits well (random scatter = good fit).
Influential Point
A point, often with an extreme x-value, whose removal substantially changes the regression line.
Extrapolation
Using a regression line to predict y for x-values far outside the range of observed data; unreliable.
Lurking Variable
An unmeasured variable that influences both the explanatory and response variables, creating a misleading association.
Correlation ≠ Causation
A strong correlation between variables does not prove that one causes the other.

Unit 4: Designing Studies

Simple Random Sample (SRS)
A sample where every group of n individuals has an equal chance of being selected.
Stratified Random Sample
Population divided into homogeneous strata; an SRS is taken from each stratum.
Cluster Sample
Population divided into heterogeneous clusters; entire clusters are randomly selected and studied.
Voluntary Response Bias
Bias that occurs when a sample consists only of individuals who choose to respond, often overrepresenting strong opinions.
Nonresponse Bias
Bias that occurs when individuals selected for a sample fail to respond and differ systematically from those who do.
Confounding Variable
A variable whose effect on the response cannot be separated from the effect of the explanatory variable.
Random Assignment
Using chance to assign experimental subjects to treatment groups; enables causal conclusions.
Placebo Effect
A response to treatment caused merely by receiving some treatment, not the treatment's active ingredient.
Double-Blind Experiment
An experiment in which neither subjects nor those measuring outcomes know which treatment was assigned.
Blocking
Grouping experimental units with similar characteristics before randomly assigning treatments within each group.

Unit 5: Probability & Simulation

Complement Rule
P(not A) = 1 − P(A).
Addition Rule (General)
P(A or B) = P(A) + P(B) − P(A and B).
Mutually Exclusive (Disjoint) Events
Events that cannot both occur; P(A and B) = 0.
Independent Events
Events where the occurrence of one does not change the probability of the other: P(A|B) = P(A).
Conditional Probability
P(A|B) = P(A and B)/P(B); the probability of A given that B has occurred.
General Multiplication Rule
P(A and B) = P(A) × P(B|A); simplifies to P(A)×P(B) when independent.
Law of Large Numbers
As the number of trials increases, the observed relative frequency approaches the true probability.
Simulation
Using random numbers to imitate chance behavior and estimate a probability through repeated trials.
Tree Diagram
A diagram showing sequential probabilities; multiply along branches, add across separate paths.
Two-Way Table
A table organizing counts by two categorical variables, useful for finding joint, marginal, and conditional probabilities.

Unit 6: Random Variables

Discrete Random Variable
A variable that takes a countable set of numeric values, each occurring with a specific probability.
Valid Probability Distribution
Every P(x) ≥ 0 and the sum of all P(x) equals 1.
Expected Value E(X)
μX = Σ x·P(x), the probability-weighted long-run average of a random variable.
Variance of a Random Variable
σ²X = Σ (x−μX)²·P(x); measures average squared deviation from the mean.
Linear Transformation Rules
E(a+bX) = a+b·E(X); Var(a+bX) = b²·Var(X). Adding a constant shifts the mean but never changes the variance.
Combining Independent Random Variables
μ(X±Y) = μX±μY always; σ²(X±Y) = σ²X+σ²Y only if X and Y are independent.
Binomial Setting (BINS)
Binary outcomes, Independent trials, fixed Number of trials n, Same probability p of success each trial.
Binomial Probability Formula
P(X=k) = C(n,k)·p^k·(1−p)^(n−k); mean np, SD √(np(1−p)).
Geometric Setting (BITS)
Binary outcomes, Independent trials, constant probability p, trials continue until first Success.
Geometric Probability Formula
P(X=k) = (1−p)^(k−1)·p for k=1,2,3,…; mean 1/p, SD √((1−p)/p²).

Unit 7: Sampling Distributions

Sampling Distribution
The distribution of a statistic's values across all possible samples of a given size from a population.
Statistic vs. Parameter
A statistic describes a sample and varies from sample to sample; a parameter describes a population and is a fixed (usually unknown) number.
Unbiased Estimator
An estimator whose sampling distribution has a mean equal to the true value of the parameter it estimates.
Mean and SD of p̂
μp̂ = p; σp̂ = √(p(1−p)/n).
Mean and SD of x̄
μx̄ = μ; σx̄ = σ/√n.
Central Limit Theorem
As n increases, the sampling distribution of x̄ becomes approximately Normal, regardless of the population's shape (rule of thumb n≥30).
Large Counts Condition
np≥10 and n(1−p)≥10; ensures the sampling distribution of p̂ is approximately Normal.
10% Condition
Sample size n should be less than 10% of the population size to treat observations as approximately independent.
Standard Error
An estimate of a statistic's standard deviation using sample data when the true population parameter is unknown (e.g., s/√n).
Effect of Increasing n
Larger n decreases the variability (standard deviation) of a sampling distribution but does not affect bias.

Unit 8: Inference: Confidence & Tests

Confidence Interval
statistic ± (critical value)×(standard error); estimates a population parameter with a stated level of confidence in the method.
Confidence Level (interpretation)
The long-run proportion of intervals, built the same way from repeated random samples, that would capture the true parameter.
Margin of Error
(critical value)×(standard error); the ± amount added/subtracted from the statistic to build the interval.
Null Hypothesis (H0)
The claim of 'no effect' or 'no difference' that is assumed true unless the sample data provide convincing evidence against it.
P-value
The probability of obtaining a test statistic at least as extreme as the one observed, assuming H0 is true.
Type I Error
Rejecting a true null hypothesis; occurs with probability α (the significance level).
Type II Error
Failing to reject a false null hypothesis; occurs with probability β.
Power of a Test
1−β; the probability of correctly rejecting a false null hypothesis (detecting a real effect).
One-Sample t-Test Statistic
t = (x̄ − μ0)/(s/√n), with degrees of freedom n−1.
Chi-Square Test Statistic
χ² = Σ(observed − expected)²/expected; used with categorical count data.
Press 1–4 to answer · Enter for next

Unit 1: Exploring Data

Shape, Center, and Spread
Every distribution of quantitative data is described by its shape (symmetric, skewed, unimodal/bimodal), center (mean or median), and spread (range, IQR, or standard deviation), along with any outliers.
Measures of Center
The mean is the arithmetic average and is sensitive to outliers/skew, while the median is the middle value and resists extreme values.
Measures of Spread
Spread describes how variable the data are; range and IQR are resistant-friendly, while variance and standard deviation use every value.
Five-Number Summary, Boxplots, and Outliers
The five-number summary (min, Q1, median, Q3, max) is displayed visually with a boxplot, and the 1.5×IQR rule flags potential outliers.
z-Scores and Standardization
A z-score converts a raw value into the number of standard deviations it lies from the mean, letting you compare values from different distributions.
Comparing Distributions and Density
When comparing two or more distributions, always compare shape, center, spread, and outliers in context rather than describing them separately.
Key fact
$x̄ = (Σx)/n$ ; s $= √[Σ(x−x̄)²/(n−1)]$
Key fact
IQR $=$ Q3 − Q1; outliers lie beyond Q1 − 1.5·IQR or Q3 + 1.5·IQR
Key fact
z $=$ $(x − μ)/σ$
Key fact
Mean is pulled toward the tail in skewed data; median resists skew

Unit 2: Modeling Distributions

Density Curves
A density curve is a smooth curve that models the shape of a distribution, where the total area underneath equals 1 and area corresponds to proportion of data.
The Normal Distribution
The Normal distribution is a symmetric, bell-shaped density curve fully described by its mean $μ$ and standard deviation $σ$, following the Empirical (68-95-99.7) Rule.
Standardizing and z-Scores
Any Normal distribution can be converted to the standard Normal distribution N(0,1) by standardizing values with z $=$ $(x − μ)/σ$.
Normal Probability Calculations
Calculator functions normalcdf and invNorm let you compute probabilities/proportions and find values from percentiles for any Normal distribution.
Assessing Normality
A Normal probability plot (or comparing the data to the 68-95-99.7 rule) helps determine whether a data set is approximately Normally distributed.
Key fact
Empirical Rule: 68% within $±1σ$, 95% within $±2σ$, 99.7% within $±3σ$ of $μ$
Key fact
z $=$ $(x − μ)/σ$ ; x $= μ$ + $zσ$
Key fact
Total area under any density curve $=$ 1
Key fact
Standard Normal: N(0,1)

Unit 3: Describing Bivariate Data

Scatterplots and Association
A scatterplot displays the relationship between two quantitative variables, revealing direction, form, strength, and unusual features (outliers).
Correlation (r)
The correlation coefficient r measures the strength and direction of a linear relationship between two quantitative variables, ranging from -1 to 1.
Least-Squares Regression Line
The least-squares regression line ŷ $=$ a + bx minimizes the sum of squared residuals and is used to predict y from x.
Residuals and Residual Plots
A residual is the difference between an observed and predicted value (residual $=$ y − ŷ), and residual plots help check whether a linear model is appropriate.
r² (Coefficient of Determination)
r² is the proportion of variation in y that is explained by the linear regression on x, and is always between 0 and 1.
Outliers, Influential Points, and Transformations
Outliers and especially influential points (often high-leverage x-values) can dramatically change the regression line, and nonlinear data may require transformation.
Key fact
b $=$ r(sy/sx); line passes through ($x̄$, ȳ)
Key fact
r² $=$ r²; proportion of variation in y explained by x
Key fact
residual $=$ y − ŷ; residuals sum to 0 for LSRL
Key fact
-1 ≤ r ≤ 1; r measures strength/direction of LINEAR association only

Unit 4: Designing Studies

Sampling Methods
Good sampling methods use chance to select a subset of a population so the sample represents the whole population without bias.
Bias in Sampling
Bias occurs when a sampling method systematically favors certain outcomes, making the sample not representative of the population.
Observational Studies vs. Experiments
An observational study measures variables without imposing treatments, while an experiment deliberately imposes treatments to study cause-and-effect.
Principles of Experimental Design
Well-designed experiments use control, randomization, and replication to reduce bias and allow cause-and-effect conclusions.
Blocking and Matched Pairs
Blocking groups experimental units by a variable expected to affect the response before randomizing treatments within each block, increasing precision.
Scope of Inference
The scope of conclusions you can draw depends on whether random sampling and/or random assignment were used in a study.
Key fact
SRS: every possible sample of size n is equally likely
Key fact
Random assignment → causation; random sampling → generalizability
Key fact
3 Principles of experimental design: Control, Randomize, Replicate
Key fact
Block what you can, randomize what you cannot

Unit 5: Probability & Simulation

Basic Probability Rules
Probability quantifies the likelihood of an event, always between 0 and 1, and follows specific rules for combining events.
Independence and Mutual Exclusivity
Two events are independent if the occurrence of one does not change the probability of the other; this is a different concept from mutually exclusive events.
Conditional Probability
Conditional probability P(A|B) gives the probability of event A occurring given that event B has already occurred, restricting the sample space to B.
The General Multiplication Rule
The multiplication rule finds the probability that two events both occur: P(A and B) $=$ P(A)×P(B|A), simplifying to P(A)×P(B) when independent.
Simulation
Simulation uses random number generators (or physical devices) to imitate chance behavior, estimating probabilities that are hard to calculate directly.
Two-Way Tables and Venn Diagrams
Two-way tables organize data by two categorical variables, and Venn diagrams visually represent unions, intersections, and complements of events.
Key fact
P(A or B) $=$ P(A) + P(B) − P(A and B); if disjoint, P(A or B) $=$ P(A) + P(B)
Key fact
P(A|B) $=$ P(A and B)/P(B)
Key fact
P(A and B) $=$ P(A)×P(B|A); if independent, $=$ P(A)×P(B)
Key fact
Independent ≠ Mutually Exclusive: disjoint events with P>0 are never independent

Unit 6: Random Variables

Discrete Random Variables & Probability Distributions
A discrete random variable X takes a countable set of numeric values, each with a probability; the probability distribution lists every value with P(X $=$ x).
Expected Value (Mean) of a Random Variable
The expected value $μX$ is the long-run average outcome of a random variable if the chance process were repeated infinitely many times — a probability-weighted average, not a typical single outcome.
Variance and Standard Deviation of a Random Variable
Variance measures the average squared deviation of a random variable from its mean, quantifying spread; standard deviation is its square root, back in the original units.
Combining Random Variables
New random variables formed by adding, subtracting, or scaling other random variables follow predictable rules for their mean and variance.
Binomial Distributions
A binomial random variable counts the number of successes in a fixed number of independent trials, each with the same two possible outcomes and the same probability of success.
Geometric Distributions
A geometric random variable counts the number of trials needed to get the FIRST success in a sequence of independent trials with constant probability of success.
Key fact
E(X) $= Σ$ x·P(x); Var(X) $= Σ (x−μ)²·P(x); σX =$ √Var(X)
Key fact
Linear transform: E(a+bX) $=$ a+b·E(X); Var(a+bX) $=$ b²·Var(X)
Key fact
Binomial: $μ =$ np, $σ =$ √(np(1−p)); $P(X=k) = C(n,k)p^k(1−p)^(n−k)$
Key fact
Geometric: $μ =$ 1/p, $σ =$ √((1−p)/p²); $P(X=k) = (1−p)^(k−1)p$

Unit 7: Sampling Distributions

Sampling Distributions: Core Idea
A sampling distribution describes the distribution of a statistic (like a sample mean or sample proportion) computed from every possible sample of a given size — it tells us how much a statistic varies from sample to sample.
Sampling Distribution of a Sample Proportion
When repeatedly sampling from a population with true proportion p, the sample proportion p̂ has a predictable mean, standard deviation, and (under conditions) approximately Normal shape.
Sampling Distribution of a Sample Mean
When repeatedly sampling from a population with mean $μ$ and standard deviation $σ$, the sample mean $x̄$ has its own sampling distribution with predictable mean and standard deviation.
The Central Limit Theorem (CLT)
The Central Limit Theorem states that as sample size n increases, the sampling distribution of the sample mean x̄ approaches a Normal distribution, regardless of the shape of the population distribution.
Standard Error and the Effect of Sample Size
Standard error is the estimated standard deviation of a sampling distribution when the population parameter is unknown; it shrinks as sample size grows, following an inverse square-root relationship.
Key fact
$μp̂ =$ p; $σp̂ =$ √(p(1−p)/n)
Key fact
$μx̄ = μ; σx̄ = σ/√n$
Key fact
CLT: shape of x̄'s sampling distribution → approximately Normal as n grows (rule of thumb n≥30)
Key fact
Large Counts (proportions): np≥10 and n(1−p)≥10; 10% condition needed for independence in both cases

Unit 8: Inference: Confidence & Tests

Confidence Intervals: Logic and Structure
A confidence interval uses a sample statistic plus a margin of error to estimate an unknown population parameter, with a stated level of confidence in the method used to produce it.
Confidence Interval for a Proportion and a Mean
The one-proportion z-interval and one-sample t-interval are the two core confidence interval procedures for estimating a population proportion or population mean.
Hypothesis Testing: Logic and Structure
A hypothesis test uses sample evidence to assess a claim about a population, weighing the null hypothesis (no effect/no difference) against an alternative hypothesis.
Type I and Type II Errors, and Power
Every hypothesis test risks two types of mistakes: rejecting a true null hypothesis (Type I) or failing to reject a false null hypothesis (Type II); power measures the test's ability to correctly detect a real effect.
t-Tests, Chi-Square Tests, and Choosing a Procedure
Different inference procedures fit different data types: t-tests compare means, chi-square tests analyze categorical count data, and each has its own test statistic and distribution.
Key fact
CI form: statistic ± (critical value)×(SE); proportion: $p̂$ ± $z*√(p̂(1−p̂)/n)$; mean: $x̄$ ± t*(s/√n), $df=n−1$
Key fact
Test statistic: (statistic − hypothesized value)/(standard error); one-sample t: $(x̄−μ0)/(s/√n)$
Key fact
$χ² = Σ(observed−expected)²/expected$
Key fact
P(Type I error) $=α$; P(Type II error) $=β$; Power $=1−β$
Common mistakes for each unit — read the mistake, then make sure you know why it's wrong.

Unit 1: Exploring Data

Watch out
‘Skewed left’ describes the direction of the longer tail, not where most data pile up
Watch out
Standard deviation can never be negative; a computed negative SD means an arithmetic error
Watch out
Adding a constant to all data shifts the mean but leaves the standard deviation unchanged — only multiplying changes spread
Watch out
Comparing distributions requires comparative language in context, not two separate one-group descriptions

Unit 2: Modeling Distributions

Watch out
The Empirical Rule only applies to (approximately) Normal distributions, not all distributions
Watch out
A Normal probability plot with a clear curve — not a straight line — signals the data are NOT Normal
Watch out
Area under a density curve represents proportion, not probability of a single exact point (P(X $=$ x) $=$ 0 for continuous variables)
Watch out
z-scores can be negative; a negative z simply means the value is below the mean, not an error

Unit 3: Describing Bivariate Data

Watch out
Correlation does NOT imply causation — a strong r only shows association
Watch out
A high r² does not fix a nonlinear relationship; always check the residual plot first
Watch out
Extrapolation beyond the observed x-range is unreliable even with a high r²
Watch out
Correlation r is not resistant — a single outlier can drastically change its value

Unit 4: Designing Studies

Watch out
A large sample size cannot fix bias from a poor sampling method (e.g., a huge voluntary response sample is still biased)
Watch out
Random sampling and random assignment are different: sampling allows generalization, assignment allows causal claims
Watch out
Observational studies, no matter how large, cannot establish causation because of possible confounding variables
Watch out
Blocking is done BEFORE random assignment; it's not the same as simply comparing groups after the fact

Unit 5: Probability & Simulation

Watch out
Mutually exclusive events are NOT independent (if both have positive probability) — a very common mix-up
Watch out
P(A|B) ≠ P(B|A) in general — order and conditioning direction matter
Watch out
'At least one' problems are usually easiest using the complement rule: P(at least one) $=$ 1 − P(none)
Watch out
Sampling without replacement changes probabilities from draw to draw, creating dependence, even if it 'feels' independent for large populations (approximately independent only when population is much larger than sample)

Unit 6: Random Variables

Watch out
Variances add for independent RVs even when you SUBTRACT the variables: Var(X−Y) $=$ Var(X)+Var(Y), not Var(X)−Var(Y)
Watch out
Standard deviations never add or subtract directly — always combine variances first, then take the square root at the end
Watch out
A binomial variable requires a FIXED number of trials n; if the question asks 'how many trials until the first success,' that's geometric, not binomial
Watch out
Geometric random variables start at X $=$ 1 (you need at least one trial), never at X $=$ 0

Unit 7: Sampling Distributions

Watch out
Increasing sample size reduces variability ($σ/√n$) but never removes bias from a flawed sampling method
Watch out
$σx̄ = σ/√n$, NOT $σ/n$ — a common algebra slip that badly underestimates spread
Watch out
The CLT describes the SHAPE of the sampling distribution of x̄, not the shape of the original population data
Watch out
To halve the standard error, you must quadruple n, not just double it, because of the square root in $σ/√n$

Unit 8: Inference: Confidence & Tests

Watch out
'Fail to reject H0' does NOT mean H0 is proven true — it only means there's insufficient evidence against it
Watch out
A P-value is NOT the probability that H0 is true; it's the probability of data this extreme (or more) GIVEN H0 is true
Watch out
Decreasing $α$ to reduce Type I error risk INCREASES the risk of Type II error (lowers power) — you can't minimize both simultaneously without changing the sample size
Watch out
Confidence level describes the long-run success rate of the METHOD across many samples, not the probability that one specific interval contains the true parameter