Large Counts Condition: AP Statistics Examples & Checks

Learn the large counts condition for AP Statistics with correct interval and test checks, worked one- and two-proportion examples, and current exam context.
Teal successes and gold failures beneath a normal curve illustrating the large counts condition

The large counts condition is a screening check for normal-based inference about proportions. The number 10 appears in two different ways: a confidence interval checks the successes and failures actually observed, whereas a one-proportion hypothesis test checks the counts expected if its null hypothesis were true. Mixing those quantities is one of the easiest ways to write an incorrect AP Statistics solution. This guide explains the distinction, works through one- and two-sample examples, and shows when a different method is preferable.

A successful check does not prove that the sample was random, independent, or representative. It says only that a particular normal approximation is reasonably supported by the counts. Always identify the inferential procedure first, then check its matching counts, then examine the study design, and only then calculate and interpret a result.

What the condition is trying to approximate

Suppose a school asks whether each student would enroll in an optional workshop. Label a yes response a success and a no response a failure. If a random sample includes n students and x say yes, the sample proportion, written p-hat, is x divided by n. The population proportion p is the unknown fraction of all eligible students who would say yes under the same conditions. Repeating the sampling process would produce different values of p-hat. The distribution of those hypothetical values is the sampling distribution.

The data themselves are not bell-shaped; each response belongs to one of two categories. The question is whether the sampling distribution of the proportion is sufficiently close to a normal curve for a z procedure to provide a useful approximation. If one category is rare, the underlying binomial counts may be lopsided and the normal curve may misrepresent a tail area. Requiring enough successes and enough failures helps screen for this problem. A large total sample alone is not enough.

Consider a proposed model with 100 trials and a success probability of 0.02. It predicts only two successes and 98 failures on average. The success side is much too sparse for the familiar introductory large-count rule, despite the three-digit sample size. By contrast, a model with 100 trials and probability 0.50 predicts 50 of each outcome. The approximation is generally more plausible because neither category hugs its boundary of zero.

The threshold of 10 is a rule of thumb rather than a law of nature. An approximation does not suddenly become perfect when a count changes from 9 to 10. Accuracy depends on the probability, sample size, method, and which tail is important. In introductory AP-style problems, however, an explicit at-least-10 condition supplies a transparent decision rule. Apply the threshold stated by the problem or course and avoid rounding 9.6 upward to claim it was met.

First decide: interval or test?

A one-proportion confidence interval uses observed counts

A confidence interval estimates p from sample data. The usual one-proportion z interval checks x observed successes and n minus x observed failures. Both should be at least 10 under the common classroom rule. Because p-hat equals x/n, the same check can be written n times p-hat at least 10 and n times (1 minus p-hat) at least 10. These products simply recover the observed counts. The true p is unknown; inserting an invented population proportion into the interval check would miss the purpose of estimating it.

In an exam response, show the arithmetic and say what the categories mean: “The random sample includes 60 students who would enroll and 140 who would not, so both observed counts exceed 10.” Writing only “large counts satisfied” hides the evidence and makes it hard to see whether the right version of the condition was used. It is also essential to name the population to which the proposed interval applies.

A one-proportion z test uses null-expected counts

A hypothesis test asks what sample results would look like if a stated null hypothesis were true. Let the null assert p = p0. For the usual one-proportion z test, the condition checks n times p0 expected successes and n times (1 minus p0) expected failures. These need not equal the observed x and n minus x. The null-based standard error and reference distribution use p0, so the approximation should be checked under that same model.

For example, suppose n = 120, the null says p0 = 0.25, and 42 sample members are successes. The observed counts are 42 and 78, but the test’s expected counts under the null are 30 and 90. Both expected counts exceed 10. Substituting p-hat = 0.35 into this test check would evaluate a different model, not the one whose tail probability becomes the p-value. The memorable distinction is interval: observed; test: expected under the null.

Why the distinction affects the answer

Imagine 80 observations with four successes. For a standard z confidence interval, the observed success count of four fails the condition. If the null for a separate test is p0 = 0.05, the expected count is also four and that test check fails. If a different test uses p0 = 0.25, however, the null predicts 20 successes and 60 failures, so its expected-count check passes even though only four successes were observed. The interval and tests address distinct inferential questions; their checks cannot be merged into one universal pair of numbers.

This contrast does not imply that every null-based test will be an excellent choice when the observed result is extreme. A very extreme statistic may call for additional scrutiny of numerical approximation and assumptions. It means only that the conventional test condition is evaluated under the null model. If a course or question specifies an exact binomial procedure, follow that procedure rather than forcing the z method.

Success and failure are labels, not value judgments

A “success” is whichever outcome the analyst has chosen to count. In a manufacturing study, a defective item may be called a success because defects are the event of interest. In a health survey, a reported symptom may be the success. Changing the label reverses the numerical proportion but does not change the need to examine both categories. Define the outcome before calculating so that a reader can tell whether x means yes responses, defects, or something else.

Each observation should belong to one of two clearly defined categories and contribute once. If a respondent can select several options, the simple yes/no model needs a carefully specified event, such as “selected option A.” If people answer repeatedly, or outcomes cluster within classrooms, independence may fail even when the two counts exceed 10. The large-count check addresses the approximation’s shape; it does not validate how the observations were collected.

Worked example: one-proportion confidence interval

A school takes a random sample of 200 eligible students. Sixty say they would join a voluntary online review class. Let p mean the proportion of all eligible students at that school who would say yes under the survey conditions. The point estimate is p-hat = 60/200 = 0.30. Because this is an interval, the large-count check uses the 60 observed yes responses and the 140 observed no responses. Each is at least 10, so this part of the normal-approximation assessment is satisfied.

Now inspect the design. Were the 200 selected at random from a defined list of eligible students, or did the teacher share a link and wait for volunteers? Only the former has a defensible route to population inference under the usual textbook framework. If the students were sampled without replacement from a finite school, the familiar 10% guideline asks that the sample be no more than about one tenth of the population when we use the simple independence approximation. A school with at least 2,000 eligible students would meet that particular numerical guideline for a sample of 200; a smaller school would need more care. The large-count calculation cannot settle either issue.

Under the standard one-proportion z interval, the estimated standard error is the square root of [0.30 times 0.70 divided by 200], approximately 0.0324. For a 95% interval the usual z multiplier is 1.96, giving a margin near 0.0635. The interval is therefore about 0.2365 to 0.3635, or 23.65% to 36.35%. The arithmetic is meaningful only after we establish that the sampling process and approximation reasonably match the method.

A contextual interpretation is: “Using this sampling procedure, we are 95% confident that the proportion of eligible students at this school who would choose the online review class lies between about 23.7% and 36.4%.” The population proportion is treated as fixed, not as an object randomly moving around this particular interval. The confidence level describes the long-run performance of the method across repeated random samples. Nor would a narrow numerical interval correct volunteer bias if the selection process were flawed.

Worked example: one-proportion hypothesis test

A workshop organizer claims that 25% of eligible students choose the workshop. A random sample of 120 finds 42 who choose it. Let p be the population proportion of eligible students who choose the workshop. To investigate whether it exceeds the claim, state H0: p = 0.25 and Ha: p > 0.25 before computing a tail probability. The observed proportion is 42/120 = 0.35. The direction of the alternative comes from the research question, not from a decision made after seeing which direction gives a smaller p-value.

The null model expects 120 times 0.25 = 30 successes and 120 times 0.75 = 90 failures. Both are above 10. Those expected counts support the ordinary one-proportion z test’s normal approximation under H0. The observed 42 successes and 78 failures describe the sample, but they are not the pair used in the null-based condition check. Again, verify how students were sampled, approximate independence, and any finite-population guideline relevant to the design.

Under H0, the standard error is the square root of [0.25 times 0.75 divided by 120], about 0.0395. The test statistic is (0.35 minus 0.25) divided by 0.0395, approximately 2.53. The right-tail probability of a standard normal value at least this large is about 0.006. At a prespecified significance level of 0.05, that is evidence against the null in favor of a population workshop-choice proportion above 25%, subject to the design assumptions.

The p-value is not the probability that the null is true. It is the probability, calculated under the null model, of obtaining a statistic at least as extreme in the specified direction. Statistical significance also does not by itself tell us that a roughly ten-percentage-point sample difference is educationally important. That judgment requires context, potential costs, and the consequences of being wrong. If the original question instead asked whether the population proportion differs from 0.25 in either direction, the alternative would be two-sided and the p-value would change.

Why both categories, rather than just the total n?

The binomial count cannot fall below zero or rise above n. When p is near one half, outcomes around the center are more balanced and a bell-shaped approximation is often plausible with a moderate sample. When p is near zero or one, the count is pushed toward a boundary and the distribution becomes skewed. A symmetric normal curve extends beyond those boundaries and can assign implausible area to negative counts or proportions greater than one. Enough information on both sides limits that mismatch.

Suppose 500 items are inspected and an assumed defect rate is 0.01. The expected number of defects is only five, while expected nondefects number 495. The sample is large in ordinary language but the rare event is still sparse. By contrast, if the expected defect rate were 0.10, 500 items would yield 50 expected defects and 450 expected nondefects. The difference comes from the event probability, not merely n. This is why replacing the condition with an unrelated slogan such as “sample size at least 30” can be misleading.

Ten is not a guarantee of exact coverage for every interval or perfect Type I error control for every test. Even with both counts at least 10, the simple Wald confidence interval can have performance limitations, particularly near the ends of the proportion scale. A score-based interval or exact method may be preferable in more demanding analysis. Conversely, a count below 10 does not make the data worthless: it means the stated normal-based method requires reconsideration, while descriptive results remain informative.

Comparing two independent proportions

Many questions compare a proportion in one group with a proportion in another. Suppose group A contains n1 observations and x1 successes, while group B contains n2 observations and x2 successes. A standard z confidence interval for p1 minus p2 requires the observed successes and failures in each group to be sufficiently numerous: x1, n1 minus x1, x2, and n2 minus x2. Inspecting only the total number of successes across both groups could conceal a sparse category in one group.

For example, 60 of 150 students in group A choose a review workshop; 39 of 130 in group B choose it. The four observed counts are 60 and 90 for A, and 39 and 91 for B. All exceed 10. The estimated proportions are 0.40 and 0.30, so the estimated difference p1 minus p2 is 0.10. The groups must also be independent and recruited in a way appropriate to the target populations. A large-count pass does not turn two self-selected groups into a randomized comparison.

For a conventional 95% confidence interval, use the unpooled estimated standard error: the square root of [0.40 times 0.60 divided by 150, plus 0.30 times 0.70 divided by 130]. This is approximately 0.0567. Multiplying by 1.96 gives a margin near 0.111, and the approximate interval runs from negative 0.011 to positive 0.211. It includes zero, so this interval does not clearly separate the population proportions at the 95% confidence level. That does not prove equality; it shows that differences on either side of zero remain plausible under this procedure.

A standard test of H0: p1 = p2 often uses a pooled estimate because the null assumes a common population proportion. In this example the pooled rate is (60 + 39)/(150 + 130), or 99/280, approximately 0.354. The null-based expected counts are calculated separately for each group with that common rate; they are all comfortably above 10 here. Do not automatically carry over the observed-count interval check as if it were the pooled-null test check. The method and its reference model determine which quantities matter. Penn State’s two-sample inference lesson distinguishes unpooled interval estimation from pooled-null testing.

If the groups were paired rather than independent—for instance, the same students were surveyed before and after an intervention—the preceding two-independent-proportion procedure would not match the data structure. The difference within each matched pair is the relevant unit. Likewise, students nested in the same classrooms may have correlated responses. Large counts cannot solve a design mismatch. Specify whether groups are independent, paired, randomized, or observational before using an off-the-shelf formula.

Randomness, independence, and scope of inference

Random selection is not the same as random assignment

A random sample helps justify generalization to the sampled population because eligible members had a defined chance of selection. A voluntary website poll, a single convenient classroom, or a social-media survey may overrepresent certain viewpoints. Even 10,000 responses on each side cannot by themselves remove selection bias. You may report the observed fraction among respondents, but inferring the fraction in a much wider population requires a defensible sampling design or a more sophisticated model with explicit assumptions.

Random assignment in an experiment has another role: it helps compare treatment groups by distributing other factors across them. An experiment with randomly assigned volunteers can support a causal comparison within the studied population, but it does not automatically make those volunteers a random sample of every student elsewhere. A survey with a random sample can support population description yet not necessarily a causal claim about a program. Distinguishing selection and assignment prevents an apparently correct z calculation from carrying an unjustified conclusion.

Independence and finite populations

When observations are independent, knowing one outcome does not change the probability model for another. Sampling without replacement creates some dependence: after one person is drawn, the pool changes. Introductory courses commonly use a sample no larger than about 10% of the population as a practical guideline for treating such dependence as small in standard calculations. It is a separate condition from large counts. A sample of 200 from 1,000 students could have 100 successes and 100 failures, yet fail the usual 10% numerical guideline.

Other forms of dependence cannot be checked with that rule. Siblings in one household, repeated responses by one person, students sharing a classroom, and measurements taken from a clustered design may be correlated even if n is tiny relative to a national population. If an exercise states that a random sample was obtained, use that information. In real research, describe the recruitment and any clustering rather than assuming every row is independent merely because the data are in a spreadsheet.

What a successful condition check actually licenses

The large counts condition supports a normal approximation for a specified proportion procedure. It does not confirm that the measure is valid, that missing responses are harmless, that a causal effect exists, or that the study generalizes outside its population. Think of inference as a chain: define the question and parameter, evaluate the design, check the model conditions, compute the statistic, and make a conclusion no broader than the weakest link allows. A calculation can be arithmetically flawless while its real-world interpretation remains unwarranted.

When the large-count check fails

First say which count fails and by how much. “Only eight observed defects were found among 150 products, so the observed-success side of a usual one-proportion z interval is below 10” is precise. “Sample too small” is not: perhaps n is large but the event is rare. Do not quietly apply the normal z formula and hope that rounding hides the issue. Nor should you pronounce all analysis impossible. The raw counts and sample proportion can still be reported descriptively.

For a one-sample test with a sparse null-expected count, an exact binomial test can calculate probabilities from the binomial model rather than using the usual normal approximation. For an interval, a Wilson score interval or an exact binomial interval may be considered instead of the elementary Wald z interval. They are not interchangeable, and “exact” does not mean shortest or universally best. Follow the method specified by the course or analysis plan, and describe why an alternative was chosen. For a small two-by-two comparison, Fisher’s exact test is one possible test of association, not a universal fix for every one-sample interval problem.

Planning a sample when the event is rare

If a plausible planning value p is known, the expected-success criterion requires approximately n times p at least 10, while the expected-failure criterion requires n times (1 minus p) at least 10. Both inequalities must hold. Solving them gives a minimum n at least as large as the greater of 10/p and 10/(1 minus p). For an anticipated event rate of 0.02, the success side requires about 500 observations. At that sample size the expected failures number 490, so the rare category controls the requirement.

That calculation is only a screening step for an approximation. It is not a complete sample-size plan. A study seeking a narrow interval, sufficient power to detect a small effect, or protection against nonresponse may need many more people. The planning p can also be wrong. If 500 people are sampled but the event occurs in only 1%, an interval based on the realized sample has around five observed successes and does not meet the usual observed-count rule. Planning expectations cannot guarantee a future data set.

Increasing n reduces standard error roughly with the square root of sample size when other aspects of the design are comparable. Doubling n does not halve uncertainty. Recruiting more participants through a biased channel is also not a statistical repair for the sampling design. Before collecting data, define the outcome, target population, desired precision or power, recruitment mechanism, possible clustering, and likely nonresponse. Those choices matter as much as a numerical count threshold.

Write a strong AP Statistics explanation

Begin by naming the procedure and parameter in context. For an interval, state the number who meet the defined outcome and the number who do not. For a test, write the null and alternative, then calculate the two expected counts from the null proportion. For two independent samples, show the counts for both groups, and clarify whether the task is an interval or a pooled-null test. Avoid a bare formula with no interpretation of success or population.

A concise interval justification might read: “From a random sample of 200 eligible students, 60 would choose the online class and 140 would not. Both observed counts exceed 10, so the large-count check for a one-proportion z interval is met. If the eligible population contains at least 2,000 students, the sample also meets the usual 10% guideline for sampling without replacement.” Notice that the last sentence is conditional; a writer should not invent a school enrollment if it is not supplied.

A test justification might read: “Under H0: p = 0.25, a sample of 120 has 30 expected successes and 90 expected failures. Both are at least 10, supporting a normal approximation for the one-proportion z test, assuming the sampling and independence conditions also hold.” This wording identifies the model from which the expectations come. It avoids incorrectly using the 42 observed successes in the null check.

Finish by interpreting the result in ordinary language. An interval estimates a population proportion or difference between proportions, not the probability that a fixed parameter is randomly inside one particular interval. A p-value assumes the null model and measures how unusual a result at least as extreme would be, not how likely the null itself is. State whether a test is one- or two-sided, compare with the chosen significance level, and distinguish statistical evidence from practical importance. Clear writing often reveals conceptual errors before any arithmetic is checked.

Five practice checks

1. Sparse observed success in an interval

Among 150 randomly selected products, eight are defective. The observed counts are eight defects and 142 nondefects. The success side is below 10, so the conventional one-proportion z interval fails its large-count check despite n = 150. Describe the sample rate of about 5.3%, but consider an interval method suited to sparse counts if an inferential interval is needed. More failures cannot compensate for too few successes.

2. Moderate null in a test

A study tests H0: p = 0.40 with n = 50, and the observed sample has 12 successes. The z test’s null-expected counts are 20 and 30, both above 10. The observed 12 and 38 are data used to calculate p-hat and the statistic, but not the pair used for this null-model screening check. Whether the test supports population inference still depends on the selection mechanism and approximate independence.

3. Rare null in a test

A study tests H0: p = 0.03 with n = 200. It expects six successes and 194 failures under H0. Six is below 10, so the usual z test fails its null-based large-count check. Even if the sample unexpectedly has 20 successes, the null reference model is still the one being checked. An exact binomial test may be useful if its other assumptions fit the study.

4. A two-group interval

Group A has 11 successes in 30 observations. Group B has nine successes in 40. The observed four counts are 11, 19, nine, and 31. The nine in group B fails the at-least-10 check for a standard two-proportion z interval. Combining the groups to display 20 total successes would hide the sparse category in the second sample. Examine each sample on its own.

5. Lots of counts, poor selection

A website poll receives 10,000 yes votes and 8,000 no votes from visitors who chose to participate. Both counts are large, but the voluntary response mechanism does not establish representation of a wider student or community population. A safe descriptive statement is that 10,000 of the 18,000 respondents, about 55.6%, voted yes. Population generalization needs further assumptions or a better sample design.

Common mistakes in one place

Check both categories, not only successes. Do not substitute a blanket n greater than 30 rule for proportion-specific counts. Use p-hat or observed x for an interval and p0 for a one-proportion z test. Avoid calling the threshold an exact guarantee of coverage or error rates. Do not use the same calculation blindly for paired and independent groups. A correct condition statement includes numbers, the category definitions, and the method. Above all, do not let a successful numerical check overshadow nonrandom selection, dependence, or a conclusion broader than the study permits.

Where it fits in current AP Statistics

College Board’s AP Statistics revision takes effect for the 2026–27 school year. The current course overview organizes the subject into five units; inference for categorical data and proportions belongs to Unit 3, whose multiple-choice weighting is listed as 15%–25%. Older resources may call this area Unit 6 because they follow the previous nine-unit structure. The unit label changed, but the statistical distinction between observed counts for an interval and null-expected counts for a z test remains important. Check the current College Board course overview and revision summary for the framework used by your teacher.

For the 2027 exam, College Board describes a fully digital AP Statistics exam in Bluebook with four free-response questions. An inference response benefits from a named parameter, correct conditions, a transparent calculation, and a contextual conclusion. Memorizing “both at least 10” without deciding whether counts should be observed or null-expected is fragile preparation. If you are choosing an AP course rather than solving a proportion question, Sly Academy’s updated AP subjects guide explains the broader course options. Your teacher’s current materials remain the authority on notation and methods expected in class.

Calculate carefully, then interpret

A calculator is helpful only after the method is chosen. Keep the entire product and division inside the square root when evaluating a standard error. In the interval example, calculate 0.30 times 0.70 divided by 200 before taking the square root; dropping the division by n produces a grossly incorrect margin. Sly Academy’s scientific calculator can check arithmetic while studying, but it does not choose the formula or replace written context. Follow College Board’s current exam calculator policy when preparing for test day.

Keep extra digits in intermediate work, round at the end, and identify whether an endpoint is a decimal proportion or percentage. The interval endpoint 0.2365 means 23.65%, not 0.2365%. For a test, distinguish a one-sided from a two-sided tail and state the significance level specified in the question. These habits make a condition check part of sound statistical reasoning, rather than a memorized line detached from its purpose.

Sources and further study

For the formal one-proportion distinctions see Penn State STAT 200 on one-sample inference. For two independent proportions, pooled tests, and unpooled intervals see its two-sample lesson. OpenStax’s binomial-distribution chapter provides background on the binary count model. For exam-format details consult the current College Board exam page. These references support the method choices; they do not replace checking the exact conditions supplied in an individual exercise.

More Sly academy Content

Calculate Your AP Score
Support Us