Subject-Specific Teaching
Teaching Scientific Literacy: Questions, Evidence, Models, and Uncertainty

Scientific literacy is the ability to ask answerable questions, interpret how evidence was produced, use and critique models, describe uncertainty, and form conclusions that do not reach beyond the evidence. It is not a store of isolated facts or a habit of accepting every statement labeled scientific. The National Academies framework for K-12 science education connects scientific knowledge with practices such as analyzing data, developing models, constructing explanations, and arguing from evidence.
Teach a repeatable evidence routine
Give students a short sequence they can apply to a graph, investigation, news item, or research summary. Model one step at a time before asking for a complete evaluation.
| Part | Questions to ask | Evidence in a student response |
|---|---|---|
| Question and claim | What was asked? What, exactly, is being claimed? Is the claim descriptive, comparative, causal, or predictive? | Restates a bounded claim without strengthening it. |
| Data and measurement | What was measured, with what unit or category, by whom, when, and how? What was the sample? | Names the relevant values and the measurement process. |
| Model and omissions | What representation or model is being used? What does it include, simplify, or leave out? | Uses the model for its stated purpose and names one consequential omission. |
| Uncertainty and confidence | What variation, measurement uncertainty, missing data, or sampling limit matters? How confident should the conclusion be? | Calibrates language to the evidence instead of treating uncertainty as ignorance. |
| Cause and alternatives | Does the design support causation, or only an association? What other explanation fits the pattern? | Separates correlation from causation and names a plausible alternative. |
| Source and review | Who produced the work? What relevant expertise, methods, funding, conflicts, publication venue, and review process are reported? | Explains why a source feature affects confidence; it does not score credibility from a logo alone. |
| Conclusion | What conclusion is warranted now? What new evidence could change it? | States a restrained answer tied to cited evidence and limits. |
Claims, data, evidence, and models are different
A claim is a proposed answer. Data are recorded observations or measurements. Data become evidence when a reasoned argument connects selected data to a claim. A model is a purposeful representation of a system or process. A diagram, equation, physical replica, simulation, or verbal explanation can be a model.
Ask students to name the model’s purpose before judging it. A map that omits individual trees may still be useful for showing land-cover zones; it would be a poor tool for counting trees. All models leave things out
is only the start. Students should identify an omission and explain whether it matters for the question.
Uncertainty is information, not a flaw
Measurements vary because instruments have limits, conditions change, samples differ, and people make classification choices. A range, margin of error, confidence interval, or repeated-measure spread describes a particular source of uncertainty; the terms are not interchangeable. The NIST guide to measurement uncertainty provides a technical reference for teachers selecting age-appropriate language.
Confidence should match the design and the size and consistency of the pattern. The evidence supports this explanation under these conditions
is often more accurate than either this proves it
or we cannot know anything.
Statistical confidence also does not account automatically for biased sampling, weak measures, omitted variables, or errors in the underlying data.
Correlation, causation, authorship, and review
When two variables change together, the pattern is a correlation. A causal conclusion needs a design that addresses time order, comparison conditions, alternative causes, and the mechanism connecting cause and effect. Even a strong correlation can arise from a third variable, reverse direction, selective measurement, or chance.
Source review is more than sorting websites into trusted and untrusted groups. Check the named authors and their relevant expertise; the original data and method; the publication date; funding and conflicts; corrections; and whether the claim appears in the source or only in a summary. Authorship assigns contribution and responsibility. Peer review is expert scrutiny before publication, but it does not certify that a finding is error-free. The U.S. Office of Research Integrity overview of reporting and review can support this distinction.
Scientific consensus is the broad convergence of relevant experts after multiple lines of evidence and continued scrutiny. It is not a vote on one paper, and it does not require unanimity. A dissenting expert is not, by itself, evidence that the field is evenly divided; consensus can still change when better evidence or explanations emerge.
Complete low-risk data example: paper helicopter drop time
This practice case uses a constructed data set for teaching; it does not report measurements from a real investigation or student outcomes. The question is: Under the described conditions, did the longer-wing paper helicopter take longer to reach the floor?
Source and method: Mentor Teaching created the complete practice data below to accompany this worksheet. The hypothetical method states that one adult dropped each paper helicopter five times from the same marked height indoors and timed each drop to the nearest 0.1 second with the same stopwatch. Helicopter S had 8 cm wings; Helicopter L had 12 cm wings. Paper, body dimensions, paper clips, dropper, timer, and location were held constant.
| Trial | Helicopter S: 8 cm wings | Helicopter L: 12 cm wings | Difference, L minus S |
|---|---|---|---|
| 1 | 1.2 s | 1.6 s | 0.4 s |
| 2 | 1.3 s | 1.5 s | 0.2 s |
| 3 | 1.1 s | 1.7 s | 0.6 s |
| 4 | 1.4 s | 1.6 s | 0.2 s |
| 5 | 1.2 s | 1.8 s | 0.6 s |
| Mean | 1.24 s | 1.64 s | 0.40 s |
Limits and alternatives: There were only five drops per design, reaction time affects stopwatch readings, trials were not described as randomized or blinded, and only two wing lengths and one drop height were tested. Small differences in folds, release angle, air movement, or rotation could also affect time. The data do not show that every longer-wing design falls more slowly or that wing length alone caused the difference.
Sample response and revision
First response: Long wings make paper helicopters fall more slowly because every long-wing time was higher.
Feedback: Cite the size of the pattern, limit the conclusion to the tested designs and conditions, and include measurement limits and an alternative explanation.
Revised response: In these five indoor drops per design, Helicopter L had a mean time of 1.64 seconds, 0.40 seconds longer than Helicopter S. All five recorded L times were higher than the paired S times, so the measurements support the claim that this longer-wing design took longer under the tested conditions. The small sample and 0.1-second hand timing limit confidence. Fold differences, release angle, or air movement could partly explain the pattern. More randomized drops, another timer, and several wing lengths would test whether wing length is the main cause.
Copyable scientific-literacy worksheet
Scientific claim or question:
1. Claim type (descriptive, comparative, causal, or predictive):
2. Exact claim in my own words:
3. Data source and original source link or citation:
4. Who or what was sampled, and how many:
5. What was measured or classified, how, and in what units:
6. Relevant data values or patterns:
7. How these data do or do not support the claim:
8. Model or representation used:
9. What the model includes:
10. What the model omits, and why that omission matters:
11. Uncertainty, variation, or missing data:
12. Correlation or causation? Why:
13. One plausible alternative explanation:
14. Author, relevant expertise, funding, and conflicts reported:
15. Publication and review status; corrections or later evidence:
16. My restrained conclusion:
17. Evidence that could change my conclusion:Reasoning rubric
| Criterion | 4: Specific and warranted | 3: Mostly warranted | 2: Partial | 1: Unsupported |
|---|---|---|---|---|
| Claim and evidence | States a bounded claim, cites the most relevant values, and explains the connection. | States a suitable claim and cites relevant values; connection has a minor gap. | Gives a claim or data but makes only a general connection. | Misstates the claim, omits evidence, or cites unrelated information. |
| Measurement and data | Accurately describes sample, measure, units, method, and consequential data limits. | Describes most measurement features and one relevant limit. | Mentions the method or a limit without explaining its effect. | Treats recorded values as complete or error-free without examining how they were produced. |
| Model and alternatives | Explains the model’s purpose, a consequential omission, and a plausible alternative explanation. | Explains two of those three features accurately. | Names a model limit or alternative without connecting it to the claim. | Treats the model as the system itself or ignores alternatives. |
| Uncertainty and conclusion | Calibrates confidence, distinguishes correlation from causation, and states what evidence could change the conclusion. | Uses restrained language and addresses uncertainty or causation correctly. | Uses a caveat but overstates or understates what follows from it. | Claims proof, dismisses all evidence, or gives no conclusion. |
Common reasoning errors
| Error | Why it fails | Correction prompt |
|---|---|---|
| Cherry-picking | Selects only values or studies that support the preferred answer. | What does the complete set show, including contrary or missing evidence? |
| Correlation as causation | Does not rule out reverse direction, third variables, or chance. | What design feature supports a causal inference, and what alternatives remain? |
| Authority alone | A credential or institution cannot substitute for relevant methods and evidence. | What data and method support this particular claim? |
| Peer review as proof | Review can identify problems but cannot guarantee correctness. | Has the result been replicated, corrected, or challenged by later evidence? |
| Uncertainty as ignorance | A known range of uncertainty can still support a decision. | What conclusion remains warranted within the stated range? |
| Consensus as unanimity | Broad expert agreement can coexist with unresolved details and dissent. | Which conclusions converge across independent evidence, and which remain open? |
A restrained conclusion
End each task with a conclusion proportional to the evidence: state what the data support, the conditions to which the conclusion applies, a material uncertainty or alternative, and what further evidence would matter. Score the reasoning shown in that task. Do not use one response to infer a student’s general judgment, identity, or future decision-making.
Sources
- National Research Council, A Framework for K-12 Science Education: practices, models, evidence, and cause-and-effect reasoning.
- National Institute of Standards and Technology, Guidelines for Evaluating and Expressing the Uncertainty of NIST Measurement Results: measurement and uncertainty terminology.
- U.S. Office of Research Integrity, Reporting and Reviewing Research: authorship, publication, and peer review.
- OECD, PISA 2025 Science Framework: evaluating scientific inquiry and using scientific information for decisions.
What changed: This revision adds a complete data case, a copyable worksheet, a reasoning rubric, and clearer checks for models, uncertainty, sources, and causal claims.