Collecting enough data for the IB Biology IA means gathering sufficient relevant evidence to identify a pattern, measure variation, perform appropriate processing, and support a valid conclusion. The IB does not prescribe one universal minimum number of trials or independent-variable levels. However, for many school laboratory investigations, five or more levels with about five independent repeats per level is a sensible planning benchmark, not an official rule.
The appropriate quantity depends on the design, biological variability, expected effect, available equipment, and statistical analysis. A large table does not automatically constitute strong data. Your measurements must also cover a meaningful range, represent genuine replication, include uncertainty, and be recorded with enough precision and context to answer the research question.
What the current IB Biology IA assesses
Under the Biology course with first assessment in 2025, the IA is formally called the scientific investigation. It is an open-ended investigation in which the student gathers and analyses data to answer an individually formulated research question. The report has a maximum overall word count of 3,000 words, contributes 20% of the final Biology grade, and is assessed using four criteria worth six marks each:
- Research design: 6 marks
- Data analysis: 6 marks
- Conclusion: 6 marks
- Evaluation: 6 marks
The Data analysis criterion considers how clearly and precisely relevant data are recorded, processed, and presented. It also considers treatment of uncertainties and whether the processing enables the research question to be addressed. Students can examine these expectations through the current IB Biology IA rubric grader and compare different approaches in the Biology IA exemplar library.
Sufficiency is therefore judged in context. The question is not simply, “Did I collect 25 numbers?” It is, “Can these data support accurate processing, interpretation, and a justified biological conclusion?”
How many independent-variable levels should you use?
For an investigation involving a continuous independent variable such as temperature, pH, concentration, light intensity, or salinity, aim for at least five well-spaced levels when feasible. Five to seven levels often provide enough points to reveal a trend, curvature, optimum, threshold, or plateau.
This is a widely used practical recommendation rather than a fixed IB requirement. Certain valid investigations use two categorical groups, database data, field sampling, or repeated measurements over time, so a five-level structure would not always be appropriate.
Choose a meaningful range
The range is the distance between the lowest and highest values tested. It should be wide enough to produce a biologically meaningful response without creating unsafe, unrealistic, or uncontrollable conditions.
For example, testing catalase activity at 20°C, 21°C, 22°C, 23°C, and 24°C technically provides five levels, but the range may be too narrow to reveal the enzyme's temperature response. Testing 10°C, 20°C, 30°C, 40°C, and 50°C is more likely to reveal a trend, provided the method can maintain those temperatures reliably.
Select useful intervals
Your interval is the difference between adjacent levels. Equal intervals often make trends easier to interpret, but they are not compulsory. Unequal intervals may be justified when pilot data indicate that the response changes rapidly within one part of the range.
A useful range should:
- include the values relevant to the research question
- avoid clustering all measurements in a narrow region
- be achievable with the available apparatus
- include a control or baseline where biologically appropriate
- provide enough distinct points for the intended graph or model
A pilot investigation is the most reliable way to test the range before committing time and materials to full collection. RevisionDojo's guide to designing an effective science IA experiment provides a broader planning framework.
How many repeats or trials are enough?
For a typical controlled laboratory investigation, three repeats per level should usually be regarded as a lower practical starting point, while five independent repeats per level generally produce a more defensible estimate of biological variation. More may be needed if the response is highly variable, the expected effect is small, or the statistical test requires a larger sample.
This leads to the common “5 by 5” plan:
| Design feature | Practical starting point | Why it helps |
|---|---|---|
| Independent-variable levels | 5 or more | Reveals the form of a trend |
| Independent repeats per level | About 5 | Estimates variability and reduces dependence on one result |
| Total raw outcomes | About 25 | Supports means, spread, graphs, and basic statistical analysis |
This table is planning advice, not an official IB threshold. Twenty-five poor measurements do not become reliable merely because they fit the pattern, while a carefully justified field or database investigation may require a completely different sample size.
Genuine repeats versus repeated readings
A genuine replicate applies the same treatment to a new, independent experimental unit. Measuring five separate potato cylinders at each sucrose concentration captures variation among cylinders. Reading the same cylinder's mass five times mainly measures balance or handling variation.
| Type | Example | What it estimates |
|---|---|---|
| Biological replicate | Five independently prepared leaf discs | Biological and preparation variability |
| Technical replicate | Measuring one sample three times | Measurement or instrument variability |
| Pseudoreplication | Treating repeated readings of one sample as five independent samples | Artificially inflates the apparent sample size |
Biological replicates are usually more important when your conclusion concerns biological organisms or materials. Technical repeats can still be useful when measurement precision is a major concern, but they should be identified accurately rather than presented as independent sample size. Research on replication likewise emphasizes that independent biological samples capture different information from repeated measurements of the same sample.
Randomizing the order of treatments can also reduce bias. If every low-temperature trial is conducted first and every high-temperature trial last, gradual changes in room conditions, reagent age, or organism condition could become confused with the independent variable.
Record raw data correctly during collection
Raw data are the original measurements and observations obtained before calculations such as means, rates, percentages, or standard deviations. Record them immediately rather than reconstructing a table from memory after the experiment.
A strong raw-data table should include:
- a specific table number and descriptive title
- the independent variable in the first column
- separate columns or clearly labelled rows for repeats
- units in column headings, not after every value
- instrument uncertainty or resolution where relevant
- consistent decimal places matching the measuring instrument
- all collected results, including unexpected values
- brief qualitative observations in a separate table or adjacent notes
For example, a heading could read Mass after osmosis / g (±0.01 g). If the same balance is used throughout, the uncertainty belongs in the heading rather than being repeated in every cell.
Do not replace raw trials with means. A mean hides variation and prevents the reader from checking calculations, identifying anomalies, or judging the consistency of the measurements. Keep raw and processed data in separate, clearly labelled tables.
Useful qualitative observations might include unexpected cloudiness, damaged tissue, incomplete submersion, colour change, condensation, or visible contamination. Such observations can later explain anomalous values and support a more specific evaluation. The RevisionDojo guide to presenting IA data and results gives further examples of table and graph conventions.
Why thin data limits Data analysis marks
Thin data are data that are too sparse, narrow, inconsistent, or dependent to support the intended conclusion. Examples include only two temperatures for a claimed continuous relationship, one measurement per treatment, or many readings taken from the same biological sample and incorrectly treated as independent.
The current Data analysis criterion rewards relevant processing that addresses the research question, clear and precise communication, and appropriate consideration of uncertainties. Sparse data restrict what you can legitimately do in each area:
- A mean from one result is impossible, and a mean from two results is unstable.
- Spread and reliability cannot be evaluated convincingly without repeats.
- A trend based on two or three levels may miss a plateau, optimum, or nonlinear relationship.
- Outliers cannot be identified responsibly when there is little comparison data.
- Statistical tests may have low power or fail to meet their assumptions.
- A conclusion may overstate what the evidence actually demonstrates.
There is no automatic rule that fewer than 25 values caps the criterion at a particular mark. Nevertheless, major omissions or processing that cannot adequately address the research question are incompatible with the strongest performance descriptors. In practice, weak collection creates a ceiling because later analysis cannot manufacture information that was never measured.
Before collecting data, decide what processed values and graph you expect to produce. If you want means with standard deviations at each concentration, you need enough independent observations at every concentration to make those summaries meaningful. The RevisionDojo statistics guide for science IAs can help you connect the design to an appropriate analysis.
A practical data-collection workflow
Before the main experiment
- Define the experimental unit, independent variable, dependent variable, and controlled variables.
- Run a pilot using a reduced number of levels and repeats.
- Adjust the range, timing, concentrations, and measurement method.
- Decide how many independent repeats are feasible and scientifically defensible.
- Prepare blank raw-data and qualitative-observation tables.
- Record instrument resolution and determine how uncertainty will be reported.
During collection
- Follow the same protocol for every trial.
- Randomize or alternate treatment order when sequence effects are plausible.
- Record results immediately and retain unexpected values.
- Note procedural deviations, sample losses, and equipment problems.
- Keep sufficient precision in raw values and avoid premature rounding.
- Repeat failed trials only according to a consistent, documented rule.
After collection
Check whether every independent-variable level has the planned number of valid observations. Never invent, duplicate, or quietly delete measurements to make the table balanced. If a result is excluded, preserve the original value and justify the decision using scientific evidence rather than the fact that it weakens the expected pattern.
You can compare your finished structure with sample Biology IA investigations or use Jojo AI to check whether your tables, calculations, and explanations are understandable. Feedback should improve your reasoning and communication, but the investigation and analysis must remain your own work.
Common data-collection mistakes
The most damaging mistakes are usually decisions made before the final report is written:
- treating “five by five” as an IB rule rather than a design benchmark
- using too narrow an independent-variable range
- collecting unequal repeats without explaining missing samples
- confusing repeated readings with independent biological replicates
- recording only averages instead of original measurements
- omitting units, uncertainties, or consistent precision
- deleting anomalous results without justification
- choosing a statistical test after collection that the dataset cannot support
- collecting many values while failing to control major confounding variables
More data cannot compensate for systematic bias. If increasing temperature also changes pH, or if treatment groups are measured on different days with different equipment, additional repetitions may simply reproduce the same design flaw.
Conclusion
Effective IB Biology IA data collection combines an appropriate independent-variable range, enough distinct levels, genuine repeats, controlled conditions, and precise raw-data recording. For many continuous laboratory investigations, five or more levels with approximately five independent repeats per level is a useful benchmark, but it is a recommendation rather than an official numerical requirement.
Plan backward from the graph, uncertainty treatment, and statistical analysis you intend to use. RevisionDojo's IA exemplars, rubric tools, and Jojo AI can help you review the clarity of your design. To strengthen the same data-handling skills for examinations, attempt questions first and then use the IB Biology past paper video solutions to compare your interpretation with a worked approach.
Sources and referenced URLs
- Official IB Biology subject brief, first assessment 2025
- IB Biology guide for first assessment 2025
- NIST guidance on completely randomized experimental designs
- Nature Methods explanation of biological and technical replication
- RevisionDojo IB Biology IA rubric grader
- RevisionDojo Biology IA exemplar library
- RevisionDojo guide to designing an effective science IA experiment
- RevisionDojo guide to presenting IA data and results
- RevisionDojo guide to statistical analysis in a science IA
- RevisionDojo sample Biology IA investigation
- RevisionDojo IB Biology past paper video solutions