How is a math curriculum evaluated? A research-based framework

Marc Colomer
12/02/2026|9 min read
How is a math curriculum evaluated? A research-based framework

When a school adopts a math curriculum program, a common question often arises, both at home and at school: “Okay, but… does it work?” It’s a fair question. The problem is that the answer cannot be reduced to a simple yes or no.

In this article, we analyze the complexity of evaluating a math program from a broad educational research perspective.

A broad view: what does it mean to “evaluate well”?

The National Research Council (NRC) warned years ago that evaluating a math curriculum should not be reduced to a single result or a single type of study. In its 2004 report, it proposes understanding evaluation as a set of complementary evidence: impact studies, implementation studies, content analysis, and analysis of the theoretical framework (that is, what theory of learning and teaching supports the program). In fact, that same report reviewed 698 studies of 19 math curricula and concluded that only 147 met minimum quality standards —and that even these were not enough on their own to establish the effectiveness of a program: no single evaluation can settle the question.

This idea is key: an educational program is not just a set of materials. It is an example of how students learn and of the teaching practices that make that learning possible. That’s why, when evaluating it, it makes sense to ask questions such as:

  • Impact: Do results change? For which students and under what conditions?
  • Implementation: Is it being used as intended? What adaptations are made, and why?
  • Content: Do the activities really promote rich mathematical learning? Are the learning trajectories coherent?
  • Theoretical coherence: Is the program aligned with what research tells us about learning math?

In other words, impact studies are important, but they need to be complemented with other more qualitative studies in order to understand in depth how, when, why, and for whom the program has an impact.

The key role of implementation: what happens between the program and the results

When a new program arrives, it’s common to expect an automatic improvement in results. But inside and outside the classroom, things aren’t that simple (or that fast). What ends up happening looks more like a network of interconnected factors: structural elements (curriculum, school resources, education policy, inequalities, etc.); teacher factors (knowledge, beliefs, attitudes, identity, etc.); student factors (abilities, attitudes, identity, well-being, etc.); family and social context (expectations, support, math anxiety, resources, etc.).

And all these factors influence one key piece that is often taken for granted: implementation. That is, the way the program actually reaches the classroom: how consistent it is, how well it is carried out, whether it reaches all students, and how it is adapted to the school’s reality.

The scientific evidence is clear on this: adopting a new curriculum, on its own, barely moves results, while programs that transform what happens in the classroom —tutoring, cooperative work, classroom organization— have substantially larger effects (Slavin & Lake, 2008; Pellegrini et al., 2021).

What makes the difference? That the material is accompanied by a real change in teaching practice. Combining materials with classroom-focused training and implementation improves outcomes compared to applying only one of the two (Lynch et al., 2019), especially when changing pedagogical practices —structured lesson plans together with student materials and teacher coaching (Piper et al., 2018).

The conclusion is clear: a program can be good and still not produce the expected change if implementation is not strong enough, or if it does not generate any structural change in math teaching.

Evaluating impact: not all evidence says the same thing

A key aspect to consider when discussing impact evidence is that not all studies have the same level of rigor. Research design is fundamental, as it determines the extent to which we can attribute observed changes to the program rather than to other factors (such as socioeconomic status or the teacher’s prior motivation). Correlation, then, does not imply causation.

In the United States, the ESSA law (Every Student Succeeds Act) classifies evidence into four levels from a scientific quality perspective: from strong evidence (randomized controlled trials or RCTs) to promising evidence (correlational studies) or demonstration of reasoning (theoretical foundation and ongoing evaluation).

This common paradigm has led to the creation of different institutions designed to ensure scrutiny of any research that aims to demonstrate one of these levels of scientific quality. The most important are the What Works Clearinghouse (WWC) and Evidence for ESSA, two repositories that collect, review, and classify this evidence according to its methodological rigor. By applying common criteria and regardless of who develops each program, they allow us to situate, compare, and classify results in an impartial way.

That said, having a good classification does not mean evidence is abundant. When Bellwether (2023) reviewed the 59 highest-quality elementary and secondary math curricula, 42 of them had no published effectiveness study —an evidence gap of roughly 70%—. And among the studies that did exist, only 10% were randomized controlled trials.

Key variables in quantitative studies

In impact studies using quantitative methods, there are two indicators that often appear and are important to understand well:

  • The p value tells us whether the observed difference could be due to chance.
  • The effect size (such as Cohen’s d or Hedges’ g) tells us how large this difference is.

How do we interpret effect size?

There are different interpretation frameworks. For example, Hattie (2009) popularized the idea of a “benchmark” around 0.4 standard deviations to indicate a clearly meaningful impact.

However, Kraft (2020) proposes more realistic benchmarks for applied education research: less than 0.05 indicates a small effect, between 0.05 and less than 0.20 a medium effect, and 0.20 or more a large effect, also emphasizing that interpretation depends on essential aspects of the study design.

These benchmarks are not arbitrary. Kraft (2023) compiled 3,426 effect sizes from 973 randomized controlled trials in education and observed that the median is around 0.10 standard deviations, that more than a third of the studies don’t reach 0.05, and that only 10% exceed 0.57. Evans and Yuan (2022) find a nearly identical distribution (median of 0.10) in low- and middle-income countries. In education research, then, an effect around 0.10 would already be considered meaningful.

In fact, when reviewing the evidence-based programs analyzed by WWC or Evidence for ESSA, we see that effects fall within modest but significant ranges that vary depending on the type of program. Comprehensive math curricula tend to have effects between 0.09 and 0.11 SD, but effects tend to be larger in targeted educational interventions (0.19–0.51 SD) or in high-impact tutoring programs (0.20–0.40 SD).

What variables do we want to measure?

All this focus on p values and effect sizes assumes a prior decision that is rarely made explicit: what we have chosen to measure. And that points to an even more basic question: what do we mean by “improving” in math?

Are we talking about getting better standardized test results? And if so, on which tests and with what types of tasks? Or are we referring to deeper mathematical competence: the ability to reason, solve problems, and make sense of what you are doing?

 

Maybe we also want to measure less visible but equally relevant changes: attitudes toward math, confidence, and students’ math identity. Or changes in teaching practice: how instruction happens, what learning opportunities are created in the classroom, what interactions take place.

And there is still another dimension that is often left out of short-term studies: long-term effects. Academic decisions, learning pathways, or even future career opportunities.

All of these perspectives make sense, but some are more difficult to measure than others. For example, measuring the impact of a math program on students’ future career opportunities requires many years of follow-up, is very costly, and is often unfeasible. In this sense, most studies rely on standardized test results, but it is important to understand that there are other relevant aspects that these tests are not able to measure.

On top of this, there is an important technical nuance: measuring with a test aligned with the program itself is not the same as measuring with a standardized, independent test. Evaluations tailored to the program tend to show larger effects, so results obtained with program-specific tests should be read with caution. In short, what is measured conditions both the answer and how much improvement seems to occur.

Educational evidence with a critical eye

At this point, one last word of caution is in order: even when a study reports an effect size, that number is not a fixed property of the program. There are several factors that can increase or decrease it, and knowing them helps to read the evidence with judgment.

First of all, what is measured matters. Evaluating with a test designed to match the program is not the same as evaluating with a standardized, independent test: measures aligned with the program’s own content tend to show systematically larger effects. That’s why it makes sense to give more weight to independent tests and to view results obtained with program-specific evaluations with caution.

What and who the program is compared against also matters. Comparing a program against a group that receives nothing produces larger effects than comparing it against an equally active alternative. And studies that select especially favorable schools, or that monitor implementation with an intensity that is hard to replicate, obtain higher effects than those working with representative samples. In those cases, the effect should be understood as a ceiling, not as what can be expected at scale.

It also matters when things are measured. Many improvements fade over time (the so-called fade-out): a year later, typically only between a third and a half of the initial effect remains (Bailey et al., 2017). Since almost no studies measure beyond one school year, the effects reported are almost always immediate estimates that overstate cumulative impact. Hence the value of measuring not only right after the intervention, but also in the medium term.

Finally, it is worth remembering that educational evidence has a strong bias toward positive results: null or negative results are rarely published. On top of this, there is a very consistent pattern: pilot effects almost always exceed those observed when the program is scaled up —coaching, for example, drops from 0.28 to 0.10 SD when moving from pilot studies to large-scale studies (Kraft, Blazar & Hogan, 2018)— and studies with small samples report effects that are much larger than those of large studies (Evans & Yuan, 2022).

As a practical rule, the most reliable effects are those from independent studies, measured with standardized tests, compared against real alternatives, and with follow-up beyond one school year.

None of this means that the material doesn’t matter. On the contrary: it matters especially when it is designed to bring about well-grounded instructional change; when it structures good practice, offers high-value activities, and supports teachers’ professional development. That’s where a good program stops being a stack of pages and becomes a lever for improvement.

Closing the loop: evidence to improve

We can see it: evaluating an educational program is not about finding a simple verdict, but about understanding a complex process. It is about building a map that helps us answer what works, for whom, under what conditions, with what supports, and with what implementation costs.

That’s why it makes sense to combine different types of evidence and to keep investing in both rigorous studies and implementation research. Because in the end, the important question isn’t only “Does it work?” but what needs to happen for it to work here, with these teachers and these students?

References and materials

Bailey, D. H., Duncan, G. J., Odgers, C. L., & Yu, W. (2017). Persistence and fadeout in the impacts of child and adolescent interventions. Journal of Research on Educational Effectiveness, 10(1), 7–39.

Bellwether Education Partners. (2023). Rounding up: An analysis of math curriculum effectiveness studies. Bill & Melinda Gates Foundation.

Evans, D. K., & Yuan, F. (2022). How big are effect sizes in international education studies? Educational Evaluation and Policy Analysis, 44(3), 532–540.

Hattie, J. (2009). Visible learning: A synthesis of over 800 meta-analyses relating to achievement. Routledge.

Kraft, M. A. (2020). Interpreting effect sizes of education interventions. Educational Researcher, 49(4), 241–253. https://doi.org/10.3102/0013189X20912798

Kraft, M. A. (2023). The effect size benchmark that matters most: Education interventions often fail. Educational Researcher, 52(3), 183–187.

Kraft, M. A., Blazar, D., & Hogan, D. (2018). The effect of teacher coaching on instruction and achievement: A meta-analysis of the causal evidence. Review of Educational Research, 88(4), 547–588.

Lynch, K., Hill, H. C., Gonzalez, K. E., & Pollard, C. (2019). Strengthening the research base that informs STEM instructional improvement efforts: A meta-analysis. Educational Evaluation and Policy Analysis, 41(3), 260–303.

National Research Council. (2004). On evaluating curricular effectiveness: Judging the quality of K-12 mathematics evaluations. Washington, DC: The National Academies Press.

Pellegrini, M., Lake, C., Neitzel, A., & Slavin, R. E. (2021). Effective programs in elementary mathematics: A best-evidence synthesis. AERA Open, 7(1), 1–29.

Piper, B., Zuilkowski, S. S., Dubeck, M., Jepkemei, E., & King, S. J. (2018). Identifying the essential ingredients to literacy and numeracy improvement: Teacher professional development and coaching, student textbooks, and structured teachers’ guides. World Development, 106, 324–336.

Slavin, R. E., & Lake, C. (2008). Effective programs in elementary mathematics: A best-evidence synthesis. Review of Educational Research, 78(3), 427–515.

U.S. Department of Education. (2016). Non-regulatory guidance: Using evidence to strengthen education investments. Washington, DC: Author. https://www.ed.gov/sites/ed/files/policy/elsec/leg/essa/guidanceuseseinvestment.pdf

Related reads