An Independent Study Evaluates the Impact of Innovamat in Public Schools in Rio de Janeiro

This independent study by Germina shows that, after two years of implementation, students who used Innovamat achieved better results in mathematics than those who did not. The article accompanies a technical article and an Executive Summary prepared by the research group that led the study
How did it all begin?
A few years ago, some members of Fundação Lemann, one of the most important foundations in education in Brazil, visited our offices. Their interest was to help improve mathematics education in public schools in Brazil. And the challenge was no small one. In Brazil, many students finish compulsory education with a mathematics level below what is expected. As the Germina report we share here reminds us, in 2023 only 16.5% of 9th grade students (3rd year of middle school) reached adequate learning in mathematics.
After that visit came many conversations. At Innovamat, we connected with that desire to improve mathematics learning in Brazil. We wanted to build a relationship and find a true fit, but we knew we had to take the step of assuming that the proposal had to adapt to a reality very different from ours.
That is how the journey in Rio de Janeiro began.
A pilot study with a rigorous evaluation
The first agreement came with the Municipal Department of Education of Rio de Janeiro, within the Technological Educational Gymnasiums, a network of public schools in the city. We started with a small pilot of 5 schools. The initial goal was to understand whether the project fit the reality of the country, the schools, the teachers, and the students.
From the start, we agreed on something important: the pilot had to be accompanied by an independent impact evaluation. With the help of Fundação Lemann, the evaluation was funded and it was decided to work with Germina, a Brazilian research institution whose study proposal and team gave us a lot of confidence. The evaluation was carried out independently, with a formal agreement with the Municipal Department of Education of Rio and approval from Brazil’s national ethics committee.
If we wanted to know what was happening, we needed data external to Innovamat, an independent look, and a design that would allow the results to be interpreted with caution.
Before measuring, we had to understand how change could happen
The evaluation did not start directly with a test. First we worked on co-designing a theory of change: an organized explanation of how we expected the program to generate impact, based on our experience in hundreds of schools and on the evidence we had accumulated in previous studies.
And here there was a central idea we shared from the beginning: educational improvement is not immediate.
Innovamat does not improve results through a one-time or isolated intervention with students. It acts mainly through the teacher. The program offers manipulative materials, teaching guides, a digital platform, training, and support. But real change happens when teachers are able to adapt the proposal to the needs of their classroom. That is when the activities come to life and get students participating in a different way in the mathematics classroom.
That is why Germina distinguished between short-term impacts and long-term impacts:
- In the short term, we expected to see changes in teaching practices, classroom dynamics, participation, and student motivation.
- In the long term, we expected to see improvements in performance on mathematics tests.
The hypothesis was that changes could appear in the classroom during the first year; but learning results would probably need more time.
Based on this theory of change, Germina designed the study independently and using data independent of Innovamat: mainly the results of the ADR, the diagnostic assessment of the municipal network of Rio de Janeiro.
How was the study designed?
The study had two phases. The first was the initial 2024 pilot. Germina designed a stratified randomization to select the treated schools and compare them with a control group (schools that do not use Innovamat) as similar as possible. In practice, this meant using public data from the schools—such as previous performance, socioeconomic level, size, proportion of students from different ethnic groups, and other variables—to form comparable groups and assign the treatment randomly within those groups. This random assignment helps reduce possible selection bias, that is, it avoids having only schools or people with particular characteristics that could influence the results, such as greater motivation or a special interest in mathematics. In addition, using stratified randomization makes it possible for the groups being compared to be as balanced as possible. In this way, the results obtained can be interpreted with greater confidence.
In the end, 5 schools started with Innovamat and 34 schools remained in the control group. This first phase allowed them to observe two things: what happened during the first year of implementation and what happened after two years in those same schools.
The second phase came in 2025, when the program was expanded to 25 more schools. In this case, the selection was not random, but administrative: it was made by the Municipal Department of Education of Rio itself. For this reason, these results must be read with more caution. In addition, for these schools we only have data from the first year of implementation so far. To know whether the long-term effects are replicated there, we will have to wait for the data from the 2026-2027 school year.
The main analysis was done with a methodology called differences in differences (DiD, by its acronym in English). Simply put: it compares how students in schools with Innovamat evolve with how students in similar schools without Innovamat evolve. In addition, Germina complemented these data with teacher surveys: they obtained 1,049 responses from 393 schools at three different moments. And also qualitative interviews to better understand the changes in teaching practice.
First result: the class changes
The first learning was very clear: in the short term, teachers described changes in the way they taught.
Teachers in schools with Innovamat reported using more physical and digital resources, proposing more active learning situations, and giving students more structured feedback. In the executive summary, Germina highlights three figures: +28 percentage points in the use of digital resources, +20 points in active learning practices, and +25 points in structured feedback. But there are many more nuances, and that is why we recommend reading the technical article in more detail. This is important because it confirms an essential part of the theory of change. Before seeing improvements in tests, we needed to see that the program was changing what was happening in the classroom. And that is what the data show.
Second result: after two years, students learn more
In the first year of the pilot, Germina did not detect a significant effect on learning at the aggregate level. This did not surprise us, as it was consistent with the initial hypothesis. Changing teaching routines, adapting materials, and transforming classroom dynamics takes time.
However, the most relevant part came in the second year. In the 5 schools of the first phase, 2nd and 3rd grade students showed significant improvements in mathematics. In 2nd grade, the estimated effect was +0.16 standard deviations, equivalent to approximately +9 SAEB points (Brazil’s national assessment system, which measures school performance and the quality of the education system through standardized, census-based exams). In 3rd grade, the effect was +0.20 standard deviations, equivalent to approximately +12 SAEB points. Both results were statistically significant. Converting them to SAEB points helps make the result more intuitive, although Germina warns that it should be interpreted as an approximation: the ADR and SAEB are different tests.
As Germina notes in its technical article, these effect sizes are relevant if we compare them with those usually found in rigorous evaluations of educational interventions. Germina compares them with the expanded Kraft (2023) database, which compiles results from randomized trials with standardized learning tests in six international repositories, including the What Works Clearinghouse and the Education Endowment Foundation. In that comparison, the median of the effects is usually between +0.03σ and +0.14σ, depending on the source. The effect observed in 3rd grade (+0.20σ) exceeds the 70th percentile in five of the six databases analyzed, and the 2nd grade effect (+0.16σ) exceeds it in four. Simply put: these are not only positive results, but, in context, they are above what is usual in this type of study.
It is also important to be cautious: in the first phase, only 5 schools were using Innovamat. The results are positive, promising, and statistically significant, but it will be important to confirm that the effect remains when we have the second year of data from the 25 schools in the second phase.
Third result: gaps are also reduced
The study also analyzed equity. Following common practices in educational policy research in Brazil, Germina groups as PPI racialized people from groups that have historically suffered greater educational inequalities (the groups classified as Preta, Parda, and Indígena in the National Census of Brazil).
Considering the two phases and the years available, the program was associated with a significant reduction of 0.07 standard deviations in the gap between PPI and non-PPI students.
This result seems especially relevant to us. Innovamat did not include a specific racial equity intervention in this pilot, so the improvement seems to occur indirectly: by rebuilding learning and offering richer, more accessible, and more participatory mathematics experiences, the program benefits proportionally more those who started from a more vulnerable situation.
General discussion
For us, the most important point to highlight is that the results are consistent with other previous studies we had carried out internally (studies library), which provides evidence in favor of using the program in the Rio de Janeiro context. In other words, in a context quite different from previous studies and with independent research, we reached very similar conclusions: Innovamat supports teachers in generating changes in classroom practices toward mathematics more centered on deep conceptual learning and the development of competencies. This takes time, training, and familiarity with the program on the part of the teaching team. But these changes pay off: in the second year of implementation, students show better performance on mathematics tests. In addition, these gains seem to have more impact for students from more vulnerable groups and thus help reduce the equity gap.
This study does not close the conversation about the impact of Innovamat on mathematics learning. On the contrary: it opens it with more rigor. The initial sample is small, the second phase still needs a second year of follow-up, and the results will have to be compared with other tests, such as Prova Rio, SAEB 2025, and future mathematics fluency assessments. But, in part thanks to these positive results, in the 2026-2027 school year Innovamat is already being implemented in 159 public schools in Rio de Janeiro. There is a great opportunity to keep researching and learning about implementation challenges, changes in the classroom, and how all of this relates to the results.
This study reinforces a simple but very important idea: the way we teach matters. And it also reminds us that in education, to know whether something works, it is not enough to implement it: we must support it, evaluate it well, recognize its limits, and keep learning from the information.
Having the analysis of external institutions is very positive and tells us we are moving in the right direction. And that is why we are open for the scientific community and the best researchers to keep working with us, examining our proposal to continue doing top-level joint research.
References
Kraft, M. A. (2023). The Effect-Size Benchmark That Matters Most: Education Interventions Often Fail. Educational Researcher, 52(3), 183-187.

