The ConMat and Their Validity

Marc Colomer
Marc Colomer|03/04/2025|11 min read
In collaboration with: Arnau Lagarda
The ConMat and Their Validity

As Grace M. Hopper pointed out, measuring precisely has immeasurable value. It is the first step to detect the state of things and see what needs improvement.

In this article, we will discuss the ConMat, our reference assessment that aims to measure students’ mathematical competence. For us, this means knowing the mathematical concepts presented in the curriculum for each grade while mastering essential skills such as problem solving, communication, reasoning, and the ability to make connections.

But having a measurement tool is not enough; we also need to validate that the tool is valid and measures correctly. What is necessary to ensure this is really the case? In collaboration with Judith Peñafiel and Arnau Lagarda, statistics experts from the Biomedical Research Institute of Bellvitge (IDIBELL), we have been working to answer this question. In this article, we present three aspects necessary to evaluate the validity of a test:

Define a rigorous process to design the tool. That is, systematization, rigor, and with experts involved. We will call this aspect creation process.

Validate that the test is able to measure optimally; that is, it is appropriate for the level of the participants, is consistent, and has temporal stability (measures similarly if done more than once). We will call all of this internal validity.

Determine that the test measures what it intends to measure. A good way to know this is by identifying its relationship with other tests that, a priori, measure something similar. We will call this external validity.

Here we will see, then, what these three aspects look like in the case of the ConMat.

Creation process

Above all, it is necessary to define the theoretical framework very well. In this case, the ConMat are inspired by the framework of the TIMSS assessment (Trends in International Mathematics and Science Study), which distinguishes between content domains and cognitive domains, crossing both dimensions. Additionally, within the cognitive domains, we find a subdivision into two categories: content, which encompasses concepts (facts, vocabulary, etc.), and procedures (operations, reading graphs, use of measurement tools, etc.); and another subdivision for the four processes: problem-solving, reasoning and proof, connections, and communication and representation. The following tables show the distribution of questions in different grades according to these divisions.

Figure 1. ConMat question structure.

* ConMat2 is administered at the beginning of 3rd grade.

** In the latest edition of ConMat7, we have added a new question (content domain: Number and Operations; cognitive domain: procedural).

Once the theoretical framework is defined, it is necessary to be systematic and rigorous to ensure the expected result. In the case of ConMat assessments, Innovamat’s expert evaluation team follows the following steps each time they design a ConMat test:

  1. Creation Process: curriculum content and the defined theoretical framework are taken into account.

  2. Pilot: the test is administered to hundreds of students, discussions are held with students and teachers, and results are evaluated using statistical analysis.

  3. Iteration: the test is revised according to statistical analysis*, which determines if questions are too easy or difficult and if they adequately discriminate between students of different levels.

  4. Pilot Repetition: the test is piloted again. If the results are positive, the process is considered complete. If necessary improvements are identified, the iteration is repeated.

  5. Final Production: the test is translated into all corresponding languages and voilà, we have it!

*Statistical analysis of the pilot (dropdown): to analyze the pilot results and revise the test, the following are considered:
  • The criteria of Cohen and Swerdlik (pages 248 and 249, 2009), according to which items should have around 62.5% correct answers. This value is obtained by calculating 50% —on average, only half of the population is expected to answer each item correctly— plus the probability of getting it right by chance (25%, considering that the test is multiple choice with 4 options) divided by two. That is, 62.5 = 50 + 25/2. We consider items with correct answer percentages between 42.5% (62.5% − 20%) and 92.5% (62.5% + 30%) acceptable.
  • Also taken into account are the results of item modeling through item response theory. To better understand what item response theory is, see the following section.
  • Finally, qualitative observations made on the day of the test are considered to define possible improvements.

Internal Validity

During the creation process, each ConMat test is validated to ensure it is adapted to the students’ level. This validation is conducted with a small sample of schools that have helped us with the pilot. The question is, therefore, whether the test is appropriate for the vast majority. In this section, internal validity is evaluated with data from more than 100,000 students from 3rd grade to 7th grade, which are the grades that had ConMat assessments available in the 2024-2025 school year.

One of the best ways to determine if a test measures well and is appropriate for the population taking it is through what is called item response theory to the item, or item response theory (IRT), in English. In general terms, IRT is based on the idea of relating the general —estimated— level of participants with the skills measured by each question (item) on the test. This type of analysis is used by assessments such as PISA or TIMSS to validate that the questions are appropriate and to calculate the estimated level of each student.

However, there are more aspects to internally validate how robust a test is. On one hand, the Cronbach’s alpha parameter, which is a measure that helps us know if a set of questions within a test measure the same concept and are consistent with each other.

Finally, it is also important to look at the temporal stability of the test; that is, whether the level of students, measured with the test on several occasions, has similar results over time.

To carry out these analyses with maximum rigor, we have collaborated with Judith Peñafiel and Arnau Lagarda, from IDIBELL, who, externally, have analyzed the ConMat from the end of last school year (2023-2024). You can see the technical document with all the results here. But let’s go step by step and look in detail at how the internal validity of the ConMat has been analyzed.

Item Response Theory

IRT aims to relate the general —estimated— level of participants in a test with the skills measured by each item on the test. In the case of ConMat, the 3PL (Three-Parameter Logistic) model has been used and validity criteria have been defined according to the technical article by Bichi and Talib (2018). To better understand what IRT is about, we should know that each item or question generates a curve like this, called an item characteristic curve (ICC, for its acronym in English).

Figure 2

Adaptation of the figure from Bichi and Talib (2018), showing an example of ICC. On the x axis we see the estimated level of the student, which is standardized so that the average level is 0. The red sigmoid indicates the probability (y axis) that a specific student, according to the estimated level shown on the test, correctly answers this specific item. A probability equal to 1 indicates that the student will definitely answer the item correctly, according to the model, while a probability of 0 indicates that they definitely will not answer it correctly.

Let’s now look at what type of curves we want to obtain to consider a question valid.

In the 3PL model, there are three key factors we need to define:

  • Item discrimination (also called parameter a): allows us to see if the item can discriminate between different levels. This happens when the probability of higher-level students answering the item correctly is greater than that of lower-level students. In figure 2, the factor a will be greater the steeper the slope of the curve; that is, the more markedly it discriminates between students of different levels (see more examples in figure 3).

  • Item difficulty (also called parameter b): indicates at what point the curve (ICC) has the steepest slope. This is what we call average in figure 2 (or average in English) and indicates from what point there is greater discrimination between higher-level and lower-level students. The further to the left we have the average of the slope, the easier the item will be (see more examples in figure 3).

  • Random guessing (or guessing in English, also called parameter c): this last parameter indicates the probability that students with a very low level will guess the correct answer. In the case of ConMat, since the items are multiple-choice with 4 possible answers, we would expect this value to be around 0.25. In figure 2, the factor c will be equal at the extreme left point of the curve. Although this value can be important to see if it is necessary to revise the answer choices for a specific item, in this article we will not go into too much detail about this; it is not as relevant as the other factors for showing how valid a test is (but see figure 3 to understand its potential usefulness).

Figure 3.

Examples of ICC for some ConMat4 questions. The red dotted line indicates parameter b or item difficulty. We can see that the items in each column have different difficulties, but the rows have similar difficulties. For example, we observe that question 9 (b = -1.50) is considerably easier than question 12 (b = 0.52). That is, students with a standardized level of 0 will answer question 9 correctly with a very high probability, but they will be less likely to answer item 12 correctly. However, regarding the item discrimination, we see a big difference between the items in the first row. Item 9 (a = 2.07) has a steeper slope than item 21 (a = 0.56) and, therefore, is more precise in discriminating between levels. In other words, the difference between the probability of getting each item correct when comparing levels -1.5 and 0 (x axis) is greater in item 9 (from 50 % probability at level -1.5 to almost 100 % probability of getting it right at level 0) than in item 21 (from 50 % at level -1.5 to less than 75 % at level 0). That is, although both items are of similar difficulty, item 21 will be more precise in discriminating between students at levels -1.5 and 0. Finally, we see the factor c or random guessing. In this case, the elements in the bottom row have a higher c value than those in the top row. This means that no matter how low the student’s level is, in items 22 and 12, low-performing students will have chances of getting the answer right. In the case of question 22, this value is higher than 25 %, which indicates that one of the 4 answers is very unlikely to be selected by any student. On the other hand, questions 21 and 29 have c values lower than 25 %, which may indicate that there is a trick answer that many low-level students fall for.

Criteria according to Bichi and Talib (2018) to determine if an item is appropriate or not.

Item discrimination (a). If a is less than or equal to 0.64, it means the item should be eliminated or needs revision. From 0.64 to 1.34 means the item is already sufficiently valid, although it could be slightly revised if deemed necessary. From 1.35 and above, it is considered a good item for discriminating between levels.

Item difficulty (b). If b is less than -2, it indicates the item is very easy; from -2 to -1, it is easy; from -1 to 1, it is moderately difficult; from 1 to 2, it is difficult; and greater than 2, it is very difficult. Keep in mind that the student level values are standardized so that 0 is the average level of all responses.

From these curves that we obtain for each question or item, we can create what is called test information curve (TIC). This curve shows how much information—or precision—the test as a whole provides about different student ability levels (θ). It combines the information from each item and shows at which point it best discriminates between latent traits (or levels). Therefore, it helps us get an overall view of how valid the test is and whether it provides enough information to discriminate between different levels. Let’s look at an example!

Figure 4.

Example of a test information curve (TIC) from ConMat5. The x-axis indicates the standardized student level (θ), while the y-axis indicates how much information the test can provide about each level [I(θ)]. In other words, how good the test is at identifying students with a certain level θ. On the other hand, the dotted line shows the standard error of measurement (SE), which is inversely related to the information curve. The more information the test provides about the level, the lower the standard error will be, indicating that the measurement is more precise. In general, we would want the test to be very good at discriminating students around a level of 0 and for the curve to be wide enough to discriminate students with levels closer to the extremes. Taking into account the criteria of Bichi and Talib (2018), the extremes from -2 to 2 should include the majority of the population (considering that items below -2 are considered very easy, and items above 2, very difficult). Therefore, it is important that within this range the solid red curve provides sufficient information I(θ) and, above all, that the I(θ) value is above the value of the dotted curve [SE(θ)].

What have been the results of the ConMat assessments?

To answer this question, we present a summary table of the results. It is based on the parameters of Bichi and Talib (2018) defined earlier.

Table 1. Qualitative summary of IRT results considering factors a (level discrimination) and b (item difficulty). Items are considered not optimal when they require revision (a≤0.64) and are good or very good when a≥1.35. Ideally, items should be distributed across different levels, but there should be few (or no) items at the “very easy” and “very difficult” levels if the test’s objective is to discriminate levels for the majority of the population. Below each test (row), the sample size (N) is indicated, meaning how many students completed the assessment. The numbers within each cell indicate the number of items in the corresponding test and column.

* One of the two items (16) has a a negative value (-0.67) due to a technical error in the correction of the result, and, therefore, it is not related to the item needing revision.

On the other hand, in the following figure you can see a summary of all the TICs, which generally show how good the test is at identifying students with a certain estimated level (θ).

Figure 5.

TIC of each ConMat test. The x axis indicates the estimated level of students in the IRT model (indicated by the symbol θ), standardized with a mean of 0. The y axis on the left shows how much information the test provides according to the level, I(θ). The higher it is, the more information it provides. The y axis on the right shows the standard error of measurement of the test, SE(θ). The lower it is, the more precise the test is at discriminating between students of level θ.

Both the table and the TIC curves show that all ConMat tests have good validity, as they are well adapted to the level of the majority of the population that has taken the test and allow discrimination between levels. Let’s summarize this:

  1. As shown in the first column of table 1, the majority of items in all tests are good at discriminating between students of different levels and do not need to be revised.

  2. There is sufficient variation between the level of the items and how good they are at discriminating at each level (table 1, columns 3-7), which makes the tests good at discriminating between students of different levels in the range of -2 to 2 (figure 5). However, since there are few “very easy” items (b < -2, table 1 – column 3) or “very difficult” items (b > 2, table 1 – column 7), the test is considerably less precise at discriminating between students with levels at the extremes. In other words, if a student has a very low or very very low level, the test will probably not be able to differentiate it. The same will happen with students of very high or very very high level.

  3. In general, the tests provide more information around level 0, that is, around the average level of the population that has taken the test (figure 5). However, except for ConMat6, the peak of the TIC curves is slightly shifted to the right (the clearest case is ConMat3, where the peak is around 0.7). This means either that the test is somewhat too difficult, or that the slightly difficult items are the ones that allow better discrimination between student levels.

  4. The TIC curves for ConMat5, ConMat6, and ConMat7 have a higher peak (they provide more information around level 0) than the TIC curves for ConMat3 and ConMat4. This result is likely related to the fact that student maturity plays an important role. The more mature they are, the more consistent their responses will be and the easier it will be to measure their mathematical competence. In contrast, at early ages, factors such as reading comprehension or being less trained in answering test questions for 1 hour can cause some randomness and make the item response theory (IRT) somewhat less precise.

Cronbach’s Alpha

It is widely used in psychology, education, and social research to check the reliability of a questionnaire or test. When a test has a Cronbach’s alpha greater than 0.7, it is often considered sufficiently consistent (Taber, 2018). In the case of ConMat tests, all tests have shown a value greater than 0.7, with a range from 0.78 in ConMat3 to 0.85 in ConMat6, which obtained the highest value. As in the previous analyses, we see that the tests are somewhat less robust at younger ages, although according to Taber’s criteria (2018), they are sufficiently robust at all ages.

Temporal Stability

It allows us to confirm that the responses to a test are not completely random, but rather measure students’ mathematical knowledge in a robust way. In this case, the objective is to administer the test more than once and see if student responses are consistent over time. For ConMat tests, the test is administered twice, at the end of one school year and at the beginning of the next. In this case, the tests from the end of the 2023-2024 school year were compared with the tests from the beginning of the 2024-2025 school year. All tests were compared except for ConMat6, due to the transition between educational stages.

To check temporal stability, one can look at either the correlation between the two times the test was taken, or the differences in results between both administrations using a statistical test, such as a paired T-test.

In this case, all tests have similar results. The correlation (Spearman) between both tests is greater than 0.7 in all cases, except for ConMat7, which is 0.65. For the rest, the range goes from 0.71 in ConMat3 to 0.8 in ConMat5. These values indicate a strong correlation between the two tests. On the other hand, the differences between the results at the end of the school year and the beginning of the next, in all cases, are small with effect sizes (Cohen’s d) less than 0.15.

External Validity

Once we have ensured that a rigorous design process is followed and that the test has internal validity, we have one final step. Are we really measuring students’ mathematical competence? A good way to answer this question is to correlate the ConMat results with the results of other external mathematics tests. How? We have obtained data from external tests through questions from previous editions of official tests and designed an adaptation that follows the appropriate structure. This is what we call mock exams, which we offer to schools to conduct a simulation of the official test. In the 2023-2024 school year, this was done with the official New Jersey (USA) tests in 3rd and 4th grade, with the TIMSS mathematics test in 4th grade (in Mexico) and 5th grade [in Catalonia (Spain)], and with the basic skills test in Catalonia for 6th grade. We can correlate the data obtained through this test with the data from the ConMat tests that the same students completed at the end of the school year. In other words, we are trying to see if those who score high on one test also tend to score high on the other, and vice versa. Below we show the correlation of the ConMat tests with each external test adaptation.

Note for interpreting the graphs: in the correlation graph, we will see a scatter plot, where each point represents a student with their score on both tests. The linear model line is a straight line that attempts to represent the general trend of the data. If the line rises from left to right, it means there is a positive relationship (students with good scores on one test also get good scores on the other). If the line goes down, it means the relationship is negative. The gray area represents the 95% confidence interval, which indicates the margin of error of the prediction made by the model. The narrower this band, the more confidence we have in the model. If the area is very wide, it means there is a lot of variability and the relationship between the two tests is not entirely clear. Finally, we will see the R parameters (Pearson correlation). The R indicates the strength and direction of the relationship between the two variables. An R greater than 0.7 is considered a high positive correlation.

In all cases, we see a high positive correlation (R > 0.7, except in the 4th grade TIMSS, where the correlation also approaches 0.7). That is, the results that students obtain in the ConMat are very similar to the results they obtain in adaptations of different external tests.

Conclusion

With all the data obtained and the analyses performed, we have evidence that the ConMat assessments are valid and useful for measuring students’ mathematical proficiency. In summary, we can say that:

  1. There is a rigorous process for creating the ConMat assessments that ensures the resulting test has the expected properties.

  2. The analysis of internal validity, through item response theory (item response theory), shows that the test can discriminate between different levels and is adapted to the general level of the population taking it. Additionally, the test is robust across questions, as shown by Cronbach’s alpha, and has good temporal stability, as demonstrated by the comparison between the end-of-year ConMat and the beginning of the following year.

  3. The external validation shows that the test measures skills similar to other mathematics tests designed by external experts to Innovamat.

References

Bichi, A. A., i Talib, R. (2018) Item Response Theory: An Introduction to Latent Trait Models to Test and Item Development. International Journal of Evaluation and Research in Education, 7, 142-151. https://doi.org/10.11591/ijere.v7i2.12900

Cohen, R., i Swerdlik, M. (2009). Psychological Testing and Assessment: An Introduction to Tests and Measurement (7th ed.). Nova York: McGraw Hill.

Taber, K. S. The Use of Cronbach’s Alpha When Developing and Reporting Research Instruments in Science Education. Res Sci Educ 48, 1273-1296 (2018). https://doi.org/10.1007/s11165-016-9602-2