LEVANTE is not designed for group comparisons

One of the most exciting aspects of the LEVANTE project is its global reach. As LEVANTE datasets are labeled by country, it may be tempting to compare two datasets and interpret these as population differences between two countries. This is not a valid use of LEVANTE data. Sites differ in numerous ways that make direct comparisons of absolute levels of performance meaningless and misleading. We expect that researchers in the LEVANTE community are committed to using data ethically, and therefore will plan and interpret their analyses thoughtfully to avoid these pitfalls. 

LEVANTE isn’t PISA

One way to bring this issue to light is to consider an entirely different yet also global project: the Programme for International Student Assessment (PISA). Administered every three years, the PISA measures 15-year-olds in areas such as math, reading, and science across more than 90 countries every three years. Critically, PISA uses representative sampling, first identifying a representative sample of schools and then representative 15-year-old students within those schools. This supports their aims to guide improvement and reform in education policy through direct cross-national comparisons.

LEVANTE sites also examine academic skills and competencies in many countries. Unlike PISA, LEVANTE is a federated cohort study, meaning that our various partner sites each manage their own study design and data collection. While PISA and many other large, multi-site studies are highly standardized in sampling, recruitment, and administration, LEVANTE functions as a collection of many inexpensive projects run by autonomous research teams at different sites. Each site works toward independent goals with a high degree of flexibility while remaining united through a shared infrastructure and harmonized core measures. The result is a large sample with a balance of standardization and rich, varied data. This is aligned with LEVANTE’s aims, which are to support a basic science understanding of variability in learning across contexts. 

The federated nature of LEVANTE makes it fundamentally different from PISA when it comes to planning and interpreting analyses. For example, average math scores tend to be higher in data collected by the pilot site at the Max Planck Institute for Evolutionary Anthropology in Germany than those collected by the pilot site at the Universidad de los Andes in Colombia.  While the participants in Colombia were from a mix of urban and rural schools and completed their assessments in a group testing format, participants in Germany were recruited from an existing database where they self-nominated to participate and completed tasks on home devices. In this case, both sampling biases and administration conditions likely impact scores, invalidating inferences about absolute country-level differences. Finally, as neither sample is representative of the overall population of Germany or Colombia, in our presentation of these data we refer to the samples by dataset names (mpieva-de and uniandes-co) rather than country names (Germany and Colombia).

Sites using LEVANTE may target different populations, recruit samples that are not representative of the general population, and administer tasks in differing formats. A hypothetical Site A and Site B are depicted to highlight some potential differences. Image created with StoryTribe.

Pitfalls to avoid

Representative sampling of any national population is not the norm for most LEVANTE sites. Thus, variability in sampling, as well as differences in administration across sites significantly impacts the types of inferences that can be made with the data. Specifically, it often compromises the validity of cross-national comparisons. 

Furthermore, there are communities that are invested in group comparisons as a way to make inferences that are motivated by ideology, not science. For example, direct cross-national comparisons have been used to perpetuate racist ideologies. This is an invalid and harmful use of the LEVANTE data.

We offer the following “don’ts”:

  • DON’T interpret absolute performance and intercepts across groups as meaningful differences.
  • DON’T ascribe group differences in any scores to inherent national or group characteristics.
  • DON’T describe sites by country alone — use dataset names or descriptive indicators instead.

What to do instead

Rather than country or group comparisons, the LEVANTE data are better suited for characterizing learning processes and trajectories across the sample and/or within subgroups. LEVANTE is well-suited for understanding when key factors might vary in their impact on developmental outcomes across sites, and may also be useful for certain types of comparisons that are intentionally built into the study design. Two examples:

  1. The LEVANTE team has been working on analyses on the mental rotation task. First, we sought to understand if mental rotation scores were comparable across sites, finding scalar invariance. This finding suggests that we are using a comparable ruler across sites to assess children on their spatial cognitive abilities. Next, we investigated whether angular disparity effects (performance decreases with angle size) replicate across sites, and separately, the extent to which ability scores and response times differed by sex within sites. These approaches leverage the strengths of the LEVANTE dataset to test substantive questions without overinterpretation of group comparisons.
  2. Some researchers are also exploring uses of the LEVANTE tasks as low-cost pre- and post- measures for evaluating intervention efficacy on learning outcomes in math and reading. Although results are not normed to national representative samples, such analyses may be informative when comparisons are made relative to control schools in the same region that are carefully selected for having similar characteristics. 

We offer the following “do’s” as recommendations:

  • DO consider whether scores even share a scale across sites—to date, evaluations of measurement invariance suggest this is largely so in LEVANTE, except for certain language tasks (see our preprint), but we will continue to monitor this question across new sites. 
  • DO pursue research questions that consider the degree of variability in learning processes and trajectories within datasets from each site.
  • DO resist misuses of developmental and cognitive science to perpetuate racism and racial hereditarian research programs(see References/Resources below).

Conclusions

The federated structure of LEVANTE means that it is actually composed of numerous small studies, each with their own and sometimes wildly differing study designs (despite overlapping measures). This introduces nuances to ethical and responsible use of the LEVANTE data that must be attended to by researchers pursuing analyses and interpretation. Most crucially, we highlight that LEVANTE’s federated structure invalidates many inferences stemming from direct cross-national differences. However, this same feature makes LEVANTE a flexible and exciting tool for the study of developmental variability.

References and Resources

Bird, K. A., Jackson, J. P., Jr., & Winston, A. S. (2024). Confronting scientific racism in psychology: Lessons from evolutionary biology and genetics. American Psychologist, 79(4), 497–508. https://doi-org.stanford.idm.oclc.org/10.1037/amp0001228  

Frank, M. C., Baumgartner, H. A., Braginsky, M., Kachergis, G., Lightbody, A. A., Sparks, R. Z., Zhu, R., Carlson, S. M., Graham, S., Lipina, S. J., Newcombe, N. S., Odgers, C. L., Pianta, R. C., Siegler, R. S., Snowling, M., Yoshikawa, H., Cubillo, A., & Dodge, K. A. (2025). Learning Variability Network Exchange (LEVANTE): A Global Framework for Measuring Children’s Learning Variability Through Collaborative Data Sharing. Child Development. https://doi.org/10.1111/cdev.70011

Kachergis, G., O’Reilly, F., Braginsky, M., Xiao, X., Lightbody, A. A., Shannon, K. A., Watson, Z., Zhang, L., Zhu, R., Abutto, A.B., Wanjing A.M., Long, B., Murray, T., Yeatman, J.D., Sulik, M.J., Obradovic, J., Jenkins, N., Ansari, D., Perfetti, M.C., … Frank, M. C. (2025, December 17). Creation and validation of the LEVANTE core tasks: Internationalized measures of learning and development for children ages 5-12 years. https://doi.org/10.31234/osf.io/r4dhw_v1

Sear, R. (2026, August 24). ‘National IQ’ datasets do not provide accurate, unbiased or comparable measures of cognitive ability worldwide. Retrieved from osf.io/preprints/psyarxiv/26vfb_v2