PluraMath Preprint and Dataset Released

We released PluraMath, a human-curated benchmark extending mathematical reasoning evaluation to 18 underrepresented languages across six language families.

The work evaluates 27 reasoning LLMs and finds a persistent performance gap between high-resource and underrepresented languages. I contributed as a co-author and Serbian-language contributor.

Read the paper · Explore the code · View the dataset