Mapping how students learn

Reverse cognitive offloading by gaining visibility into productive struggle. Explore the datasets capturing the real student learning journey at scale.

Pedagogical research

How can we engineer productive struggle and harness peer communities to drive intrinsic motivation? By analyzing interaction data from over 50,000 students, we continually pressure-testing established cognitive frameworks to reveal which pedagogical strategies thrive at scale and how to do them effectively.

AI for education

Building AI that avoids cognitive offloading requires models trained to identify misconceptions and guide learning—not spoon-feed answers. This demands a foundational dataset of authentic student workings, rigorously annotated by expert teachers.

Assessment design

If backward design dictates that how we assess shapes how students learn, open-ended evaluations are key to reducing cognitive offloading. How can we leverage scalable systems—like advanced AI models—to parse and grade this complex reasoning without crushing teachers with overhead?

It takes an ecosystem to drive meaningful change.

We welcome collaboration with institutions, companies, and non-profits dedicated to student success.

Data & Tools

Longitudinal cohort engagement & in-class assessment data

cross 1,000+ live cohorts and 50,000+ students, we rigorously track full-semester attendance, active participation, and progressive assessments. This massive scale allows us to decode exactly how sustained engagement translates into measurable academic mastery.

The sample dataset provided captures a uniquely high-retention cohort from the Jan–Jun 2025 semester. Because this data successfully validates scalable pedagogical strategies, we are actively seeking research partners to collaborate with us in uncovering further cognitive learning insights.

Datasets include: 

STEM Q&A solutions (video and images)

A pedagogically sound library of over 600,000 step-by-step Mathematics, Physics, and Chemistry video solutions.Every solution is 100% expert-verified and delivers unbroken, step-by-step explanations with zero logical gaps. This high-signal structure makes it an ideal dataset for fine-tuning foundational models to advance complex reasoning capabilities.

This dataset currently powers foundational models at top AI laboratories. To test our data fidelity, we have made public a set of 1,000 videos and corresponding metadata for evaluation. This evaluation package also includes a comprehensive user guide and English-language solution images showcasing fully worked-out answers. Summary of content creation and quality control processes are provide in the information provided. 

Datasets include: 

Mapping students' misconception from handwritten workings on Practice questions

Processing over 50,000 handwritten student submissions monthly from prescribed question sets allows us to map common misconceptions in Mathematics (and soon, Sciences). After extracting and transcribing text from these images, we manually sample annotate the specific errors and cognitive gaps within the students' answers.

We are currently fine-tuning open-weight AI models—both internally and through partnerships—to accurately detect these nuanced mistakes. This capability lays the groundwork for AI tools that encourage students to write out their thought processes. By promoting this productive struggle, we are unlocking the ability to evaluate open-ended assessments at scale.

This initial sample features 100 annotated misconceptions, representing an early iteration of our methodology. We will continuously update this dataset as our annotation pipeline progresses.

Datasets include: 

Q&A pairs dataset for OCR fine-tuning

This repository provides a 1,000-sample dataset designed to train and fine-tune AI Optical Character Recognition (OCR) models specifically for the education technology purpose. The sample is subset of database containing over 20 million student question digital input.

The core value of this dataset lies in its authenticity: it features real student inputs, split between handwritten submissions and digital text (typed or screenshots) with different lighting conditions and clarity. This makes it highly valuable for developing consumer AI products in the education sector that need to accurately parse imperfect, real-world user uploads.

Datasets include: 

Connect with Us

We welcome collaboration with institutions, companies, and non-profits dedicated to student success.​