cross 1,000+ live cohorts and 50,000+ students, we rigorously track full-semester attendance, active participation, and progressive assessments. This massive scale allows us to decode exactly how sustained engagement translates into measurable academic mastery.
The sample dataset provided captures a uniquely high-retention cohort from the Jan–Jun 2025 semester. Because this data successfully validates scalable pedagogical strategies, we are actively seeking research partners to collaborate with us in uncovering further cognitive learning insights.
A pedagogically sound library of over 600,000 step-by-step Mathematics, Physics, and Chemistry video solutions.Every solution is 100% expert-verified and delivers unbroken, step-by-step explanations with zero logical gaps. This high-signal structure makes it an ideal dataset for fine-tuning foundational models to advance complex reasoning capabilities.
This dataset currently powers foundational models at top AI laboratories. To test our data fidelity, we have made public a set of 1,000 videos and corresponding metadata for evaluation. This evaluation package also includes a comprehensive user guide and English-language solution images showcasing fully worked-out answers. Summary of content creation and quality control processes are provide in the information provided.
Processing over 50,000 handwritten student submissions monthly from prescribed question sets allows us to map common misconceptions in Mathematics (and soon, Sciences). After extracting and transcribing text from these images, we manually sample annotate the specific errors and cognitive gaps within the students' answers.
We are currently fine-tuning open-weight AI models—both internally and through partnerships—to accurately detect these nuanced mistakes. This capability lays the groundwork for AI tools that encourage students to write out their thought processes. By promoting this productive struggle, we are unlocking the ability to evaluate open-ended assessments at scale.
This initial sample features 100 annotated misconceptions, representing an early iteration of our methodology. We will continuously update this dataset as our annotation pipeline progresses.
This repository provides a 1,000-sample dataset designed to train and fine-tune AI Optical Character Recognition (OCR) models specifically for the education technology purpose. The sample is subset of database containing over 20 million student question digital input.
The core value of this dataset lies in its authenticity: it features real student inputs, split between handwritten submissions and digital text (typed or screenshots) with different lighting conditions and clarity. This makes it highly valuable for developing consumer AI products in the education sector that need to accurately parse imperfect, real-world user uploads.
