Executive functioning skills help young children pay attention, manage their behavior, remember information, and shift between tasks. These skills are associated with children’s future academic performance and educational attainment, income, and health as adults.
But executive functioning can be difficult to measure at scale, often requiring additional testing time, specialized administration, and trained staff members. And many measures overburden young learners.
Technology-based assessments could offer a new, low-burden way to measure young learners’ executive functioning. Schools are increasingly using these assessments to test children’s academic skills, including in early childhood classrooms. But the technology also records metadata about how the student took the assessment, including how long they took to respond to each question and whether they completed assessment activities. Schools generally do not have easy access to or use these data.
In this study, MDRC researchers analyze whether metadata can be used to measure prekindergarten students’ executive functioning. They constructed six metadata-based measures of executive functioning.
The team found promising initial evidence that some of these measures capture executive functioning, but the results also suggest that the types of metadata used and how they are combined matter considerably. The strongest measures showed modest relationships with established executive functioning assessments, while other measures showed little or no relationship. The findings suggest that, with additional work, schools could use metadata as a low-burden way to identify children who may be struggling with executive functioning.
Key Takeaways
Metadata-based scores were better at identifying children who are struggling with executive functioning than at distinguishing between children with average scores and children with above-average skills. The measures also generally differentiated children in the lower third of the skill distribution from those in the upper portions of the distribution. But their measures were less effective at distinguishing between children with average executive functioning versus children with high executive functioning.
Using existing literature and assessments, as well as educator input, the research team constructed six metadata-based measures of executive functioning:
- response speeds that are not unusually fast
- response speeds that are not unusually slow
- response speeds that are neither unusually fast nor unusually slow
- average response time on correct answers
- task completion
- a “two-vector” score combining response time and accuracy
All six measures passed the threshold for statistical reliability (Cronbach’s alpha > 0.60 or that had confidence intervals that include 0.6), meaning that students performed similarly on a given measure across performance tasks.
Measures that combined response time and accuracy and the average response time on correct answers had the strongest relationship with the established measures of executive functioning. These measures were better predictors of executive functioning skills than measures based on response speed alone. One measured the average response time for questions a child answered correctly. The other, a two-vector score, combined accuracy with response time.
Unusually fast responses were not a useful indicator of weaker executive functioning among prekindergarten students. Previous research has suggested that among older students, unusually fast responses can indicate disengagement or guessing and therefore show lower executive functioning. The researchers found that a lack of unusually slow responses (i.e., responding in the typical amount of time or longer) among prekindergarten students did not indicate stronger executive functioning.
The “not unusually fast” measure had correlations ranging from -0.24 to -0.01 with established executive functioning. The measure also showed weak or negative relationships with the other metadata measures. In short, unlike for older kids, “not too fast” doesn’t tell us much about young students’ executive functioning.
This could reflect an important difference between the use of metadata in assessments for young students and for older students, whose data are used in most of the existing literature. Prekindergarten tasks often ask children to complete simple, single-step activities, such as identifying a letter or shape. In those situations, a quick response may be appropriate.
Metadata-based scores were more strongly correlated with existing executive functioning measures than with academic skill scores on the assessment tools. This shows that the metadata measures are measuring something different than the academic skills that the assessment tools are intended to measure. But the two-vector score is an exception, with strong correlations with both executive functioning and the academic skills measured by the assessment tools. This reflects the way the measure was constructed, which includes the child’s performance on the underlying task, and highlights the challenge of distinguishing between constructs when using this hybrid scoring approach.
Potential Implications for Policymakers and Practitioners
Additional validation and translational work are needed to bolster these findings and enable their uptake into policy and practice settings, but SUMI sees the following potential applications:
Early childhood educators could use metadata as another way to gain insights into children’s behavior and skills and to identify children who need additional support. The current evidence does not support replacing direct assessments of executive functioning with metadata. But with additional work, including future studies that validate these measures, metadata could help educators identify children who may benefit from a more comprehensive assessment of their skills without requiring every child in a classroom to complete additional executive functioning testing.
Assessment developers could design technology-based assessments to better use metadata. Digital assessments of prekindergarten and older students already collect information about children’s response behavior. These data could hold valuable information. Assessment developers could intentionally preserve and improve the usability of these data so they can be used for purposes beyond scoring academic skills.
Future Research
Validate the measures in larger and more diverse samples. Larger samples could help researchers determine whether these findings replicate across children, prekindergarten settings, and technology-based assessment platforms and across demographic characteristics (e.g., race, sex, and language).
Explore additional types of metadata. The measures in this study primarily capture attention and engagement. Future assessments could collect information such as how response times change when children encounter a new type of item or how many times a child taps the screen before selecting a response. These data could provide information about cognitive flexibility and other aspects of executive functioning.
Examine whether metadata predict long-term student outcomes. Established measures of executive functioning have been shown to predict outcomes that matter for children’s long-term success, such as kindergarten readiness, later academic achievement, and educational and economic outcomes in adulthood. A future study could test whether metadata measures of executive functioning are similarly predictive of future success.
Methods, Data Sources, and Measures
The researchers analyzed data from 428 prekindergarten children in 18 prekindergarten programs across five states. These students participated in a nine-week pilot in fall 2024 using two technology-based assessments: one from Khan Academy Kids (279 children) and the other from Kibeam (149 children) These assessments were under development at the time of the pilot study and are primarily used to assess language, literacy, and math.
Using existing research, established executive functioning assessments, and input from educators who participated in focus groups, the team identified behaviors during a digital assessment that might signal executive functioning. They used assessment metadata—including response times, accuracy, and task completion—to construct six metadata-based measures. These measures primarily capture children’s attention and focused engagement rather than the full range of executive functioning skills.
For 59 students across Khan Academy Kids and Kibeam, the researchers also tested executive functioning using three established measures, the HTKS-R measure (Head-Toes-Knees-Shoulders-Revised), Pencil Tap, and the Preschool Self-Regulation Assessment Assessor Report. The researchers used Pearson correlations to examine the relationship between scores on the metadata measures and the three established measures, a form of concurrent validity. The researchers then used calibration analyses to determine whether the metadata measures placed children in roughly the same parts of the skill distribution as the established assessments. As a test of discriminant validity, they also used Pearson correlations to examine the relationship between scores on the metadata measures and scores on the academic skills measured by the Khan Academy Kids and Kibeam tools.
Research Team
Amy Taub
Principal Investigator, MDRC
Emily Hanno
Former PI, Overdeck Family Foundation
Sophie Litschwartz
Former Co-PI, Cambium Assessments
Emerging Insights
Measuring skills that drive student upward mobility