Skip to main content
Grantee Research Can Higher-Order Thinking Skills Associated with Upward Mobility Be Measured Using Statewide Assessment Data?
Last updated on
Display Date

Statewide English language arts exams can capture higher-order thinking skills at scale, but it’s not easy to do, especially as exams change from year to year and vary in terms of their cognitive demand.

Higher-order thinking skills (HOTS) are a core subset of durable skills that include analytical thinking, reasoning, and problem-solving abilities. These skills can set students up for lifelong success but are complex and burdensome for districts and schools to measure at scale and therefore have been challenging for researchers to examine. 

In this study, the researchers code and sort more than two decades of statewide 3rd-to-10th-grade English language arts (ELA) test items from Massachusetts into those that capture HOTS and those that capture lower-order thinking skills (LOTS) using Webb’s Depth of Knowledge framework. The researchers then link the coding with information about student performance on those test items to develop new tools for capturing HOTS, examine these tools’ measurement properties, and determine whether the tools can capture HOTS as distinct from LOTS using data districts already collect statewide. 

The researchers find that measuring HOTS and distinguishing them from LOTS using preexisting ELA item-level data is challenging given that these exams were designed to measure unidimensional ELA ability. That said, it is possible to use these tests to distinguish HOTS from LOTS under certain conditions. En route to that conclusion, the researchers add considerable nuance to the pursuit of using existing data to measure cognitive or durable skills. On the conceptual side, they find that LOTS tend to develop first and are foundational to HOTS and that HOTS are predictive of long-term outcomes even after accounting for LOTS. More pragmatically, the researchers conclude that the testing regime matters when measuring HOTS and that certain types of questions and grade-level exams were more likely than others to capture HOTS.

Key Takeaways

Webb’s Depth of Knowledge framework can be used to distinguish thinking skills (both HOTS and LOTS) from question difficulty, as reflected by variation in difficulty ratings even among HOTS questions. Example items coded as capturing HOTS and LOTS with both lower and higher levels of difficulty are provided in figure 1. The Depth of Knowledge framework has been used to design many widely administered assessments, including statewide tests in 26 states and the Measures of Academic Progress test. On the ELA tests examined for this project, HOTS questions were most likely to be on the reading subscale and to be open ended.

 

Figure 1. Examples of HOTS and LOTS Items with Low and High Difficulty

Figure 1. Examples of HOTS and LOTS Items with Low and High Difficulty

Note: These items are sourced from the Massachusetts Department of Elementary and Secondary Education’s publicly available 2023 Grade 3 assessment, which permits reproduction for non-commercial purposes. The items are based on a passage from A Vacation in Ruins by Precious McKenzie, which is not shown due to its copyright. Item difficulties are generated using Item Response Theory Three-Parameter Logistic and Graded Response Models with all items (including those without text released). The rationale for each item's categorization as HOTS or LOTS are as follows: (LOTS, Low Difficulty) item asks students to summarize an event in the text; (LOTS, High Difficulty) item asks students to identify the meaning of a word in context; (HOTS, Low Difficulty) item asks students to identify/make inferences about explicit or implicit themes; and (HOTS, High Difficulty) item asks students to explain, generalize, or connect ideas using supporting evidence.

 

Student performance supports that HOTS and LOTS are separate and developmentally related skill sets. The researchers found that HOTS and LOTS appear to develop sequentially, with HOTS building off the foundation of LOTS. It was unusual for students to answer a HOTS question correctly if they answered a LOTS item in the same topical domain incorrectly. Inversely, students who performed well on LOTS questions were far more likely to correctly answer related HOTS questions. Further supporting that HOTS and LOTS are distinct but both important, the team found that each was a significant predictor of postsecondary educational attainment, even after controlling for the other.

In general, the researchers found that it was difficult to distinguish HOTS from LOTS using preexisting statewide ELA exams. To be a useful vehicle for measuring and distinguishing between HOTS and LOTS, a test needs to have a nontrivial proportion of more cognitively demanding (HOTS) questions but not too many. The researchers found that Massachusetts ELA assessments from before the Common Core era fall in this window more reliably than assessments from after the introduction of Common Core–aligned statewide assessments. Newly aligned assessments included a larger share of HOTS items than the exams from the earlier testing regimes, making it a useful measure of HOTS but harder to use to disentangle HOTS from LOTS.

Distinguishing HOTS from LOTS was challenging even with the pre–Common Core–aligned assessments, but there were some grades and years among these assessments for which there was evidence that the new tools could capture HOTS as distinct from LOTS. For these grade-years, the team also found that the tools could be used to compare across groups of interest based on social class or race or ethnicity, allowing researchers to study inequality. Future research could examine what characterizes these tests and how commonplace they are beyond Massachusetts.

Multiple states’ contemporary ELA assessments are likely similar to Massachusetts’s pre–Common Core assessments in terms of their cognitive demand, suggesting that they might be suitable for measuring HOTS and distinguishing them from LOTS.

The authors also found that it was easier to detect HOTS from ELA ability in the later-grade assessments studied in this context (compared with 3rd grade, for example), with 10th grade being the strongest, though this finding is especially preliminary.

Body

Potential Implications for Policymakers and Practitioners

Additional validation and translational work are needed to bolster these findings and enable their uptake into policy and practice settings, but SUMI sees the following potential applications:

Use existing data to learn about HOTS, when possible, while exercising caution. For policymakers, the findings suggest that existing statewide assessment systems can sometimes provide meaningful information about HOTS—without requiring entirely new testing systems—but only under certain design conditions that are not always well understood. This opens the door to more cost-effective strategies for incorporating durable skills into accountability and reporting frameworks and for using historical data to learn about HOTS development. That said, given the challenges of using existing data for these purposes, practitioners and researchers alike should carefully examine the measurement properties of new tools before using them for these ends.

Measure and promote both higher- and lower-order skill development. Despite its initial focus on the value of HOTS, this research ultimately suggests that LOTS are important building blocks for HOTS development. The findings further suggest that LOTS are an important predictor of postsecondary attainment in their own right. Therefore, policymakers and practitioners should not prioritize learning about and promoting higher-order learning at the expense of these foundational skills. For assessment developers, structured cognitive frameworks such as Webb’s Depth of Knowledge, especially when they are already built into assessment design, could also be scaffolded into reporting to add value for states, district, leaders, and parents seeking to measure and understand students’ durable skills.

Consider the development of domain-specific HOTS. The findings from this research are more consistent with the idea that HOTS are specific to the domains or subject areas in which they are being applied rather than domain general. In other words, it appears more likely that there are HOTS specific to reading that may not generalize to science or math. Researchers and policymakers should therefore learn more about how these skills may differ across domains. Practitioners should explore efforts to tailor the development of these skills to specific subjects.

Future Research Directions

Future research could explore the following:

  • the characteristics of assessments that lend themselves to distinguishing HOTS from LOTS and how widely those characteristics are found beyond Massachusetts tests
  • whether there are grades or assessment contexts that particularly lend themselves to HOTS or LOTS measurement
  • the points at which knowledge is developed enough to build on HOTS in different subjects and at different developmental stages
  • the use of AI paired with the Depth of Knowledge framework and the coding experience from this study to code items for cognitive demand and sort them into HOTS and LOTS more efficiently
  • the transferability of these approaches and insights to other states and subjects
  • additional approaches to distinguishing HOTS at scale (without the ability to measure these skills, we cannot answer other policy-relevant questions about how to efficiently and equitably develop these skills)

Methods, Data Sources, and Measures

This research team analyzed test item-level responses over two decades on the Massachusetts Comprehensive Assessment System ELA test. This encompassed 133 assessment combinations and 9 million student-grade-year observations from students in 3rd through 8th and 10th grades. The researchers also looked at two separate anonymized datasets for school characteristics.

The researchers used Webb’s Depth of Knowledge framework to code individual test items for cognitive demand, sort questions into HOTS or LOTS, and pair questions with student-level performance data. They then analyzed student performance on ELA test questions and then paired the coding with student item-level performance data to create HOTS and LOTS measures and assess the measurement properties of these new tools.

To assess the skills’ predictive validity, the researchers examined school-level HOTS and LOTS measures against key mobility outcomes, including school-level educational attainment (school-level ninth-grade cohort graduation within four years and postsecondary enrollment within 16 months of graduation).

 



Disclosures

All expressed opinions, findings, and errors are the author’s and do not represent the views of the Georgia Governor’s Office of Education and Workforce Strategy (GOEWS) or any of the participating agencies.

Body

Research Team

Beth Schueler

Principal Investigator, Stanford University

Jim Soland

University of Virginia

John Wang

Georgia Governor’s Office of Education and Workforce Strategy


Emerging Insights

Measuring skills that drive student upward mobility


External resources

The Challenge of Capturing Higher-Order Thinking Skills at Scale (PDF)

Tags Measuring skills that drive student upward mobility
Related content