By Robert Meyer and Claudia Gentile
When it comes to accountability assessments, many wonder: Can we structure accountability assessments to reduce test burden without reducing the validity and reliability of the assessments’ findings?
As we’ve discussed in previous blog posts, state and district assessment systems are expected to serve multiple purposes. They need to provide valid and reliable accountability measures, while at the same time support educators with timely information for instruction and offer diagnostic insight for struggling learners, all while minimizing burden on schools. Fortunately, it is possible to design a comprehensive assessment system that serves multiple purposes by including different types of assessments.
In this blog post, we’ll be focusing on assessments primarily designed to provide accountability data. Before addressing this question about accountability structure, it’s important to step back and explore the major uses of accountability-based assessments.
Common Accountability Measures
Federal and state accountability requirements prioritize aggregate school- and district-level uses of summative assessments. They want a collection of measures about student achievement and a collection of measures about student growth.
The most common aggregate achievement measures are:
- the proportion of students with test scores at the proficient level or higher,
- the proportion of students at each proficiency level, especially the lowest proficiency level, and
- average achievement overall.
The most common aggregate growth measures are:
- value-added measures of growth, obtained from a statistical model of student achievement that controls for one or more years of prior achievement and, in some cases, student demographic characteristics, and
- means (or medians) of student growth percentiles, based on a statistical model of achievement that controls for prior achievement.
For achievement measures, districts most often use state assessments, supplemented by additional standardized tests from national vendors. To measure growth, a wider range of strategies is available.
Expanding Options to Reliably Measure Growth
Often summative assessments—those big picture, once-a-year tests—are used to provide data on student achievement and student growth. However, accountability assessments can also be performed as interim assessments.
These tests, given throughout the year, can provide states and districts with other valuable information to inform early warning and program evaluation. They can also contribute to end-of-year measurements of achievement and growth—especially for smaller schools that experience relative instability in their growth measure metrics from year to year.
This approach uses interim assessments (growth measures for assessments administered during the school year) to:

- measure student progress at each timepoint with respect to end-of-year annual student growth.
- combine one or more of these growth measures to create progressively more reliable growth measures during and at the end of the school year (which would be reported in school report cards).
- report on a growth measure at each interim test date that is optimally predictive of the end-of-year growth measure.
One option for minimizing the number of assessments used for high-stakes purposes is to incorporate only a single interim assessment in the augmented growth measure, preferably, the mid-year/winter interim assessment.
Meeting the Challenge
When we revisit the challenge for accountability assessments—structuring accountability assessments to reduce test burden without reducing the validity and reliability of the assessments’ findings—this expansive option (using interim assessments) provides several possible benefits.
It is intriguing to explore whether accountability assessment requirements paired with the addition of interim assessments (that seek data on growth measures) can reduce test burden. We believe it may be possible and worth exploring.
This combined approach may produce substantial increases in precision of growth measures, especially for small schools, which then may allow the state and test vendors to reduce the number of items on both the summative and interim assessments.
The combined approach may also produce summative test scores with sufficient reliability to create reliable annual growth measures.
Simply put, we posit including interim assessments as a component of accountability assessment systems can increase precision while reducing test length.
The Central Comprehensive Center continues to work with state education agencies, educators and researchers to develop approaches to improving state assessment systems, both the quality and usefulness of assessment data for all stakeholders. Through requests from our states, we’re diving into the opportunities for innovation within state assessment systems and highlighting our findings in a series of blog posts.
