Showing posts with label assessment. Show all posts
Showing posts with label assessment. Show all posts

NCTM Denver 2013: Fennell and Wray's Math Specialists Get Ready Now: Common Core Assessments Are Coming

Annual Meeting - Friday, April 19, 3:30 pm

Francis (Skip) Fennell - NCTM Past President; McDaniel College, Westminster, Maryland
Jon Wray - NCTM Board of Directors; Howard County Public Schools, Ellicott City, Maryland

Skip Fennell, Jon Wray, and Beth Kobett (who was absent for this presentation) are the leads on ems&tl, the Elementary Mathematics Specialists & Teacher Leaders Project. As the name implies, the focus here is on supporting math specialists, such as district-level curriculum directors, instructional coaches, and anyone who is in a position to support mathematics teachers.

For this presentation, Fennell and Wray looked at the upcoming Common Core assessments, PARCC and Smarter Balanced (SB), and suggested ways math specialists can help teachers prepare for the tests.

Francis (Skip) Fennell

The challenge Fennell and Wray presented was essentially to focus on the upcoming assessments and respect the influence they will have on curriculum and instruction, without focusing too narrowly on the assessments and cause instruction and learning to suffer. This means, for example, not turning classroom practice into test prep, and using sample items from both PARCC and SB wisely.

Fennell and Wray used the concept of assessment literacy to describe the ability for teachers and specialists to understand a testing program. Many teachers have no formal training in assessment, so math specialists must be able to help them build their assessment literacy. Part of this is simply becoming more familiar with the schedules and formats of the upcoming PARCC and SB assessments. Both consortia offer more than just an end-of-year test, and teachers are going to need to help students interpret new kinds of technology-enabled assessment tasks.

Jon Wray

Fennell sees great potential in the CCSSM, but said, "If the Common Core becomes political, it's dead." Teachers and specialists need to work with the standards in ways that doesn't reduce them to a checklist of vaguely connected ideas. Using a number of items and task prototypes, Fennell and Wray showed examples of sample items from PARCC and SB and showed the many ways these could be used richely in lessons if the teacher provides the right support and instruction. "There are a lot of ways sample items can be used as instructional gems, " said Fennell. A list of potential questions and strategies for various tasks can be found in their slides.

The presentation wrapped up with an urging to better understand the role of formative assessment around these sample tasks. Also, encouragement was made to use materials from both PARCC and SB, regardless of the test your state has adopted. More task resources were linked to, including Illustrative Mathematics, the Institute for Mathematics and Education (especially the progressions documents), The Mathematics Common Core Toolbox, the PARCC Educator Leader Cadre Portal, and the Smarter Balanced Scientific Sample Pilot Test Portal.

The slides for this presentation are available here.

NCTM Denver 2013: Foegen et al.'s Building Progress Monitoring Measures for Algebra: Exploring Items and Scores

Research Presession - Tuesday, April 16, 8:30 am

Anne Foegen - Iowa State University
Barbara Dougherty - University of Missouri, Columbia
Vickie Spain - University of Missouri, Columbia
Jeannette Olson - Iowa State University
Subha Singamaneni - Iowa State University

One of the essential challenges in measurement is to design an instrument that captures the maximum amount of information using a minimum amount of time and resources. (Hint: You can't have both.) While I agree with Lorrie Shepard (2000) that "good assessment tasks are interchangeable with good instructional tasks" (p. 8), there can be room for other kinds of assessments for specific purposes.

This collaboration between researchers at Iowa State University and the University of Missouri, Columbia, is attempting to develop a set of algebra assessments for progress monitoring that students can take in a matter of minutes and can be graded by a teacher equally quickly. Part of the goal is to create the kind of progress monitoring instrument that would be useful for Response to Intervention (RtI). It's definitely something that is needed, as I remember once being told I had to use a simple arithmetic assessment for RtI (for high schoolers!) because an algebra version did not exist for the RtI system the district was using.

We spent time in this discussion session talking about the tasks on both the procedural and conceptual progress-monitoring instruments. While some of the problems were certainly interesting in the way they tried to address student thinking, most of the instruments looked very, very traditional. There is a heavy focus on symbol manipulation and almost nothing comes associated with even a hint of context. Yes, a math student with strong formal skills would do well on these assessments, but I couldn't help but think they're missing something in their approach. That feeling was probably best summarized with the question I asked them: "This assessment looks like something designed for students of Saxon textbooks. How would you feel if that is the case and your instrument is used to promote Saxon texts?"

I left reminding myself two things: (a) projects at the end of the first of a federally funded project are still very much at the learning/big revisions stage and (b) they wanted a test that could be taken in 5-7 minutes, and they got one. If there's something impressive about this effort, it's probably that this team seems to have come to grips with the reality of sacrificing some quality for efficiency. They're pushing themselves to make the best <10 minute algebra tests they can, and I hope their end result is useful, or, failing that, helps us better understand practical limits of assessment.

OpenComps: Validity and Causal Inference

With the start of my comprehensive exams beginning in 12 days, my studying has hit the homestretch. Thankfully, my advisor has inspired some confidence by telling me that my understanding of the math education literature is solid and I won't need any more studying in that area. That's good for my studying, and something I take as a huge compliment. So now I can focus for a while on preparing myself for the exam question Derek Briggs is likely to throw my way. Typically, one of the three people on a comps committee is tasked with asking a question related to either the quantitative or qualitative research methodology we learn in our first year of our doctoral program. Derek is a top-notch quantitative researcher, and I enjoyed taking two classes from him last year: Measurement in Survey Research and Advanced Topics in Measurement. Where this gets slightly tricky is that Derek didn't actually teach either of my first-year quantitative methods courses, so there's a potential I could get surprised by something he normally teaches in those classes that I didn't see. It's a risk I was willing to take after working with Derek more recently and more closely in the two measurement courses last year.

It certainly won't be a surprise if Derek asks a question that focuses on issues of validity and causal inference. He mentioned it to me personally and put it in a study guide, so studying it now will be time well spent. I feel like I've had a tendency to read the validity literature a bit too quickly or superficially, so this is a good opportunity for me to revisit some of the papers I've looked at over the past couple of years. Here's the list I've put together for myself:

AERA/APA/NCME. (1999). Standards for educational and psychological testing. Washington, D.C.: American Educational Research Association. [Just the first chapter, "Validity."]

Angoff, W. H. (1988). Validity: An evolving concept. In H. Wainer & H. Braun (Eds.), Test validity (pp. 19–32). Mahwah, NJ: Lawrence Erlbaum Associates.

Borsboom, D., Cramer, A. O. J., Kievit, R. A., Scholten, A. Z., & Franic, S. (2009). The end of construct validity. In R. W. Lissitz (Ed.), The concept of validity: Revisions, new directions, and applications (pp. 135–170). Information Age Publishing.

Brookhart, S. M. (2003). Developing measurement theory for classroom assessment purposes and uses. Educational Measurement: Issues and Practice, 22(4), 5–12. doi:10.1111/j.1745-3992.2003.tb00139.x

Chatterji, M. (2003). Designing and using tools for educational assessment (p. 512). Boston, MA: Allyn & Bacon. [Chapter 3, "Quality of Assessment Results: Validity, Reliability, and Utility"]

Cronbach, L. J. (1988). Five perspectives on validity argument. In H. Wainer & H. I. Braun (Eds.), Test validity (pp. 3–17). Hillsdale, NJ: Lawrence Erlbaum.

Eisenhart, M. A., & Howe, K. R. (1992). Validity in educational research. In M. LeCompte, W. Milroy, & J. Priessle (Eds.), The handbook of qualitative research in education (pp. 642–680). San Diego, CA: Academic Press.

Gorin, J. S. (2007). Test design with cognition in mind. Educational Measurement: Issues and Practice, 25(4), 21–35. doi:10.1111/j.1745-3992.2006.00076.x

Haertel, E. H., & Herman, J. L. (2005). A historical perspective on validity arguments for accountability testing. In J. L. Herman & E. H. Haertel (Eds.), Uses and misuses of data for educational accountability and improvement (NSSE 104th., pp. 1–34). Malden, MA: Wiley-Blackwell.

Holland, P. W. (1986). Statistics and causal inference. Journal of the American Statistical Association, 81(396), 945–960. doi:10.2307/2289069

Kane, M. T. (1992). An argument-based approach to validity. Psychological Bulletin, 112(3), 527–535. doi:10.1037/0033-2909.112.3.527

Leighton, J. P., & Gierl, M. J. (2004). Defining and evaluating models of cognition used in educational measurement to make inferences about examinees’ thinking processes. Educational Measurement: Issues and Practice, 26(2), 3–16. doi:10.1111/j.1745-3992.2007.00090.x

Linn, R. L., & Baker, E. L. (1996). Can performance-based student assessments be psychometrically sound? Performance-based student assessment: Challenges and possibilities (pp. 84–103). Chicago, IL: The University of Chicago Press.

Messick, S. (1988). The once and future issues of validity: Assessing the meaning and consequences of measurement. In H. Wainer & H. I. Braun (Eds.), Test validity (pp. 33–45). Hillsdale, NJ: Lawrence Erlbaum.

Michell, J. (2009). Invalidity in validity. In R. W. Lissitz (Ed.), The concept of validity: Revisions, new directions, and applications (pp. 111–133). Information Age Publishing.

Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference (p. 623). Boston, MA: Houghton Mifflin. [Probably Chapters 1-3 and 11, if not more.]

Shepard, L. A. (1993). Evaluating test validity. Review of Research in Education, 19(1), 405–450.

Shepard, L. A. (1997). The centrality of test use and consequences for test validity. Educational Measurement: Issues and Practice, 16(2), 5–24. doi:10.1111/j.1745-3992.1997.tb00585.x

Zumbo, B. D. (2009). Validity as contextualized and pragmatic explanation, and its implications for validation practice. In R. W. Lissitz (Ed.), The concept of validity: Revisions, new directions, and applications (pp. 65–82). Information Age Publishing.

Thankfully, some of these papers I've read recently for my Advances in Assessment course so the amount of reading I have to do is appreciably less than it might look. In my typical fashion, I'll study these in chronological order with the hopes that I get a sense for how the field has evolved its thinking and practice regarding these ideas over the past several decades.

Although I have little other graduate school experience to compare it to, I feel like this reading list is representative of what sets a PhD apart, particularly one earned at an R1 university. It's not necessarily glamorous, and its relevance to the day-to-day teaching and learning in classrooms might not be immediately obvious. But without attending to issues like validity and causal inference, we have a much more difficult time being sure about what we know and how we're using that knowledge. Issues of validity should be at the heart of any assessment or measurement, and when they're attended to properly we greatly improve our ability to advance educational theories and practice.

RYSK: Greeno, Pearson, & Schoenfeld's Implications for NAEP of Research on Learning and Cognition (1996)

This is the 14th in a series describing "Research You Should Know" (RYSK).

You might have read my recent post about Lorrie Shepard's 2000 article The Role of Assessment in a Learning Culture and assumed she focused on classroom assessment because changing large-scale, standardized assessments was a lost cause. Think again. By that time, an effort to integrate new theories of learning and cognition into the NAEP was already underway, traceable back to a 1996 report titled Implications for NAEP of Research on Learning and Cognition written by by James G. Greeno, P. David Pearson, and Alan H. Schoenfeld. For years Greeno has been recognized as one of education's foremost learning theorists, while Pearson and Schoenfeld are highly-regarded experts in language arts and mathematics education, respectively.

The National Assessment of Educational Progress, sometimes called "The Nation's Report Card," has been given to students in various forms since 1969. Unlike the high-stakes assessments given by states to all students, the NAEP is given to samples of 4th, 8th, and 12th grade students from around the country, and the use of matrix sampling means no student ever takes the entire test. The goal of the NAEP is to inform educators and policymakers about performance and trends, and details about how different NAEP exams try to achieve this are described in depth at the NAEP website.

Greeno et al. tried to answer two main questions in their report: (a) Does the NAEP inform the nation "about significant aspects of the knowing and learning" (p. 2) in math and reading, and (b) What changes in NAEP would make it a better tool for informing the nation about the performance and progress of our educational system? The authors acknowledge the tradition with what they call differential and behaviorist perspectives on learning, and focus more of their attention on the ability to assess cogntiive and situative perspectives, which have strong theoretical foundations but hadn't been reflected in most large-scale assessments.

Concisely, the report says the "key features of learning in the cognitive perspective are meaningful, conceptual understanding and strategic thinking" and that the "key feature of learning in the situative perspective is engaged participation with agency" (p. 3, emphasis in original). Greeno et al. say that if students are engaged in learning activities that reflect these perspectives, then the NAEP should try to capture the effects of those experiences.

One of the main reasons I'm writing about this report is because it gives me another chance to describe current learning perspectives that go beyond the simpler "behaviorism vs. constructivism" argument I knew as a teacher and heard from others. This report does this well without burdening the reader with all the gory details that learning theorists grapple with as they try to push these theories even further. So here's my summary of their summaries of each perspective:

Differential

This perspective accepts the assumption that "Whatever exists, exists in some amount and can be measured" (p. 10). For knowledge, that "whatever" is referred to as a trait, and different people have traits in different amounts. Evidence of traits can be detected by tests, and the correlation of different tests supposedly measuring the same trait is an indication of our confidence in our ability to measure the trait. Because the person-to-person amount of a trait is assumed to be relative, it's statistically important to design tests where few people will answer all items correctly or incorrectly.

Behaviorist

Behaviorism assumes that "knowing is an organized collection of stimulus-response associations" (p. 11). To learn is to acquire skills (usually and best in small pieces) and measuring learning is seen as an analysis of behaviors which can be decomposed into responses to stimuli. Behaviorism's influence on curriculum is seen when behavioral objectives are organized as a sequence building bigger ideas out of smaller, prerequisite objectives.

Cognitive

The cognitive perspective primarily focuses "on structures of knowledge, including principles and concepts of subject-matter domains, information organized by schemata, and procedures and strategies for problem solving and reasoning" (p. 12). Learners actively construct their knowledge rather than accept it passively, and conceptual understanding is not just the sum total of facts. The early part of the cognitive revolution was reflected in the math and science reforms of the 1950s and 1960s, while Piagetian ideas and research on student understanding have pushed the perspective further. Assessments need to determine more than right and wrong answers, and research involving think-aloud protocols, student interviews, eye-tracking studies, and patterns of responses have yielded better theories about how to assess for student understanding.

Situative

The situative perspective is a social view of learning focused on "interactive processes in which people participate in practices that are organized by the societies and communities they belong to, using the technologies and natural resources in their environments" (p. 14). Knowing is no longer in the head -- instead it is seen as participation in a community, and learning is represented by increased and more effective participation. John Dewey took parts of this perspective in the early 20th century, but we owe much of the theory to Lev Vygotsky, whose work in the 20s and 30s in the Soviet Union eventually emerged and has heavily influenced learning science since the late 1970s. The situative perspective is more readily applied to interactions between people or between people and technology (which is seen as a cultural artifact with social roots), but even solitary learners can be assessed with the situative perspective if we focus on "the individual's participation in communities with practices, goals, and standards that make the individual's activity meaningful, either by the individual's adoption of or opposition to the community's perspective" (p. 14). The influence of the situative perspective on curriculum and classrooms is most easily seen in the focus on student participation, project work, small-group discussions, and authentic work in subject-area disciplines.

In summary, achievement in each perspective can be described as:
Differential/Behaviorist
- "progress a student has made in the accumulation of skills and knowledge" (p. 16)
Cogntive
- a combination of five aspects (pp. 16-18):
  1. Elementary skills, facts, and concepts
  2. Strategies and schemata
  3. Aspects of metacognition
  4. Beliefs
  5. Contextual factors
Situative
- a combination of five aspects (pp. 19-21):
  1. Basic aspects of participation
  2. Identity and membership in communities
  3. Formulating problems and goals and applying standards
  4. Constructing meaning
  5. Fluency with technical methods and representations

What Does This Mean for the NAEP?

Greeno et al. declared that the NAEP was "poorly aligned" (p. 23) with the cognitive perspective. It hadn't captured the complexity of student knowledge and they recommended a greater focus on problems set in meaningful contexts and tasks that reflected the kind of knowledge models and structures theorized in the research. As for the situative perspective, Greeno et al. went so far to say that what the NAEP had been measuring was "of relatively minor importance in almost all activities that are significant for students to learn" (p. 27). Whereas the situative perspective focuses on participation in a particular community or knowledge domain, it's impossible to escape the reality that on the NAEP, the domain is test-taking itself, a "special kind of situation that is abstracted from the variety of situations in which students need to know how to participate" (pp. 28-29). Measuring learning from the situative perspective would require a complicated set of inferences about a student's actual participation practices in an authentic domain, and the technical limitations of the NAEP limits our ability to make those inferences.

The report continues with specific details about how we might measure learning in language arts and mathematics with the NAEP from both a cognitive and situative perspective. In the conclusion, the authors first recommend some systemic changes: First, NAEP needed more capacity for attending to the long-term continuity of the test and its design. Given how important NAEP is for measuring longitudinal trends, we can't change it without a careful study of how to compare new results to old. Second, the authors wanted a national system for evaluating changes in the educational system. The NAEP alone can't tell us everything we need to know about the effectiveness of educational reforms.

As for recommendations for the test itself, Greeno et al. emphasized the need to align the assessment with ongoing research, especially in the cognitive perspective. Instead of planning for NAEP tests one at a time and contracting out various work, the development process needed to become more continuous with particular sustained attention given to progress in the cognitive and situative dimensions. More ambitiously, the authors recommended a parallel line of test development to begin establishing new forms of assessment that might capture learning in these newer perspectives. This is a critical challenge because while we know the least about assessing from the situative perspective, the situative is often the perspective that frames our national educational goals. The NAEP can't measure progress to situative-sounding goals without better measurement of learning from a situative perspective.

It has now been 12 years since the release of this report. I don't know how Greeno et al.'s recommendations have specifically been followed, but there is good news. If you read most any of the current NAEP assessment frameworks, you can find evidence of progress. The frameworks have changed to better measure student learning, particularly from the cognitive perspective. Some frameworks honesty address the difficulty in measuring the situative perspective using an on-demand, individualized, pencil-and-paper (but increasingly computer-based) test. (See Chapter One of the science framework, for example.) Will we see any radical changes any time soon? I doubt it. The information we get about long-term trends from the NAEP requires a certain amount of stability. Given the onset of new national consortia tests based on the Common Core State Standards, I think the educational system will get its fill of radical change in the next 3-5 years. With that as the comparison, we all might contently appreciate the stability and attention to careful progress reflected in the NAEP.

References

Greeno, J. G., Pearson, P. D., & Schoenfeld, A. H. (1996). Implications for NAEP of research on learning and cognition (p. 84). Menlo Park, CA.

RYSK: Shepard's The Role of Assessment in a Learning Culture (2000)

This is the 13th in a series describing "Research You Should Know" (RYSK).

In her presidential address at the 2000 AERA conference, Lorrie Shepard revealed a vision for the future of educational assessment. That message turned into an article titled The Role of Assessment in a Learning Culture, and its message is still very much worth hearing today. Lorrie Shepard remains a globally-respected expert in assessment, psychometrics, and their misuses, and I'd think she was totally awesome even if she wasn't my boss.

Shepard is often present for debates about large-scale testing, but this paper focuses on classroom assessment -- the kind, says Shepard, "that can be used as a part of instruction to support and enhance learning" (p. 4). Shepard does this by first explaining a historical perspective, then describing a modern view of learning theories, then envisioning how new assessment practices could support those theories. Impressively, she does this all in just 11 well-written pages. (In fact, given that the paper is available on the web, I wouldn't blame you at all for skipping this summary and just reading the article for yourself.)

History

Shepard highlights several major themes from history that have continued to drive our assessment practices. One is the social efficiency movement, which "grew out of the belief that science could be used to solve the problems of industrialization and urbanization" (p. 4). While this movement might have helped our economic and educational systems scale rapidly (think about Ford and the assembly line), social efficiency carries with it a belief that people have a certain innate (and largely fixed) set of capabilities, and our society operates its most efficiently when we measure people and match their capabilities to appropriate education and employment. For example, students were often given IQ tests to determine if their future path should lie on a particular academic or vocational track.

The dominant learning theories of the early and mid-1900s were associationism and behaviorism, both of which promoted the idea that learning was an accumulation of knowledge that could be broken into very small pieces. Behaviorism was also tied closely to theories of motivation, as it was believed learning was promoted when knowledge was made smaller and opportunities for positive reinforcement for learning were made greater. Much of the assessment work related to these beliefs can be traced back to Edward Thorndike, considered to be the father of scientific measurement and earliest promoter of "objective" testing. It's been 100 years since Thorndike was elected president of the American Psychological Association, and decades since his ideas seriously influenced the leading edges of learning theory. Still, as most anyone who works in schools or experienced a traditional education can attest, ideas of social efficiency and behaviorism are still evident in schools -- especially in our assessment practices.

Together, the theories of social efficiency, scientific measurement, and beliefs about intelligence and learning form what Shepard sees as the dominant 20th-century paradigm. (See page 6 of the paper for a diagram.) It's important to begin our discussion here, says Shepard, because "any attempt to change the form and purpose of classroom assessment to make it more fundamentally a part of the learning process must acknowledge the power of these enduring and hidden beliefs" (p. 6).

Modern Theories

In the next section, Shepard describes a "social-constructivist" framework that guides modern thought on learning:

The cognitive revolution reintroduced the concept of mind. In contrast to past, mechanistic theories of knowledge acquisition, we now understand that learning is an active process of mental construction and sense making. From cognitive theory we have also learned that existing knowledge structures and beliefs work to enable or impede new learning, that intelligent thought involves self-monitoring and awareness about when and how to use skills, and that "expertise" develops in a field of study as a principled and coherent way of thinking and representing problems, not just as an accumulation of information. (pp. 6-7)

These ideas about cognition are complimented by Vygotskian realizations that the knowledge we construct "is socially and culturally determined" (p. 7). Unlike Piaget's view that development preceded learning, this modern view sees how development and learning interact as social processes. While academic debates remain about the details of cognitive vs. social (and vs. situative vs. sociocultural vs. social constructivist vs. ...), for practical purposes these theories can coexist and are already helping teachers view student learning in ways that improve upon behaviorism. However, Shepard says, since about the 1980s this has left us in an awkward state of using new theories to inform classroom instruction, while still depending on old theories to guide our assessments.

Improving Assessment

If we wish to make our theories of assessment compatible with our theories of learning, Shepard says we need to (a) change the form and content of assessments and (b) change the way we use and regard assessment in classrooms. Some of the potential changes in form are already familiar to most teachers, such as a greater use of open-ended performance tasks and setting assessment tasks in real-world contexts. Furthermore, Shepard suggests that classroom routines and related assessments should reflect the need to socialize students "into the discourse and practices of academic disciplines" (p. 8) as well as foster metacognition and important dispositions. Shepard does not go into much more detail here because others have already given attention to these ideas, but gives us this simple yet powerful idea (p. 8):

"Good assessment tasks are interchangeable
with good instructional tasks."

Next Shepard pays special attention to negative effects of high-stakes testing. Shepard could be called a believer in standards-based education, but recognizes how "the standards movement has been corrupted, in many instances, into a heavy-handed system of rewards and punishments without the capacity building and professional development originally proposed as part of the vision (McLaughlin & Shepard, 1995)" (p. 9). Unfortunately, Shepard's predictions have held true over the past 12 years: we've seen test scores distorted under political pressure, a corruption of "teaching to the test," and a trend towards the "de-skilling and de-professionalization of teachers" (p. 9). What's worse might be a decade of new teachers who've learned to "hate standardized testing and at the same time reproduce it faithfully in their own pre-post testing routines" (p. 10) because they've had such little exposure to better forms of assessment.

For the rest of the article, Shepard focuses on how assessment can and should be used to support student learning. First, classrooms need to support a learning culture where "students and teachers would have a shared expectation that finding out what makes sense and what doesn't is a joint and worthwhile project" (p. 10). This means assessment that is more informative and reflective of student learning, one where "students and teachers look to assessment as a source of insight and help instead of an occasion for meting out rewards and punishments" (p. 10). To do this, Shepard describes a set of specific strategies teachers should use in combination in their classrooms.

Dynamic Assessment

When Shepard wrote this article, formal ideas and theories about formative assessment were still emerging and the field had yet to settle on some of the language we now use. But if you're at all familiar with formative assessment, Shepard's description of "dynamic" assessment will sound familiar: teacher-student interactions continuing through the learning process rather than delayed until the end, with the goal of gaining insight about what students understand and can do both on their own and with assistance from classmates or the teacher.

Prior Knowledge

The idea of a pre-test to see what students know before instruction begins is not new, but Shepard says we should recognize that traditional pretests don't usually take account of social and cultural contexts. Because students are unfamiliar with a teacher's conceptualization of the content prior to instruction (and vice versa), scores might not accurately reflect students' knowledge as well as, say, a conversation or activity designed to elicit the understandings students bring to the classroom. Also, as Shepard has frequently observed, traditional pre-testing often doesn't significantly affect teachers' instruction. So why do it? Instead, why not focus on building a learning culture of assessment: "What safer time to admit what you don't know than at the start of an instructional activity?" (p. 11)

Feedback

The contrast in feedback under old, behaviorist theories and newer, social-constructivist theories is clear. Feedback under old theories generally consisted of labeling answers right or wrong. Feedback under new theories takes greater skill: teachers need to know how to ignore student errors that aren't immediately relevant to the learning at hand, while crafting questions and comments that force the student to question themselves and any false knowledge they might be constructing. (See Lepper, Drake, and O'Donnell-Johnson, 1997, for more on this.)

Transfer

While it is our hope that our students will be able to generalize the specific knowledge they have learned and apply it to other situations, our ability to accurately research and make claims about knowledge transfer turns out to be a pretty tricky business. Under a strict behaviorist perspective, it was appropriate to believe that each application of knowledge should be taught separately. Many of our current theories support an idea of transfer, and evidence shows that we can help students by giving them opportunities to see how their knowledge reliably works in multiple applications and contexts. So while some students might not agree, Shepard says teachers should not "agree to a contract with our students which says that the only fair test is one with familiar and well-rehearsed problems" (p. 11).

Explicit Criteria

If students are to perform well, they need to have clear guidance about what good performances look like. "In fact, the features of excellent performance should be so transparent that students can learn to evaluate their own work in the same way their teachers would" (p. 11). This reinforces ideas of metacognition and, perhaps more importantly, fairness.

Self-Assessment

There are cognitive reasons to have students self-assess, but other goals are to increase student self-responsibility and make teacher-student relationships more collaborative. Students who self-evaluate become more interested in feedback from others, are more aware of standards of excellence, and take more ownership over the learning process.

Evaluation of Teaching

This is another idea now heavily intertwined with formative assessment, but Shepard takes it one step farther than I normally see it. Instead of just using assessment to improve one's teaching, Shepard recommends that teachers be transparent about this process and "make their investigations of teaching visible to students, for example, by discussing with them decisions to redirect instruction, stop for a mini-lesson, and so-forth" (p. 12). This, Shepard says, is critical to cultural change in the classroom:

If we want to develop a community of learners -- where students naturally seek feedback and critique of their own work -- then it is reasonable that teachers would model this same commitment to using data systematically as it applies to their own role in the teaching and learning process. (p. 12)

Conclusion

Shepard admits that describing this new assessment paradigm is far easier than it is to implement in practice. It relies on a great deal of teacher ability and confronting some long-held beliefs. Shepard recommended a program of research accompanied by a public education campaign to help citizens and policymakers understand the different goals of large-scale and classroom assessments. Neither the research or educating the public is easy, because both are built upon a history of theories and practice that a new paradigm needs to discard. Perhaps we haven't taken on this challenge with the effort and seriousness we've needed, and I worry that now we're more apt to talk about "learning in an assessment culture" rather than the other way around, as Shepard titled this article. I sometimes wonder if she's considered writing a follow-up with that title, or if she's hoping she'll never have to. I guess the next time it comes up I'll have to ask her.

Math note: This is an article about assessment and not specific to mathematics, but I'd be remiss if I didn't share Shepard's inclusion of one of my all-time favorite fraction problems:


References

Lepper, M. R., Drake, M. F., O'Donnell-Johnson, T (1997). Scaffolding techniques of expert human tutors. In K. Hogan & M. Presley (eds.), Scaffolding student learning: Instructional approaches & issues. Cambridge, MA: Brookline Books.

McLaughlin, M. W., & Shepard, L.A. (1995). Improving education through standards-based reform: A report of the National Academy of Education panel on standards-based educational reform. Stanford, CA: National Academy of Education.

Shepard, L. A. (2000). The role of assessment in a learning culture. Educational Researcher, 29(7), 4–14. doi:10.2307/1176145

Thompson, P. W. (1995). Notation, convention, and quantity in elementary mathematics. In J. T. Sowder & B. P. Schappelle (Eds.), Providing a foundation for teaching mathematics in the middle grades (pp. 199-221). New York: State University of New York Press.

Sorting Out the Summative: When Standards-Based Grading Meets the End of the Semester

Source: Wikipedia

Many teachers who choose to use standards-based grading eventually find themselves facing the reality of their school's grading policies and tradition: the expectation of final, summative grades that are reported as percentages and letters. So regardless how hard you try to focus on quality feedback instead of grades all semester long (for good reason), there comes a time when, for reasons probably beyond your control, you have to turn levels and descriptions of student understanding into numbers. This is SBG's "Monday Morning Problem" that doesn't always get addressed in theory. But this week is finals week for my basic statistics students, so for me the time has come to convert standards-based formative grades into a summative grade, including calculating final exam grades. Here I'll try to describe the two steps I'll take to calculate my students' grades: (a) conversion of their formative scores into a summative score and (b) scoring and inclusion of the final exam into their semester grades.

Formative to Summative
Besides giving students a lot of written and verbal feedback about where they should try to improve, I've been using the simplest of measures to record their performance on class objectives: either students (a) "get it," (b) "sort of get it," or (c) "dont' get it/haven't demonstrated it." You could think of these as "green light," "yellow light," and "red light," respectively. I've tried discerning more levels of understanding in a gradebook and it only seems to lead to confusion and indecision (both for me and students), so I'm sticking to three levels, as suggested in Her & Webb (2004). If I need more detail, I can always go back to the copies of the work students have submitted and the comments I've made.

The gradebook we have for class is pretty primitive and as far as I can tell it only accepts numbers, so I mark my three levels as either a 2, a 1, or a 0. It doesn't take much explaining to students that a 1 shouldn't be viewed as "out of two" and therefore worth 50%. I do tell them, though, that in order to receive credit for the course they should average a 1 across all objectives. In other words, you can't pass the class without an average of at least some understanding of every objective.

Around here and in many other places, 70% seems to be the low end of passing grades. (We're not messing with Ds.) So if a student with all 1s should get at least a 70%, and a student with all 2s maxes out at 100, and we choose a linear function between the two, the "conversion formula" to percentages is simply:

percentage = 30 * objective score average + 40

If you feel a little dirty at this point because you know you just reduced all the various skills, knowledge, and abilities of your students into a single number, I say join the club. If you didn't feel that way I wouldn't have expected you to be using standards-based grading to begin with.

A "No Surprises" Approach to Final Exam Grades
Designing a final exam is often tricky business. It can't possibly assess everything in the course, but we generally want it to include the major topics and themes for the class and be possible to complete in the time allowed. We also have to think about difficulty. Trust me, your students are!

Teachers want their finals to be challenging, but they don't want to have that sinking feeling as they grade the exams that maybe the test was too hard. For whatever reason, sometimes students perform poorly and averaging the final exam grade into their other grades will look like a disaster. But ask yourself: What am I more confident in, my careful judgments of students' ability as demonstrated over an entire semester, or a fleeting, one-time judgement of students' ability on a single assessment during the most stressful time of the year? If you're using standards-based grading, I already know how you'll answer that question. If not, consider this example: I have a student who I know can do stats. She's turned in good work. She's asked quality questions. We've had good discussions. But I also know she has seven final exams this week. I still think she'll do fine, but I'll understand if she's not at her best. And I need a grading system that reflects that understanding.

In order to free myself to still give challenging, yet reasonable, assessments, without risking any huge surprises when grades are calculated, I perform a little statistical magic that ensures that the distribution of final grades has the same center and spread of class grades before the final. I'm sure many of you try "curving" your exam scores some other way, such as letting the top score count as the total possible, or even having a pre-set distribution in mind of how many As, Bs, Cs, etc. you'll allow (which is not a good idea, generally, for reasons described by Krumboltz & Yeh, 1996). I prefer my method because it accounts for the distribution of grades, not just the top score, and the distribution is determined by the students, not arbitrarily by me. Allow me to demonstrate with a couple examples.

Suppose before the final the average percentage grade is 85 and the standard deviation of those grades is 10. Then I grade my final exams and find that the average final exam grade is 60 with a standard deviation of 18. Ouch. But don't worry -- statistics will come to our rescue.

Provided you know a little basic descriptive statistics, the conversion is simple. For each student's final exam score, find out how many standard deviations above or below the mean they scored on the final (their final exam z-score), and match that with the same number of standard deviations above or below the mean they'd fall on the pre-final grade distribution (their pre-final z-score). Consider the following students and the class and exam statistics above:

  • Suppose Student A scores a 51 on the final exam. That's 0.5 standard deviations below the mean. (51 - 60 = -9, and -9/18 = -0.5.) So where is 0.5 standard deviations below the mean on the pre-final distribution? If that mean is 85 and the SD is 10, then 0.5 standard deviations below the mean is 80. So I record an 80 for that student instead of a 51.
  • Suppose Student B scores a 75 on the final exam. That's about 0.83 standard deviations above the mean. (75 - 60 = 15, and 15/18 = 0.83.) So where is 0.83 standard deviations above the mean on the pre-final distribution? About 8.3% above an 85, so I record their exam grade as a 93.3.
  • Suppose Student C scores a 60 on the final exam. That's the same as the mean, so zero standard deviations above or below. That conversion is super-easy: their final exam grade is the mean of the pre-final mean, an 85.
For an example of how to set up a spreadsheet to do this, see https://docs.google.com/spreadsheet/ccc?key=0Anne5Z-jCkqhdDVtemkyaGhnRWFfclJoa0dIUVQ5RVE. I recommend making a copy of it for yourself and seeing what happens as you change values.

This is not a perfect system (and comments about its imperfections are welcome in the comments), but it does take away the element of surprise if the final exam happens to be way too easy or too difficult, or if other circumstances prevent grades from working out the way you'd expect. Yes, this is a norm-referenced system instead of a criterion-referenced system, meaning that the grades students earn on the final is measured largely as how they compare to their classmates and the class average. The good news is this: both the teacher and the students have an incentive before the final to master as many objectives as possible, and that is criterion-referenced. A high pre-final average helps everyone get a high final exam average, and a small pre-final standard deviation minimizes variability in final exam scores.

References

Her, T., & Webb, D. C. (2004). Retracing a path to assessing for understanding. In T. A. Romberg (Ed.), Standards-based mathematics assessment in middle school: Rethinking classroom practice (pp. 200-220). New York, NY: Teachers College Press.

Krumboltz, J. D., & Yeh, C. J. (1996). Competitive grading sabotages good teaching. Phi Delta Kappan, 78(4), 324-326. Retrieved from http://www.jstor.org/stable/20405782

A Quick-and-Dirty Guide to Fighting the Math Wars

I just posted this to a reply to a post by David Wees on Google+, but I thought it might be useful to some if it had some permanence here.

I've been in and out of "Math Wars" debates for 10+ years, and I find it's helpful to examine the issue at a more granular level. Here's a quick list of questions I jotted down:

What is your definition of mathematics? (Someone who answers, "It's a subject you learn in school" may have very different views from someone who answers, "It's a human activity we undertake to solve problems relating to number and shape.")

What is your philosophy of mathematics? (A Hardyist and a Mathematical Maoist have very different views, as do a Platonist and a Formalist. And for all the consistency in mathematics, this is not something with which we as individuals are necessarily consistent.)

What is our goal for students learning mathematics? (Is it to prepare them for work? For more school? To gain an appreciation of mathematics? For mental exercise?)

How should we assess mathematics? (Often when we claim that students do or do not perform well in mathematics, we are basing those claims on an assessment that may not embrace a balanced view of the issues above. Or, failing that, we make those claims without regard to the biases of the assessment.)

What learning theories do we use, and how do we use them? (A difficulty with learning theories is that in most all cases we can design curriculum and pedagogy around them that show they work -- at least to a degree. The workings of the human brain aren't easy to study, explain, or leverage in a classroom.)

How do we perceive "failure" or "success" of practices of the past? (I fear sometimes we stereotype certain historical movements, such as "New Math" and the "Back to Basics" movement, and we falsely assume that those movements were implemented in every classroom with high fidelity. We also sometimes forget that as time has passed, we are trying to teach higher and higher levels of mathematics to more and more students.)

How do we avoid false dichotomies? (False dichotomies were addressed in that article and Zwaagstra was wise to try to avoid them. But it's such an *easy* trap to fall into! [I've probably done it here without realizing it.] For example, he cited a paper by Alfieri, et al. (2011) that claimed through meta-analysis that "unassisted discovery does not benefit learners." But why would a well-trained constructivist teacher believe discovery should be unassisted? That's the same as assuming that a traditional teacher only has students listen to lectures and work problems in isolation. No teacher or student thrives exclusively on either. Interestingly, Zwaagstra in the next sentence says learners should be "scaffolded," an idea developed by Jerome Bruner in support of learning in a social constructivist environment.)

What skills, abilities, and philosophies do we believe teachers need to be successful? (I'm not sure we fully comprehend the effects on the received curriculum when it's taught by a teacher with skills, abilities, and philosophies that run counter to those supported by the curriculum. In such cases it's easy to misplace blame for poor outcomes.)

I'm sure there are more that I could add, but I strongly recommend that anyone who is serious about this debate to take on these issues one by one. Only if there is some agreement, or at least some sympathy and understanding, on these issues does it become truly productive to talk about "what works."

RYSK: Butler's Effects on Intrinsic Motivation and Performance (1986) and Task-Involving and Ego-Involving Properties of Evaluation (1987)

This is the third in a series of posts describing "Research You Should Know" (RYSK).

As teachers, we care not only about what students learn, but why students learn. In a perfect world, we would all agree on what's important to learn and do and be self-motivated to learn and do those things. But our world isn't perfect, and students are motivated to learn and do things for many reasons. Understanding those reasons is important if we want students to be properly motivated and to perform well with the right attitude.

Ruth Butler earned her Ph.D. in developmental psychology from the Hebrew University of Jerusalem in 1982 and was a relatively new professor there when she teamed with veteran educational psychologist Mordecai Nisan, whose career includes time spent at the University of Chicago, Harvard University, The Max Planck Institute for Human Development, and Oxford University. Together, they sought to build upon studies that compared extrinsic vs. intrinsic motivation and positive vs. negative feedback, looking specifically at how different feedback conditions -- ones that can be manipulated by teachers -- affect students' intrinsic motivation.

For their 1986 paper, Effects of No Feedback, Task-Related Comments, and Grades on Intrinsic Motivation and Performance, Butler and Nisan expected that students who received feedback in the form of simple positive and negative comments (without elements of praise or grading/ranking) would remain motivated, while students who received grades or no feedback would generally become less motivated. To test this hypothesis, Butler and Nisan randomly assigned 261 sixth grade students to one of three groups. They gave the students two types of tasks: Task A was a quantitative "speed" task where students created words from the letters of a longer word, while Task B was a qualitative "power" task that encouraged problem solving and divergent thinking.

Butler and Nisan conducted three sessions with the groups:
  • Session 1: Students performed the tasks.
  • Session 2: Two days after Session 1 the tasks were returned.
    • Students in the first group got comments in the form of simple phrases such as, "Your answers were correct, but you did not write many answers," or "You wrote many answers, but not all were correct."
    • Students in the second group got numerical grades that were computed to reflect a normal distribution of scores from 30 to 100.
    • Students in the third group got their work returned with no feedback.
    After students reviewed their previous work, they were given new tasks and told to expect the same type of feedback when they returned for Session 3.
  • Session 3: Two hours after Session 2 students again reviewed their work and feedback (except for the third group, who got no feedback) from Session 2 and then got a third set of tasks. Students were asked to complete the tasks and were told that they would not get them back. The session ended with a survey of students attitudes towards the tasks.
When Butler and Nisan compared the students' average performance on the tasks in Session 1, all three groups scored approximately the same. That changed in Session 3. On Task A, students receiving comments and grades scored about the same in Session 3 (with an edge to the comments group for the creation of long words), but students receiving no feedback did far worse. For Task B, students receiving comments did significantly better than students who received grades or no feedback, who performed about the same. The only students doing well in Session 3 -- in fact, the only students consistently scoring higher, on average, in Session 3 than in Session 1 -- were the students who received comments.

The survey also showed attitudinal benefits for the comments group, who indicated they found the tasks more interesting and were most willing to do more tasks. Furthermore, 70.5% of students who received comments attributed their effort to their interest in the tasks, compared to only 34.4% of those graded and 43.4% of those receiving no feedback. Only 9% of students receiving comments said their effort was due to a desire to avoid poor achievement, compared to 26.7% of students receiving grades and 9.6% of the no feedback group. Lastly, 86.3% of students receiving comments wanted to keep receiving comments, while only 21% of the graded group wanted to keep receiving grades. The vast majority of graded students, 78.9%, wanted comments. The no feedback group was roughly split 50/50 on wanting comments or grades. None wanted to keep receiving no feedback.

Butler modified this study for her 1987 paper Task-Involving and Ego-Involving Properties of Evaluation: Effects of Different Feedback Conditions on Motivational Perceptions, Interest, and Performance. In it, Butler adapted a theory of task motivation used by Nicholls (1979, 1983, as cited in Butler, 1987):
  • Task involvement: Activities are inherently satisfying and individuals are concerned with developing mastery in relation to the task or prior performance.
  • Ego involvement: Attention is focused on ability compared to the performance of others.
  • Extrinsic motivation: Activities are undertaken as a means to some other end, and the focus is that goal, not mastery or ability.
Butler believed comments would promote task involvement, while grades would promote ego involvement. While both of these can be seen as intrinsic motivation, a third type of feedback needed to be considered: praise. Previous research on praise had gotten mixed results, possibly because researchers hadn't considered if the praise was task- or ego-involved. Butler's study would include ego-involving praise using comments designed to focus a student's attention on their self-worth and not on the task. Therefore, Butler hypothesized that praise and grades would generate similar results, results less desirable than task-involved comments.

The study was similar to the 1986 study, with 200 fifth and sixth graders split into four groups (comments, grades, praise, and no feedback) with subgroups in each for high- and low-achieving students. Tasks were administered in three sessions, with no feedback given after the third session. The tasks this time were divergent thinking tasks, used as Task B in the 1986 study. Praise would come in the form of a single phrase: "Very good." An attitude survey was given after Session 3.

As Butler expected, comments promoted task-involved attitudes while grades and praise promoted ego-involved attitudes. Students' interest in the tasks after Session 3 was higher for the comments group than for the grades, praise, and no feedback groups combined. Students who received praise showed more interest than those who received grades. As for performance, the comments group easily performed the best in Session 3, with both high and low groups improving their scores over Session 1, while all other groups performed about the same or worse compared to their Session 1 performance.

So what does this mean?

As a teacher who struggled with assessment and grading, it was Butler's work that most inspired me to start this RYSK blog series. Despite these results being 25 years old, there's not much evidence that Butler's findings have had a serious impact on the practice of most teachers. I suspect that few teachers know about Butler's work -- I certainly didn't. I was wrapped up in the scores and grades game, not fully aware of the impact those scores were having on my students. I knew it wasn't working, but I didn't have this kind of theoretical knowledge to support a significant change in my practice.

I'm not suggesting that we should suddenly demand a grade-free world. That's just not a realistic thing to expect given where we are now. What I would like to suggest is that teachers become more aware of how the feedback they give affects student motivation, and be careful to focus on task-involved comments whenever possible. Because students aren't likely to get this kind of feedback from standardized tests or computer-based learning systems (i.e., Khan Academy), it takes a teacher's touch to carefully craft the kind of feedback a student needs to sustain their motivation.

References

Butler, R., & Nisan, M. (1986). Effects of no feedback, task-related comments, and grades on intrinsic motivation and performance. Journal of Educational Psychology, 78(3), 210-216. doi:10.1037/0022-0663.78.3.210

Butler, R. (1987). Task-involving and ego-involving properties of evaluation: Effects of different feedback conditions on motivational perceptions, interest, and performance. Journal of Educational Psychology, 79(4), 474-482.

2004-2006: My Adventures in Standards-Based Grading (And Why I Stopped)

(cc): NASA
Strap yourself into the wayback machine, boys and girls, because we're going back in time five whole years. Life was different then: the U.S. was engaged in wars overseas, unprecedented disaster had struck our Gulf Coast, and teachers struggled to adapt to an assessment-centric school culture. It was, like, totally different than things are now.

In 2003 I began my teaching career at a medium-sized high school in Southern Colorado. My first year, as it is for many teachers, wasn't much more than a fight for survival. I spent much of my second year learning how to become something other than the teachers who taught me, and part of that meant tinkering with my assessments. By my third year, I was ready to collaborate with the school's other four math teachers in a concerted effort to improve our assessment and grading practices. I'm not sure I knew what an ideal assessment and grading system should look like, but I knew I wanted something other than the traditional quiz/test, either-you-got-it-or-you-don't system.

I think there's a reason math teachers in particular get so heavily invested in assessment and grading. Numbers are our friends. We trust them. We can sort them, scale them, manipulate them, and summarize them in ways that reveal certain truths. I was determined - perhaps obsessed - with finding a grading system that was accurate and fair. To me "accurate" meant "students earn the grade they deserve" and "fair" meant "objective and unbiased." I think I really believed that if I could just find the right scales and weights, the math would solve my assessment problems. 

While I might have been facetious in my opening paragraph, things really were different five years ago. Not many teachers were bloggers, Twitter hadn't been invented, and you couldn't search for #sbg or #sbar hashtags. Sure, standards-based grading existed, but Guskey, Marzano, Wiggins, et. al. sure weren't knocking on my door to tell me about it. I didn't know about SBG and I don't think any of my colleagues or administrators knew about it, either. It's sad that so many good ideas in education struggle to find their way from theory to practice, but that's another story for another day. This story is about my attempts at standards-based grading, my successes, failures, and frustrations I had along the way, and why I reverted back to a traditional grading system.

2004-2005: SBG(ish)
During my first semester of teaching I spent many hours after school designing quizzes and tests, thinking, "This is part of being a first-year teacher. Once I write these tests I'll never have to do this again." HA! Not only did I not reuse any of those assessments, in six years of teaching I hardly reused any of my assessments. Every semester I had new problems, assessment designs, and grading systems that made my old ones look horribly obsolete. As I went into my second year of teaching, I already knew that I wasn't going to be satisfied with a traditional system.

Before I go any further, I want to make this clear: I'm not claiming that I somehow independently invented or discovered standards-based grading. I was making this up as I went along and, as you'll soon read, it didn't necessarily work all that well. I did manage to reorganize my gradebook around concepts and skills instead of dates or arbitrarily-titled ("Chapter 8 Test") assessments. But while that part looked like SBG, sometimes very little else did. In fact, I wouldn't be surprised if you read this and decide I wasn't doing SBG at all.

As I said, I started making SBG-like changes to start my second year, as you can see in this passage from my fall 2004 Algebra 1 syllabus:
Each unit in the text will have several key objectives that should be your focus during that unit. Your score for an objective will usually be established by your performance on a quiz or test and is based on my perception of your understanding. Objectives are graded on a scale of 5 to 10. Think of 10/10 as an A+, 9/10 an A-, 8/10 a B-, and so forth. Objectives not genuinely attempted will be given a zero. If you are dissatisfied with one or more of your objective scores, it is highly recommended that you see me for 1-on-1 help, preferably before or after school. Because the CPM philosophy is mastery over time, you can expect to be quizzed or tested over each objective multiple times, with each time representing an opportunity to raise your objective score.

I got off to a good start by focusing on objectives, but then quickly got bogged down by the point system. SBG isn't about accumulating points. Sure, you'll need a way to record student performance, but a well-implemented SBG system will be formative and focused on feedback, not a point system. By the way, if you ever want to drive yourself into an insane asylum, try a 5-to-10 point grading scale based on perceptions of student understanding. I nearly drove myself crazy, scoring, re-scoring, and re-re-scoring every bit of work out of fear that somebody's 7/10 actually showed the same understanding as a classmate's 8/10. I was far more concerned with ranking and sorting than feedback. Even worse, I would have students who scored 7/10 or 8/10 over and over, earning passing scores for objectives even though I'd never actually seen them get any right answers.

By the following spring (with a new set of classes, like a college schedule) I decided that it would be far better for students to get most problems all right instead of all problems mostly right, so my syllabus now said this:
IT'S ALL ABOUT THE ROADMAP. The objective roadmap is a detailed list of objectives that together make up everything you should learn in the course. The objectives are organized by unit, but each objective will be assessed separately. To pass an objective you need to get 80% or better on the objective test. If you fail an objective, you will have to retake and pass that objective test before moving on to any objectives in the next unit.

(Here are my roadmaps for Algebra 1, Algebra 2, and Business Math.)

Almost every objective test consisted of five problems, and if it wasn't right, it was wrong. I was so sick of agonizing over partial credit I got rid of it entirely. Objectives were added to students' grades on a schedule, and students who fell behind schedule got zeros on objectives they hadn't yet attempted. Zeros have rarely been known to play nice with the statistical mean, so a student who averages 80% on 80% of the objectives (with 0% on the others) is only going to get a 64%. It was setting a high standard, sure, but somehow I convinced myself that I was a one-man-army who could rid the world of grade inflation. Anyhow, I asked students to focus less on the gradebook and more on objective progress charts like this one:
(Those are student numbers on the left. Yes, this class only had 11 students.)
Still, if this was SBG, it was really, really bad SBG. Three major problems stand out. First, I was telling students that there were a bunch of objectives and any one of them could stop them dead in their tracks. (I think I relented on that one not far into the semester.) Second, I was still way too focused on points and percentages, not formative assessment. Third, a student who was "mostly right" on all five test problems got a zero in the gradebook, even if the cause was a small procedural mistake that he/she happened to repeat on all five problems. How's that for demotivation?

Amidst all the problems of this system, there emerged a brilliant, shining light, and her name was Tara. She gave me hope. She was part of a pretty special Algebra 2 class, and I think previous struggles in math had left her without much confidence. She barely spoke, and when she did it was very quiet, and quite often it was apologetic, like "I'm sorry to make you stay after school so I can retake a test." Tara understood the grading system and used it to her advantage, retaking test after test with study sessions in between. I remember a marathon session in the days before the final exam, when Tara was within sight of her goal of passing every objective. I think she stayed after school for five hours until her goal was finally met. Tara, if somehow you happen to read this, I hope you realize you weren't a burden, you were an inspiration. You did and still make me want to be a better teacher.

2005-2006: Professional Learning Communities (PLCs) and Common Assessments
I think every teacher knows that feeling when an administrator latches on to an idea from a conference or meeting they attended and then bring it home for their staff to use. Unfortunately, in a rural district like ours with scant resources, the promotion usually went like this:
This year we're going to implement [insert acronym or edujargon here]. I've talked to a lot of other principals who are doing it and think it's very effective. ["Effective" usually means "Since we didn't make AYP, again, I need to tell the state that we're trying something new. Again."] The principals I talked to sent their teachers to summer workshops and hired a coordinator to guide the implementation of [edujargon], but since we can't afford so much as a box a tissues for this school, I'll just tell you what I remember and we'll make the rest up as we go along.

For 2005-2006, administration was pushing Professional Learning Communities (PLCs) and common assessments. The math department PLC opted to tackle Algebra 1 in the first year, and we began by brainstorming a list of concepts we thought all Algebra 1 students should know. We shuffled our list into eight groups that aligned with our textbook and other materials, and finally scheduled eight test dates spaced evenly throughout the year.

Because of some of the success I experienced the previous year, we opted to grade each concept separately instead of amassing a single score for the whole test. Students who did poorly on a concept would be remediated and given an opportunity to retest until they showed proficiency with each concept. Perhaps more importantly, since all the Algebra 1 teachers were assessing the same objectives, using the same tests, all scored the same way, we now had a basis for comparing our students' performance, and that led to a sharing of tips and techniques for teaching particular skills. Unfortunately our time was limited so I never learned as much from my colleagues as I hoped, but there is certainly an advantage in using the same SGB system across teachers in a department, or even now across schools in our ever-connected world.

We had a schedule change that allowed freshmen to take Algebra 1 all year long. With about a month left in the school year, the progress chart for one of my classes looked like this:
You can detect a problem here: too many students started slacking at the end of the year!
Even though the system was working, it was showing new cracks. Students knew I had a finite (usually just 2 or 3) set of tests for each objective, and if they failed form A, they would carefully memorize their mistakes and the right answers, then pretty much fail form B on purpose so they could then take form A again. (Ugh!) Also, after two years of using this kind of system, I was growing tired of explaining and defending it to students, parents, and administrators. I enjoyed the extra time spent with students who came in after school, but it became increasingly hard to make up for the lost time. Lastly, the paperwork was a challenge - I had to have every form of every test ready at a moment's notice for students who needed to take or retake tests. Because I didn't have my own classroom, this meant hauling around a file box to three or more rooms each day in two different buildings.

One of my biggest dissatisfactions with SBG was the quality of the assessments I was using. During instruction my students worked in groups, delving into big problems in context, problems that required a combination of reasoning and the integration of multiple skills. Because I wanted to assess every skill in isolation, the tests I wrote were nothing like that. Too often they looked like this:
(What, I couldn't think of contexts for these?)

Goodbye, SBG
I was having another major problem at the end of the 2005-2006 school year: because of budget cuts, declining enrollment, and a shuffle of veteran teachers, I found myself out of a job. Even though I knew these to be the reasons, I was pretty hard on myself and thought if I had been a better teacher, then the district would have somehow found a way to keep me.

In the interview for my new job I talked about my assessment practices and quickly got the feeling that they wouldn't work in my new school. The school day was two hours longer and few students would be able to spend time after school. I still hadn't figured out how to measure skills in isolation while wanting students to tackle larger, more comprehensive tasks. I was now the only teacher in the district teaching any of my subjects, so I had no colleagues to work with on common assessments. Frustrated and not wanting to rock the boat, I figured a simple, traditional assessment and grading system would be the easiest and best way forward. I now know that I couldn't have been more wrong! In my bitterness and self-doubt I gave up the progress I had gained over two years of hard work.

If you take anything away from this post, I hope it's this: it's not 2005 anymore. You don't have to stumble around and make the mistakes I did. You don't have to work in isolation like I did. There's a great community of bloggers discussing standards-based and other grading practices, and I haven't met one yet that doesn't want to help their fellow teachers. You might not have access to the research, but some of us do and we enjoy sharing what we think about it. So please, whether you're new to SBG or have been doing it for years, share your experiences and don't hesitate to ask questions!

We're Not All Math and English Teachers

Yesterday the Des Moines Register published an editorial applauding Colorado for reforming its teacher evaluation laws. The editorial goes on to criticize Iowa's reforms, saying the state is "moving too cautiously." Iowa requires teachers to "provide multiple forms of evidence of student learning and growth," but the Register is disappointed that the inclusion of standardized test scores is not required.

When will the public and policymakers realize that not every subject is covered by a standardized test? I feel like this is one of the most overlooked aspects of this argument. We're not all math and English teachers! Colorado spends 3 hours, per subject, each school year to measure every grade 3-10 student's achievement in math, reading, and writing (plus 3 hours for science in grades 5, 8, and 10). In a state that requires a minimum of 1080 hours of student contact time, this minimum of 9 hours of testing represents less than 1% of a school year. If only it felt like so little!

But that's only for math and English. If you want to mandate teacher evaluations based on standardized tests, you need equally rigorous tests for every subject and every teacher. Can you imagine tests for P.E., music, art, or vocational courses? What if schools could only offer classes that were backed by standardized tests?

Let's also consider the extra time required. Suppose students average 7 Carnegie units (credits) per year. At 3 hours per subject and 7 subjects, we're up to 21 hours of testing per year. If you were to require that standardized tests measure student growth, you'd need to test each student both at the beginning of the year and at the end. Doubling the testing takes us to 42 hours, or almost 4% of the year. It still may not sound like much, but students aren't going to be testing 7 hours a day for 6 straight days. If students tested 2 hours daily, the testing schedule would stretch out to 21 days, or about a month of the school year (half at the beginning, half at the end). You thought finals week was bad? Try two weeks!

Unfortunately, I have yet to work in any school that could carry on with regular learning during standardized test times. We tried everything from one test per day to four, mixed with full-length classes meeting on a rotation to all classes meeting on a shortened schedules. No matter how little testing is done, that testing time affects everything else in the day. Can you imagine a month spent like that?

This little rant has come from me, a guy who has almost always been pro-test. They have their place in the educational system and are part of an effective assessment strategy. But when people want every teacher to be measured by their students' standardized test scores, we have to think about the possible ramifications. So be careful what you wish for, Des Moines Register!

Why Don't More Teachers Practice Proper Formative Assessment?

Research that fails to impact practice is a problem in any arena, and education is no exception. Teachers have many reasons for not implementing practices based on research. Too often research is unknown to teachers, locked away in journals that teachers and schools cannot afford. Beyond the cost of access, much of the best research is generally written for higher academic audiences and not easily digested by a busy and distracted practicing teacher. Transforming research into improved practice takes time, effort and patience, and is made easier in a professional community who share ideas and experiences. Teachers are constantly improving their practice, but too often the improvements are driven by personal failures or anecdotal evidence, not the quality results of dedicated educational researchers. This is a crippling inefficiency in the field of education, one that is largely self-imposed and tied to traditional practices held in place by the inertia of our experience.

Research strongly suggests that teachers could improve student learning by using formative assessment. Chapter tests and final exams are summative: they summarize the knowledge and skills a student has acquired, and are generally assigned a fixed grade. In fact, any assignment or task assigned a fixed and lasting grade can be considered summative, at least in part. Formative assessment, in contrast, focuses on improvement rather than final measurement. Teachers use formative assessment to adapt their teaching, and students, equal partners in the process, use feedback and self-monitoring to improve their knowledge and skills.

For any classroom teacher, the concept of formative assessment should be comprehensible and implementation should not be impeded by any significant obstacles. So why don't more teachers practice proper formative assessment? I suggest two simple reasons, reasons that could be eliminated by improved understanding between teachers, administrators, teachers, and parents.

Reason 1: Teachers think they're already doing formative assessment. My early understanding of summative versus formative assessment came from my curriculum director, who simply defined the two this way: "Summative assessments are tests and quizzes you grade, and formative assessments are anything you use to guide your instruction. You are already doing formative assessment all the time." These definitions were vague and incomplete, but not necessarily incorrect. The real problem was the message that we were "already doing formative assessment." Why then should we, a room full of teachers feeling burdened yet open to new ideas, seek to improve our practice of something we were apparently already doing? Just as with students, there is danger in false praise. Even worse, I knew my questioning techniques in whole-class activities were lacking. Had I been properly introduced to formative assessment, I may have improved my questioning practices and sought better ways to assess student understanding. I shouldn't have been led to think I was doing something well when in reality, I wasn't, or denied access to information that would have helped me improve.

Reason 2: Teachers are pressured to assess performance with grades. Grading practices can influence assessment practices, and pressure to assign grades for all classroom activity can inhibit the use of true formative assessment. In a world of 24/7 access to online gradebooks, parents and students expect to see near real-time measures of progress and achievement on their computer screens. To expect the full benefits of formative assessment, students must invest themselves in the improvement process as much as teachers. Once a teacher assigns a grade to a task, the message received by the student and parent is of summation – the task is complete, learning has been measured, and it's time to move on to the next task. The teacher might not want to send this message, but what's important is the message received by the student, not what was intended by the teacher. The sophisticated give-and-take of formative assessment is best recorded and measured outside the simple percentages and averages calculated by our technically limited gradebooks.

Formative assessment is understandable and practical, but inhibited by false assumptions. Administrators falsely assume teachers already know and use it, and teachers falsely assume students are willing and able to translate a summative grade into formative feedback. Fortunately, both of these obstacles can be overcome through better a understanding of formative assessment, improved communication, and a commitment to collaboration. Teachers, students, and parents alike should welcome an increased focus on improvement, instead of the summative and often harsh dependence on grades and percentages. Summative assessments and grades might be more familiar, but that doesn't make them easier or more beneficial.

Survey: Teaching to the Test

Colorado math teachers: I need your input! I have heard many teachers use the phrase "teaching to the test" and I want to know two things:
  1. What does "teaching to the test" mean to you?
  2. How well do you know the math CSAP test? What is assessed most? And least?
To help answer these questions I've developed a series of surveys, one for each grade-level math CSAP. To take the survey for your grade, please follow the appropriate link below:
I need your responses by Monday, April 26. If you teach multiple grade levels of math, you're welcome to complete the survey for multiple grades. (Just be consistent with your name and email, please!) If you aren't a Colorado math teacher but happen to know one, please pass this post along to them so they can complete the survey,

Many thanks for your cooperation, and I'll be sure to let everyone know the survey results and how they factor into my research. If any of you submit perfect rankings, I believe you deserve some recognition. I'd offer cash prizes, but that's not exactly in my budget. Sorry!

Calling Colorado Teachers: Can I Analyze Your Assessment?

Sometime in the next week or so I need to observe a math or science teacher. I'm turning to you, my PLN, for a volunteer. I'm open to any grade level, and here's the official description of the assignment:
Analysis of the Assessment Practices of a Math/Science Teacher
For this assignment you will observe a school mathematics or science teacher to analyze opportunities to assess student understanding embedded within a classroom context. Determine the extent to which assessment is embedded in instruction, detailing the kinds of questions and tasks used during instruction.
  1. Consider what it means to assess understanding
  2. As part of your observation, be prepared to distinguish between questions that guide the discussion and questions that elicit student understanding
  3. To what extent does the teacher create opportunities to gauge student thinking and learning? When does the teacher "check in" with students?
  4. To what extent do students responses influence instruction?
Additionally, as part of this assignment include a brief interview of the teacher’s beliefs about assessment. Some questions to consider…
  1. What is the purpose of assessment?
  2. When do you assess your students?
  3. [if appropriate] How do you use information from observations and listening to students to inform instruction?
These are just suggested questions. You are encouraged to adapt these and include other questions that you consider important. The report can be set up as a 2-part narrative that summarizes the observation and the teacher’s responses to the interview. Be sure to highlight specific aspects of what was observed and discussed, so that you can report your findings to the group and class on April 13th.
To be clear, I am not wishing to observe you giving a test or quiz, or to analyze the content and design of your test or quiz. I want to observe you during an instructional task and analyze what questions you ask, how you ask them, and how your students respond. As you can see from the assignment, I'm looking for the formative processes you use to determine if your students truly understand the content.

So if you don't mind an extra body in your classroom and can spare some minutes to discuss assessment, please email me at johnson@downclimb.com or DM @MathEdnet on Twitter. If you don't feel right for the study but can arrange something with a teacher that you think would be great to observe, that would be great, too. I'm in Boulder, am willing to travel a reasonable distance, and am available any day of the week. Your help will be greatly appreciated and I look forward to working with you!

Think Your Students Can Solve This Problem? (Hint: They Probably Can't)

In his book chapter "Aspects of the Art of Assessment Design," Jan de Lange gives some examples of tasks found on large-scale standardized assessments, such as the TIMSS, PISA, and NAEP. Here's a problem that was rejected from the PISA (it would have been given to 15-year-olds):

For health reasons people should limit their efforts, for instance during sports, in order not to exceed a certain heartbeat frequency. For years the relationship between a person's recommended maximum heart rate and the person's age was described by the following formula:

Recommended maximum heart rate = 220 - age

Recent research showed that this formula should be modified slightly. The new formula is as follows:

Recommended maximum heart rate = 208 - (0.7 x age)

Question 1: A newspaper article stated: "A result of using the new formula instead of the old one is that the recommended maximum number of heartbeats per minute for young people decreases slightly and for old people it increases slightly." From which age onwards does the recommended maximum heart rate increase as a result of the introduction of the new formula? Show your work.

Question 2: The formula Recommended maximum heart rate = 208 - (0.7 x age) also helps determine when physical training is most effective: this is assumed to be when the heartbeat is at 80% of the recommended maximum heart rate. Write down a formula for calculating the heart rate for most effective physical training, expressed in terms of age. (de Lange, 2007, pp. 103-104)

So why do you think this problem was rejected from the PISA? Is it not authentic? Is it not relevant to mathematical content? Is it culturally biased? Nope. It was rejected because when it was field tested, fewer than 10% of students got it correct. As a high school math teacher, I can't identify any mathematics most of my 15-year-olds haven't "learned" (I'm using that term loosely, given the circumstances), but that doesn't mean they can do the problem.

de Lange's point is that we must be careful about excluding problems from our assessments (particularly the large-scale ones) because they are too difficult. There are political pressures from many countries to exclude such items because they make countries look bad, but that's not a great reason to not use them. Having high-quality test items that we can use to track long-term trends is more important than saving face. Without problems like these, we risk not asking ourselves why our students have difficulties with such problems.

While de Lange makes a good argument, opinions may vary. How do you feel about item difficulty on assessments? Does it matter if it's a large-scale test like the PISA versus a classroom assessment? Would you give students a test item that had a known 10% success rate?

References
de Lange, J. (2007). Aspects of the Art of Assessment Design. In A. H. Schoenfeld (Ed.), Assessing Mathematical Proficiency, Mathematical Sciences Research Institute Publications (pp. 99-111). New York: Cambridge University Press.