Showing posts with label grading. Show all posts
Showing posts with label grading. Show all posts

Sorting Out the Summative: When Standards-Based Grading Meets the End of the Semester

Source: Wikipedia

Many teachers who choose to use standards-based grading eventually find themselves facing the reality of their school's grading policies and tradition: the expectation of final, summative grades that are reported as percentages and letters. So regardless how hard you try to focus on quality feedback instead of grades all semester long (for good reason), there comes a time when, for reasons probably beyond your control, you have to turn levels and descriptions of student understanding into numbers. This is SBG's "Monday Morning Problem" that doesn't always get addressed in theory. But this week is finals week for my basic statistics students, so for me the time has come to convert standards-based formative grades into a summative grade, including calculating final exam grades. Here I'll try to describe the two steps I'll take to calculate my students' grades: (a) conversion of their formative scores into a summative score and (b) scoring and inclusion of the final exam into their semester grades.

Formative to Summative
Besides giving students a lot of written and verbal feedback about where they should try to improve, I've been using the simplest of measures to record their performance on class objectives: either students (a) "get it," (b) "sort of get it," or (c) "dont' get it/haven't demonstrated it." You could think of these as "green light," "yellow light," and "red light," respectively. I've tried discerning more levels of understanding in a gradebook and it only seems to lead to confusion and indecision (both for me and students), so I'm sticking to three levels, as suggested in Her & Webb (2004). If I need more detail, I can always go back to the copies of the work students have submitted and the comments I've made.

The gradebook we have for class is pretty primitive and as far as I can tell it only accepts numbers, so I mark my three levels as either a 2, a 1, or a 0. It doesn't take much explaining to students that a 1 shouldn't be viewed as "out of two" and therefore worth 50%. I do tell them, though, that in order to receive credit for the course they should average a 1 across all objectives. In other words, you can't pass the class without an average of at least some understanding of every objective.

Around here and in many other places, 70% seems to be the low end of passing grades. (We're not messing with Ds.) So if a student with all 1s should get at least a 70%, and a student with all 2s maxes out at 100, and we choose a linear function between the two, the "conversion formula" to percentages is simply:

percentage = 30 * objective score average + 40

If you feel a little dirty at this point because you know you just reduced all the various skills, knowledge, and abilities of your students into a single number, I say join the club. If you didn't feel that way I wouldn't have expected you to be using standards-based grading to begin with.

A "No Surprises" Approach to Final Exam Grades
Designing a final exam is often tricky business. It can't possibly assess everything in the course, but we generally want it to include the major topics and themes for the class and be possible to complete in the time allowed. We also have to think about difficulty. Trust me, your students are!

Teachers want their finals to be challenging, but they don't want to have that sinking feeling as they grade the exams that maybe the test was too hard. For whatever reason, sometimes students perform poorly and averaging the final exam grade into their other grades will look like a disaster. But ask yourself: What am I more confident in, my careful judgments of students' ability as demonstrated over an entire semester, or a fleeting, one-time judgement of students' ability on a single assessment during the most stressful time of the year? If you're using standards-based grading, I already know how you'll answer that question. If not, consider this example: I have a student who I know can do stats. She's turned in good work. She's asked quality questions. We've had good discussions. But I also know she has seven final exams this week. I still think she'll do fine, but I'll understand if she's not at her best. And I need a grading system that reflects that understanding.

In order to free myself to still give challenging, yet reasonable, assessments, without risking any huge surprises when grades are calculated, I perform a little statistical magic that ensures that the distribution of final grades has the same center and spread of class grades before the final. I'm sure many of you try "curving" your exam scores some other way, such as letting the top score count as the total possible, or even having a pre-set distribution in mind of how many As, Bs, Cs, etc. you'll allow (which is not a good idea, generally, for reasons described by Krumboltz & Yeh, 1996). I prefer my method because it accounts for the distribution of grades, not just the top score, and the distribution is determined by the students, not arbitrarily by me. Allow me to demonstrate with a couple examples.

Suppose before the final the average percentage grade is 85 and the standard deviation of those grades is 10. Then I grade my final exams and find that the average final exam grade is 60 with a standard deviation of 18. Ouch. But don't worry -- statistics will come to our rescue.

Provided you know a little basic descriptive statistics, the conversion is simple. For each student's final exam score, find out how many standard deviations above or below the mean they scored on the final (their final exam z-score), and match that with the same number of standard deviations above or below the mean they'd fall on the pre-final grade distribution (their pre-final z-score). Consider the following students and the class and exam statistics above:

  • Suppose Student A scores a 51 on the final exam. That's 0.5 standard deviations below the mean. (51 - 60 = -9, and -9/18 = -0.5.) So where is 0.5 standard deviations below the mean on the pre-final distribution? If that mean is 85 and the SD is 10, then 0.5 standard deviations below the mean is 80. So I record an 80 for that student instead of a 51.
  • Suppose Student B scores a 75 on the final exam. That's about 0.83 standard deviations above the mean. (75 - 60 = 15, and 15/18 = 0.83.) So where is 0.83 standard deviations above the mean on the pre-final distribution? About 8.3% above an 85, so I record their exam grade as a 93.3.
  • Suppose Student C scores a 60 on the final exam. That's the same as the mean, so zero standard deviations above or below. That conversion is super-easy: their final exam grade is the mean of the pre-final mean, an 85.
For an example of how to set up a spreadsheet to do this, see https://docs.google.com/spreadsheet/ccc?key=0Anne5Z-jCkqhdDVtemkyaGhnRWFfclJoa0dIUVQ5RVE. I recommend making a copy of it for yourself and seeing what happens as you change values.

This is not a perfect system (and comments about its imperfections are welcome in the comments), but it does take away the element of surprise if the final exam happens to be way too easy or too difficult, or if other circumstances prevent grades from working out the way you'd expect. Yes, this is a norm-referenced system instead of a criterion-referenced system, meaning that the grades students earn on the final is measured largely as how they compare to their classmates and the class average. The good news is this: both the teacher and the students have an incentive before the final to master as many objectives as possible, and that is criterion-referenced. A high pre-final average helps everyone get a high final exam average, and a small pre-final standard deviation minimizes variability in final exam scores.

References

Her, T., & Webb, D. C. (2004). Retracing a path to assessing for understanding. In T. A. Romberg (Ed.), Standards-based mathematics assessment in middle school: Rethinking classroom practice (pp. 200-220). New York, NY: Teachers College Press.

Krumboltz, J. D., & Yeh, C. J. (1996). Competitive grading sabotages good teaching. Phi Delta Kappan, 78(4), 324-326. Retrieved from http://www.jstor.org/stable/20405782

RYSK: Butler's Effects on Intrinsic Motivation and Performance (1986) and Task-Involving and Ego-Involving Properties of Evaluation (1987)

This is the third in a series of posts describing "Research You Should Know" (RYSK).

As teachers, we care not only about what students learn, but why students learn. In a perfect world, we would all agree on what's important to learn and do and be self-motivated to learn and do those things. But our world isn't perfect, and students are motivated to learn and do things for many reasons. Understanding those reasons is important if we want students to be properly motivated and to perform well with the right attitude.

Ruth Butler earned her Ph.D. in developmental psychology from the Hebrew University of Jerusalem in 1982 and was a relatively new professor there when she teamed with veteran educational psychologist Mordecai Nisan, whose career includes time spent at the University of Chicago, Harvard University, The Max Planck Institute for Human Development, and Oxford University. Together, they sought to build upon studies that compared extrinsic vs. intrinsic motivation and positive vs. negative feedback, looking specifically at how different feedback conditions -- ones that can be manipulated by teachers -- affect students' intrinsic motivation.

For their 1986 paper, Effects of No Feedback, Task-Related Comments, and Grades on Intrinsic Motivation and Performance, Butler and Nisan expected that students who received feedback in the form of simple positive and negative comments (without elements of praise or grading/ranking) would remain motivated, while students who received grades or no feedback would generally become less motivated. To test this hypothesis, Butler and Nisan randomly assigned 261 sixth grade students to one of three groups. They gave the students two types of tasks: Task A was a quantitative "speed" task where students created words from the letters of a longer word, while Task B was a qualitative "power" task that encouraged problem solving and divergent thinking.

Butler and Nisan conducted three sessions with the groups:
  • Session 1: Students performed the tasks.
  • Session 2: Two days after Session 1 the tasks were returned.
    • Students in the first group got comments in the form of simple phrases such as, "Your answers were correct, but you did not write many answers," or "You wrote many answers, but not all were correct."
    • Students in the second group got numerical grades that were computed to reflect a normal distribution of scores from 30 to 100.
    • Students in the third group got their work returned with no feedback.
    After students reviewed their previous work, they were given new tasks and told to expect the same type of feedback when they returned for Session 3.
  • Session 3: Two hours after Session 2 students again reviewed their work and feedback (except for the third group, who got no feedback) from Session 2 and then got a third set of tasks. Students were asked to complete the tasks and were told that they would not get them back. The session ended with a survey of students attitudes towards the tasks.
When Butler and Nisan compared the students' average performance on the tasks in Session 1, all three groups scored approximately the same. That changed in Session 3. On Task A, students receiving comments and grades scored about the same in Session 3 (with an edge to the comments group for the creation of long words), but students receiving no feedback did far worse. For Task B, students receiving comments did significantly better than students who received grades or no feedback, who performed about the same. The only students doing well in Session 3 -- in fact, the only students consistently scoring higher, on average, in Session 3 than in Session 1 -- were the students who received comments.

The survey also showed attitudinal benefits for the comments group, who indicated they found the tasks more interesting and were most willing to do more tasks. Furthermore, 70.5% of students who received comments attributed their effort to their interest in the tasks, compared to only 34.4% of those graded and 43.4% of those receiving no feedback. Only 9% of students receiving comments said their effort was due to a desire to avoid poor achievement, compared to 26.7% of students receiving grades and 9.6% of the no feedback group. Lastly, 86.3% of students receiving comments wanted to keep receiving comments, while only 21% of the graded group wanted to keep receiving grades. The vast majority of graded students, 78.9%, wanted comments. The no feedback group was roughly split 50/50 on wanting comments or grades. None wanted to keep receiving no feedback.

Butler modified this study for her 1987 paper Task-Involving and Ego-Involving Properties of Evaluation: Effects of Different Feedback Conditions on Motivational Perceptions, Interest, and Performance. In it, Butler adapted a theory of task motivation used by Nicholls (1979, 1983, as cited in Butler, 1987):
  • Task involvement: Activities are inherently satisfying and individuals are concerned with developing mastery in relation to the task or prior performance.
  • Ego involvement: Attention is focused on ability compared to the performance of others.
  • Extrinsic motivation: Activities are undertaken as a means to some other end, and the focus is that goal, not mastery or ability.
Butler believed comments would promote task involvement, while grades would promote ego involvement. While both of these can be seen as intrinsic motivation, a third type of feedback needed to be considered: praise. Previous research on praise had gotten mixed results, possibly because researchers hadn't considered if the praise was task- or ego-involved. Butler's study would include ego-involving praise using comments designed to focus a student's attention on their self-worth and not on the task. Therefore, Butler hypothesized that praise and grades would generate similar results, results less desirable than task-involved comments.

The study was similar to the 1986 study, with 200 fifth and sixth graders split into four groups (comments, grades, praise, and no feedback) with subgroups in each for high- and low-achieving students. Tasks were administered in three sessions, with no feedback given after the third session. The tasks this time were divergent thinking tasks, used as Task B in the 1986 study. Praise would come in the form of a single phrase: "Very good." An attitude survey was given after Session 3.

As Butler expected, comments promoted task-involved attitudes while grades and praise promoted ego-involved attitudes. Students' interest in the tasks after Session 3 was higher for the comments group than for the grades, praise, and no feedback groups combined. Students who received praise showed more interest than those who received grades. As for performance, the comments group easily performed the best in Session 3, with both high and low groups improving their scores over Session 1, while all other groups performed about the same or worse compared to their Session 1 performance.

So what does this mean?

As a teacher who struggled with assessment and grading, it was Butler's work that most inspired me to start this RYSK blog series. Despite these results being 25 years old, there's not much evidence that Butler's findings have had a serious impact on the practice of most teachers. I suspect that few teachers know about Butler's work -- I certainly didn't. I was wrapped up in the scores and grades game, not fully aware of the impact those scores were having on my students. I knew it wasn't working, but I didn't have this kind of theoretical knowledge to support a significant change in my practice.

I'm not suggesting that we should suddenly demand a grade-free world. That's just not a realistic thing to expect given where we are now. What I would like to suggest is that teachers become more aware of how the feedback they give affects student motivation, and be careful to focus on task-involved comments whenever possible. Because students aren't likely to get this kind of feedback from standardized tests or computer-based learning systems (i.e., Khan Academy), it takes a teacher's touch to carefully craft the kind of feedback a student needs to sustain their motivation.

References

Butler, R., & Nisan, M. (1986). Effects of no feedback, task-related comments, and grades on intrinsic motivation and performance. Journal of Educational Psychology, 78(3), 210-216. doi:10.1037/0022-0663.78.3.210

Butler, R. (1987). Task-involving and ego-involving properties of evaluation: Effects of different feedback conditions on motivational perceptions, interest, and performance. Journal of Educational Psychology, 79(4), 474-482.

2004-2006: My Adventures in Standards-Based Grading (And Why I Stopped)

(cc): NASA
Strap yourself into the wayback machine, boys and girls, because we're going back in time five whole years. Life was different then: the U.S. was engaged in wars overseas, unprecedented disaster had struck our Gulf Coast, and teachers struggled to adapt to an assessment-centric school culture. It was, like, totally different than things are now.

In 2003 I began my teaching career at a medium-sized high school in Southern Colorado. My first year, as it is for many teachers, wasn't much more than a fight for survival. I spent much of my second year learning how to become something other than the teachers who taught me, and part of that meant tinkering with my assessments. By my third year, I was ready to collaborate with the school's other four math teachers in a concerted effort to improve our assessment and grading practices. I'm not sure I knew what an ideal assessment and grading system should look like, but I knew I wanted something other than the traditional quiz/test, either-you-got-it-or-you-don't system.

I think there's a reason math teachers in particular get so heavily invested in assessment and grading. Numbers are our friends. We trust them. We can sort them, scale them, manipulate them, and summarize them in ways that reveal certain truths. I was determined - perhaps obsessed - with finding a grading system that was accurate and fair. To me "accurate" meant "students earn the grade they deserve" and "fair" meant "objective and unbiased." I think I really believed that if I could just find the right scales and weights, the math would solve my assessment problems. 

While I might have been facetious in my opening paragraph, things really were different five years ago. Not many teachers were bloggers, Twitter hadn't been invented, and you couldn't search for #sbg or #sbar hashtags. Sure, standards-based grading existed, but Guskey, Marzano, Wiggins, et. al. sure weren't knocking on my door to tell me about it. I didn't know about SBG and I don't think any of my colleagues or administrators knew about it, either. It's sad that so many good ideas in education struggle to find their way from theory to practice, but that's another story for another day. This story is about my attempts at standards-based grading, my successes, failures, and frustrations I had along the way, and why I reverted back to a traditional grading system.

2004-2005: SBG(ish)
During my first semester of teaching I spent many hours after school designing quizzes and tests, thinking, "This is part of being a first-year teacher. Once I write these tests I'll never have to do this again." HA! Not only did I not reuse any of those assessments, in six years of teaching I hardly reused any of my assessments. Every semester I had new problems, assessment designs, and grading systems that made my old ones look horribly obsolete. As I went into my second year of teaching, I already knew that I wasn't going to be satisfied with a traditional system.

Before I go any further, I want to make this clear: I'm not claiming that I somehow independently invented or discovered standards-based grading. I was making this up as I went along and, as you'll soon read, it didn't necessarily work all that well. I did manage to reorganize my gradebook around concepts and skills instead of dates or arbitrarily-titled ("Chapter 8 Test") assessments. But while that part looked like SBG, sometimes very little else did. In fact, I wouldn't be surprised if you read this and decide I wasn't doing SBG at all.

As I said, I started making SBG-like changes to start my second year, as you can see in this passage from my fall 2004 Algebra 1 syllabus:
Each unit in the text will have several key objectives that should be your focus during that unit. Your score for an objective will usually be established by your performance on a quiz or test and is based on my perception of your understanding. Objectives are graded on a scale of 5 to 10. Think of 10/10 as an A+, 9/10 an A-, 8/10 a B-, and so forth. Objectives not genuinely attempted will be given a zero. If you are dissatisfied with one or more of your objective scores, it is highly recommended that you see me for 1-on-1 help, preferably before or after school. Because the CPM philosophy is mastery over time, you can expect to be quizzed or tested over each objective multiple times, with each time representing an opportunity to raise your objective score.

I got off to a good start by focusing on objectives, but then quickly got bogged down by the point system. SBG isn't about accumulating points. Sure, you'll need a way to record student performance, but a well-implemented SBG system will be formative and focused on feedback, not a point system. By the way, if you ever want to drive yourself into an insane asylum, try a 5-to-10 point grading scale based on perceptions of student understanding. I nearly drove myself crazy, scoring, re-scoring, and re-re-scoring every bit of work out of fear that somebody's 7/10 actually showed the same understanding as a classmate's 8/10. I was far more concerned with ranking and sorting than feedback. Even worse, I would have students who scored 7/10 or 8/10 over and over, earning passing scores for objectives even though I'd never actually seen them get any right answers.

By the following spring (with a new set of classes, like a college schedule) I decided that it would be far better for students to get most problems all right instead of all problems mostly right, so my syllabus now said this:
IT'S ALL ABOUT THE ROADMAP. The objective roadmap is a detailed list of objectives that together make up everything you should learn in the course. The objectives are organized by unit, but each objective will be assessed separately. To pass an objective you need to get 80% or better on the objective test. If you fail an objective, you will have to retake and pass that objective test before moving on to any objectives in the next unit.

(Here are my roadmaps for Algebra 1, Algebra 2, and Business Math.)

Almost every objective test consisted of five problems, and if it wasn't right, it was wrong. I was so sick of agonizing over partial credit I got rid of it entirely. Objectives were added to students' grades on a schedule, and students who fell behind schedule got zeros on objectives they hadn't yet attempted. Zeros have rarely been known to play nice with the statistical mean, so a student who averages 80% on 80% of the objectives (with 0% on the others) is only going to get a 64%. It was setting a high standard, sure, but somehow I convinced myself that I was a one-man-army who could rid the world of grade inflation. Anyhow, I asked students to focus less on the gradebook and more on objective progress charts like this one:
(Those are student numbers on the left. Yes, this class only had 11 students.)
Still, if this was SBG, it was really, really bad SBG. Three major problems stand out. First, I was telling students that there were a bunch of objectives and any one of them could stop them dead in their tracks. (I think I relented on that one not far into the semester.) Second, I was still way too focused on points and percentages, not formative assessment. Third, a student who was "mostly right" on all five test problems got a zero in the gradebook, even if the cause was a small procedural mistake that he/she happened to repeat on all five problems. How's that for demotivation?

Amidst all the problems of this system, there emerged a brilliant, shining light, and her name was Tara. She gave me hope. She was part of a pretty special Algebra 2 class, and I think previous struggles in math had left her without much confidence. She barely spoke, and when she did it was very quiet, and quite often it was apologetic, like "I'm sorry to make you stay after school so I can retake a test." Tara understood the grading system and used it to her advantage, retaking test after test with study sessions in between. I remember a marathon session in the days before the final exam, when Tara was within sight of her goal of passing every objective. I think she stayed after school for five hours until her goal was finally met. Tara, if somehow you happen to read this, I hope you realize you weren't a burden, you were an inspiration. You did and still make me want to be a better teacher.

2005-2006: Professional Learning Communities (PLCs) and Common Assessments
I think every teacher knows that feeling when an administrator latches on to an idea from a conference or meeting they attended and then bring it home for their staff to use. Unfortunately, in a rural district like ours with scant resources, the promotion usually went like this:
This year we're going to implement [insert acronym or edujargon here]. I've talked to a lot of other principals who are doing it and think it's very effective. ["Effective" usually means "Since we didn't make AYP, again, I need to tell the state that we're trying something new. Again."] The principals I talked to sent their teachers to summer workshops and hired a coordinator to guide the implementation of [edujargon], but since we can't afford so much as a box a tissues for this school, I'll just tell you what I remember and we'll make the rest up as we go along.

For 2005-2006, administration was pushing Professional Learning Communities (PLCs) and common assessments. The math department PLC opted to tackle Algebra 1 in the first year, and we began by brainstorming a list of concepts we thought all Algebra 1 students should know. We shuffled our list into eight groups that aligned with our textbook and other materials, and finally scheduled eight test dates spaced evenly throughout the year.

Because of some of the success I experienced the previous year, we opted to grade each concept separately instead of amassing a single score for the whole test. Students who did poorly on a concept would be remediated and given an opportunity to retest until they showed proficiency with each concept. Perhaps more importantly, since all the Algebra 1 teachers were assessing the same objectives, using the same tests, all scored the same way, we now had a basis for comparing our students' performance, and that led to a sharing of tips and techniques for teaching particular skills. Unfortunately our time was limited so I never learned as much from my colleagues as I hoped, but there is certainly an advantage in using the same SGB system across teachers in a department, or even now across schools in our ever-connected world.

We had a schedule change that allowed freshmen to take Algebra 1 all year long. With about a month left in the school year, the progress chart for one of my classes looked like this:
You can detect a problem here: too many students started slacking at the end of the year!
Even though the system was working, it was showing new cracks. Students knew I had a finite (usually just 2 or 3) set of tests for each objective, and if they failed form A, they would carefully memorize their mistakes and the right answers, then pretty much fail form B on purpose so they could then take form A again. (Ugh!) Also, after two years of using this kind of system, I was growing tired of explaining and defending it to students, parents, and administrators. I enjoyed the extra time spent with students who came in after school, but it became increasingly hard to make up for the lost time. Lastly, the paperwork was a challenge - I had to have every form of every test ready at a moment's notice for students who needed to take or retake tests. Because I didn't have my own classroom, this meant hauling around a file box to three or more rooms each day in two different buildings.

One of my biggest dissatisfactions with SBG was the quality of the assessments I was using. During instruction my students worked in groups, delving into big problems in context, problems that required a combination of reasoning and the integration of multiple skills. Because I wanted to assess every skill in isolation, the tests I wrote were nothing like that. Too often they looked like this:
(What, I couldn't think of contexts for these?)

Goodbye, SBG
I was having another major problem at the end of the 2005-2006 school year: because of budget cuts, declining enrollment, and a shuffle of veteran teachers, I found myself out of a job. Even though I knew these to be the reasons, I was pretty hard on myself and thought if I had been a better teacher, then the district would have somehow found a way to keep me.

In the interview for my new job I talked about my assessment practices and quickly got the feeling that they wouldn't work in my new school. The school day was two hours longer and few students would be able to spend time after school. I still hadn't figured out how to measure skills in isolation while wanting students to tackle larger, more comprehensive tasks. I was now the only teacher in the district teaching any of my subjects, so I had no colleagues to work with on common assessments. Frustrated and not wanting to rock the boat, I figured a simple, traditional assessment and grading system would be the easiest and best way forward. I now know that I couldn't have been more wrong! In my bitterness and self-doubt I gave up the progress I had gained over two years of hard work.

If you take anything away from this post, I hope it's this: it's not 2005 anymore. You don't have to stumble around and make the mistakes I did. You don't have to work in isolation like I did. There's a great community of bloggers discussing standards-based and other grading practices, and I haven't met one yet that doesn't want to help their fellow teachers. You might not have access to the research, but some of us do and we enjoy sharing what we think about it. So please, whether you're new to SBG or have been doing it for years, share your experiences and don't hesitate to ask questions!

Why Don't More Teachers Practice Proper Formative Assessment?

Research that fails to impact practice is a problem in any arena, and education is no exception. Teachers have many reasons for not implementing practices based on research. Too often research is unknown to teachers, locked away in journals that teachers and schools cannot afford. Beyond the cost of access, much of the best research is generally written for higher academic audiences and not easily digested by a busy and distracted practicing teacher. Transforming research into improved practice takes time, effort and patience, and is made easier in a professional community who share ideas and experiences. Teachers are constantly improving their practice, but too often the improvements are driven by personal failures or anecdotal evidence, not the quality results of dedicated educational researchers. This is a crippling inefficiency in the field of education, one that is largely self-imposed and tied to traditional practices held in place by the inertia of our experience.

Research strongly suggests that teachers could improve student learning by using formative assessment. Chapter tests and final exams are summative: they summarize the knowledge and skills a student has acquired, and are generally assigned a fixed grade. In fact, any assignment or task assigned a fixed and lasting grade can be considered summative, at least in part. Formative assessment, in contrast, focuses on improvement rather than final measurement. Teachers use formative assessment to adapt their teaching, and students, equal partners in the process, use feedback and self-monitoring to improve their knowledge and skills.

For any classroom teacher, the concept of formative assessment should be comprehensible and implementation should not be impeded by any significant obstacles. So why don't more teachers practice proper formative assessment? I suggest two simple reasons, reasons that could be eliminated by improved understanding between teachers, administrators, teachers, and parents.

Reason 1: Teachers think they're already doing formative assessment. My early understanding of summative versus formative assessment came from my curriculum director, who simply defined the two this way: "Summative assessments are tests and quizzes you grade, and formative assessments are anything you use to guide your instruction. You are already doing formative assessment all the time." These definitions were vague and incomplete, but not necessarily incorrect. The real problem was the message that we were "already doing formative assessment." Why then should we, a room full of teachers feeling burdened yet open to new ideas, seek to improve our practice of something we were apparently already doing? Just as with students, there is danger in false praise. Even worse, I knew my questioning techniques in whole-class activities were lacking. Had I been properly introduced to formative assessment, I may have improved my questioning practices and sought better ways to assess student understanding. I shouldn't have been led to think I was doing something well when in reality, I wasn't, or denied access to information that would have helped me improve.

Reason 2: Teachers are pressured to assess performance with grades. Grading practices can influence assessment practices, and pressure to assign grades for all classroom activity can inhibit the use of true formative assessment. In a world of 24/7 access to online gradebooks, parents and students expect to see near real-time measures of progress and achievement on their computer screens. To expect the full benefits of formative assessment, students must invest themselves in the improvement process as much as teachers. Once a teacher assigns a grade to a task, the message received by the student and parent is of summation – the task is complete, learning has been measured, and it's time to move on to the next task. The teacher might not want to send this message, but what's important is the message received by the student, not what was intended by the teacher. The sophisticated give-and-take of formative assessment is best recorded and measured outside the simple percentages and averages calculated by our technically limited gradebooks.

Formative assessment is understandable and practical, but inhibited by false assumptions. Administrators falsely assume teachers already know and use it, and teachers falsely assume students are willing and able to translate a summative grade into formative feedback. Fortunately, both of these obstacles can be overcome through better a understanding of formative assessment, improved communication, and a commitment to collaboration. Teachers, students, and parents alike should welcome an increased focus on improvement, instead of the summative and often harsh dependence on grades and percentages. Summative assessments and grades might be more familiar, but that doesn't make them easier or more beneficial.

Boulder Valley Students "Earn" Zeros on Mis-administered CSAP

According to a story in Boulder's Daily Camera this afternoon:

"Sixty-seven students at Superior's Eldorado K-8 School will receive zeros on the writing portion of this year's CSAP test after an eighth-grade teacher broke the rules by having students practice on writing topics copied from a previous year's test."

Two years ago, the exact same situation happened at Runyon Elementary in Littleton. In both cases, the schools realized what happened soon after the test was administered and contacted CDE to report their concerns. In both cases, it appears CDE will give a score of zero to every affected student. This situation upset lawmakers enough in 2008 to include a provision for it in their CAP4K, more officially known as Senate Bill 08-212. It amended Colorado Revised Statute 22-7-604(3) to read:

"...the department shall identify and implement alterations in the calculation method, or other appropriate measures, to ensure that, to the fullest extent practicable, a public school is not penalized in the calculation of the school's CSAP-area standardized, weighted total score by inadvertent errors committed in the administration of an assessment."

So why doesn't that law apply to this new case in Boulder? Simple: it was repealed last year by Senate Bill 09-163, for reasons I never heard discussed. Assuming this is as honest of a mistake as BVSD claims, it doesn't seem fair to give kids zeros. After all, I'm sure the kids know more than nothing. I hope we hear an explanation from CDE and the Colorado legislature about the repeal of the CAP4K provision, and I won't be surprised to find it back in the law soon.

Do Online Gradebooks Compromise Our Teaching?

Aaron Eyler recently raised the question of online gradebooks on his blog. While Aaron's concerns centered more on "what does a grade mean" and the easier-than-ever assumptions we can make by looking at a letter in a gradebook, I've been more concerned about the negative effects online gradebooks might be having on how we teach. I'm all for running an open classroom and I like knowing that parents and students are monitoring progress, but I believe online gradebooks have some negative consequences. For example:

1. Do online gradebooks discourage formative assessments? From my standpoint, once you assign a fixed grade to an assignment, the gradebook sees it as summative. (Even if the teacher doesn't.) Formative assessments are important tools in both assessment and instruction, and often can go on for days without deserving a grade in the gradebook. Unfortunately, we get external pressure to put everything in the gradebook, whether it be from parents or administrators who want to monitor progress or athletic directors needing grades to determine eligibility. Students can also lack motivation if they aren't seeing their grades change as they work.

2. Should online gradebooks show class assignment averages? Suppose you give an assignment to ten students and the scores are 90, 90, 90, 80, 80, 70, 70, 70, 0, and 0. (The use of zeros is another issue, but they were expected in my school if students didn't turn in assignments.) All the students who turned in the assignment passed, and half the class got a B or better (80% = B). But because of the zeros the class average is 64%, which was failing at my school. Sadly, more than once a parent or administrator would contact me and say I had failed to teach the students because "the class got an F on the assignment."

3. Online gradebooks (at least the ones I've used) only accept numbers as input, significantly restricting options teachers have for giving feedback. Butler (1987) performed a study that revealed that indicating the grade earned on an assignment had negative impacts on performance. If you have a choice between grades, feedback, and grades plus feedback, going with feedback only can lead to the most improvement because students will focus on the feedback and use it to improve. With online gradebooks, the grade received on any assignment is a click away, potentially rendering the feedback less useful.

Poor grading practices can have negative effects on both assessment and instruction. I'd be surprised to find a school not using an online gradebook, whether it be popular systems like Infinite Campus and Powerschool, or systems from smaller players like Go.edustar, SME, MyGradeBook, Thinkwave, and Gradeconnect. (A Google search reveals many more!) Each gradebook has its own limitations, but my three concerns above likely will exist in all of them. What are your experiences with online gradebooks? Am I underestimating the positives? Are there negatives that I haven't listed? I'd love to hear your thoughts.

References

Butler, R. (1987). Task-involving and ego-involving properties of evaluation. Journal of Educational Psychology, 79(1987), 474-482.

Looking ahead to PhD: Focus and Vision

As the first of my classmates were congratulating me on my PhD acceptance, the inevitable question came: "So what do you want to study?" As I started to answer, I stumbled. I didn't have the 12-second, sure-of-myself answer I was supposed to have, despite having thought hard on the question when I wrote my application essay.

I learned of my acceptance five days ago and ever since I've been thinking about how I'm going to make the most of this opportunity. CU-Boulder's School of Education is home to some of the finest researchers in the field, and my experience in the master's program has been excellent. I've always been intellectually curious, and in the PhD program I'll learn how to apply methodologies to that curiosity in ways that are both personally satisfying and helpful to others. To clearly answer the "What do you want to study?" question, I need to develop "focus." The best demonstration of "focus" that I've seen lately is Gary Vaynerchuk's "Linchpin" video he made for Seth Godin's blog:

Linchpin: GaryVee from Seth Godin on Vimeo.


Even if you think P.Diddy's hair was an odd example, I think you get the point. As I prepared my application essay, I tried to focus on one specific thing in both curriculum and instruction (my major):

Curriculum Focus: I'm interested in the intersection of policy and practice, specifically when and how state and national standards become the curriculum that is experienced by students. I studied the math wars for my undergraduate thesis ten years ago, and my new area of interest is how our standards' increased demand for statistics (including all forms of data analysis and visualization, uncertainty, and probability) is causing change in textbooks, course sequences, standardized tests, and traditional perspectives on school mathematics.
Instruction Focus:Teachers' views of standardized assessment are affecting classroom assessment practices, and in turn poor classroom practice in assessment and grading has an inordinate influence over how teachers teach and how students learn. Additionally, negativity towards assessment deters teachers from climbing the mountains of potentially helpful data produced by assessments. I'm interested in understanding both the theoretical and practical problems teachers have with assessments and grading, and how to develop pragmatic solutions that are favorable to practicing teachers.

It's good to have focus. Unfortunately, I've never been comfortable with the pigeon-holing that happens when someone declares a specialty. I remember talking to other undergraduates after we had declared our majors, frustrated with the feeling that because we had chosen a field, people assumed we were now ignorant of everything else. I might be overreacting, and I hope to "crush it" (as GaryVee says), but not at the expense of stifling my curiosity, or losing sight of something bigger, which I'm calling "vision." Vison is big. I was afraid my vision would be too broad or vague in a specialized world, but my attitude was helped greatly by Aaron Eyler's blog post about connecting research with K-12 teachers. Eyler is personally frustrated with how little research makes its way into the hands of K-12 teachers, and the little that does is presented in a "cookie-cutter" fashion that has little of the impact intended by the researcher. Why don't teachers and administrators do a better job keeping up with research, and why don't researchers spend more time working with K-12 educators?

So what's my vision? In short, I want to help teachers be better teachers. It's that simple. I hope to interact with teachers (or future teachers) whenever possible, whether it be teaching methods classes, visiting schools, giving presentations, or whatever else that might engage me with a community of teachers. When I do research, I want to always think of teachers as my audience, not some journal editor or professor who might be refereeing my work. I want to write things that teachers will want to read, and explore ways of delivering research to teachers in helpful ways. That's my vision, and I should be confidently unapologetic about believing you can't focus without vision.

Competition

A new semester has begun and with it comes new readings. For my Assessment in Math and Science class we read two seminal articles in assessment research, including Black and Wiliam's Inside the Black Box (1998). In discussing room for improvement in assessment, the authors state:

"Approaches are used in which pupils are compared with one another, the prime purpose of which seems to them to be competition rather than personal improvement; in consequence, assessment feedback teaches low-achieving pupils that they lack 'ability,' causing them to come to believe they are not able to learn." (p. 142)

For me, I think competition has always been a healthy part of education, but I understand the authors' concern. Some grading practices, such as grading on a curve (using statistical normal curves that assure some students receive high grades and others receive low grades) are by nature competitive. Unfortunately, that kind of competition not only dissuades students from working cooperatively, but it can sabotage good teaching because of the expectation that some students will fail (Krumboltz & Yeh, 1996), and thus must be used with extreme caution.

Black and Wiliam did not specifically mention grading on a curve, which made me wonder how competitive grading was in general, and what the real source of that competition might be. The questions I posed to the class (a weekly requirement on our course message board) were:

Do you think students are subjected to competition by the teacher and their system of assessment, or is competition a natural reaction of students? Can a teacher design and enforce a competition-free system? If so, how?

After numerous responses and further thought, I think competition is so ingrained in who we are as human beings (both in an innate and culturally-driven way) that attempting to eliminate it in education is not only impossible, but it would be misguided and potentially harmful to try. Blaming the competition for those students who don't compete well isn't a particularly helpful tautology. Surely the issues for struggling students go deeper than that, and fixing the problem by eliminating competition isn't an efficient way of handling the problem.

I suggest teachers and schools try to find ways to utilize the benefits of healthy competition in a fair and voluntary way. Students who wish to compete would always have an outlet, and those who don't hopefully wouldn't feel that anything is being forced upon them. Just as competitive sports or other competitive school activities are voluntary and shouldn't be imposed on all students, the competitive aspects of education shouldn't either, unless students choose to participate.

References

Black, P., & Wiliam, D. (1998). Inside the black box: Raising standards through classroom assessment. Phi Delta Kappan, 80(2), 139-148.

Krumboltz, J. D., & Yeh, C. J. (1996). Competitive grading sabotages good teaching. Phi Delta Kappan, 78(4), 324-326.

Cases in the Ethics of Grading: Jodi Warren and Mr. Kennedy

The following is my fourth and final (as far as I know) in a series of hypothetical cases meant to raise questions about grading practices. I'd like to recognize Kenneth Strike and Jonas Soltis for their book "The Ethics of Teaching," which inspired the style and structure of this case. Enjoy and discuss!

Louis Kennedy and Jodi Warren are two teachers within the Metro City School District. Louis is a popular fifth grade teacher at Worthington Elementary and Jodi is a new, 23-year-old sixth grade English teacher at Carter Middle School. Most of the students at Worthington Elementary matriculate to Carter Middle, and many of Jodi's students had Mr. Kennedy as their fifth grade teacher.

At her very first day at the school, Jodi Warren listened to her principal preach some of the reforms they were instituting that year at Carter Middle. The biggest change was the enforcement of course requirements and credits that must be earned before being moved on to high school. "No more social promotion!" the principal cried. "We will teach these kids to be responsible and prepare them for high school!" Jodi wanted to be tough her first year of teaching, so this was a philosophy she was willing to get behind, even at the sixth grade level. She'd heard horror stories of teachers who were pushovers and let the kids run the classroom, and she was determined to not be one of those teachers.

Several months passed and Jodi stuck to her plans of holding kids responsible and grading rigorously but fairly. Many students were earning grades of C or lower, but Jodi did not seem too concerned, as she had always viewed a C as an average grade. At the end of October, she faced her first parent teacher conferences. Parents were not happy. Conference after conference, Jodi listened to parents make comments like, "This is the worst grade my student has ever gotten," and "My son was doing much better in Mr. Kennedy's class last year." Jodi tried to explain the new expectations at Carter Middle, and how it was part of a plan to better prepare students for the rigor of high school. "Well, my daughter has As and Bs in all of her other classes," some would reply. Jodi also learned that a number of the parents had already complained to the principal about her ineffectiveness as a teacher, and it did not appear that her principal gave her much support. Jodi left school that night with serious doubts about her future as a teacher, wondering if her teaching really had been so poor despite her best efforts.

The next day Jodi attended a district workshop where she had an opportunity to work with other teachers, including Louis Kennedy. Jodi was curious to hear if Louis had experienced problems similar to those she was having with students Louis had taught the year before. Maybe she could gain some insight as to what she was doing wrong, or pick up advice from a popular fellow teacher. When Jodi asked about trying to be a tough grader, Louis gave her a surprising reply. "We give A-F grades in fifth grade, but I never take it too seriously. They don't count for anything. They don't end up on a transcript for college or factor into a GPA. A few years ago I had given a couple students Ds and Fs, and I found out that it only created animosity between myself and the students. They weren't happy in class anymore and parents assumed I was doing a bad job. Now I never give a grade below a C, everyone seems to be happier, and I have an easier job connecting with struggling students."

Jodi is now unsure what she should do. If she suddenly raises her students' grades, she'll become the easy pushover she did not want to be. If she keeps grading the way she is, parents will think she is not a good teacher and that could put her future as a teacher at risk. If she complains about Mr. Kennedy's easy grading practices she'll be seen as whiny and a tattle-tale, a trait she despises in her students.

Questions
  1. Does it make a difference that Jodi is a new teacher, and has several years of teaching before she receives tenure?
  2. Given Jodi's choices (raise grades and be a pushover, be firm on grades and risk her career, or complain about Mr. Kennedy) which would you choose and why?

Cases in the Ethics of Grading: Mrs. Lemon and Troy Mann

The following is my third in a series of hypothetical cases meant to raise questions about grading practices. I'd like to recognize Kenneth Strike and Jonas Soltis for their book "The Ethics of Teaching," which inspired the style and structure of this case. Enjoy and discuss!

Troy Mann is a struggling student in Mrs. Lemon's Algebra 2 class. By all accounts from Troy's previous two math teachers, he barely passed Algebra 1 and Geometry, and his Geometry grade might have been artificially inflated to compensate for his weak skills coming out of Algebra 1 and suspicions of a possible undiagnosed learning difficulty. Still, in both classes, Troy's modus operandi was the same: "coast" through as much of the semester as possible, do as little work as possible, and then apply just enough sincere effort to avoid being ineligible for sports. Troy understood if he failed Mrs. Lemon's Algebra 2 class he would be ineligible for the beginning of the second semester, meaning he would miss most of basketball season. Troy expected to start on the varsity team, and Mrs. Lemon happened to be one of the biggest basketball boosters at the school.

Using athletic participation as a motivator, Mrs. Lemon and Troy began to meet almost every day after school to improve his math grade. Unfortunately for Troy, his lack of effort earlier in the semester and in previous classes left him unprepared for the rigors of Algebra 2, and most of his time with Mrs. Lemon was spent relearning skills from Algebra 1. After six weeks Troy was showing great improvement on his Algebra 1 skills, but was running out of time in the semester to master Algebra 2 skills well enough for the final exam. Unable to fully catch up, Troy badly failed the final exam and received an F for the course.

When school resumed after winter vacation, Mrs. Lemon was called to the principal's office to meet with Troy's parents. This was the first time Mrs. Lemon had met Troy's father, but had talked previously on several occasions with Troy's mother, who worked in the school district administration office. The principal explained to everyone that a failing grade would mean Troy would be ineligible for basketball. Both of Troy's parents emphasized how important sports are to Troy, and without sports he may not have any motivation to be successful in school. Troy's parents also expressed their difficulty in understanding how Troy could spend all those hours after school, learning and showing steady improvement, and still fail the course.

Mrs. Lemon explained that to pass the course students must show mastery of Algebra 2 objectives, and while Troy should be commended for remediation of his weak math skills, it was not worthy of Algebra 2 credit. When the principal asked if there was any way of making Troy eligible for the second semester (implying Mrs. Lemon should change the grade to passing), Mrs. Lemon bluntly suggested that he ask the athletic director to ease the eligibility standards instead of compromising her own academic credibility. With that, the meeting ended and Troy's parents left.

Before letting Mrs. Lemon go back to class, the principal stopped her and said, "I want to tell you a little story. When I was in college, I really struggled in some of my classes and for one class I failed the final and was going to fail the class. The professor for that class, instead of sticking me with that grade, brought me over to his house, made me dinner, and went over the test item-by-item. I didn't retake the test – he was just helping me understand the material. He changed my final exam grade and I passed the class. Think about sitting down with Troy and doing him the same favor."

Questions:
  1. State content standards are meant to guide curricula, but should mastery of that content be the sole determinant of a student's grades? Should Troy's grade reflect his effort and skill improvement?
  2. Was it ethical for the principal to ask Mrs. Lemon to consider changing Troy's grade in the presence of Troy's parents? What's more important, the perception of a principal's fairness and neutrality or his/her willingness to be open and helpful to students and parents?
  3. What do you think was the principal's intent of telling the personal story after the meeting?
  4. If you were Mrs. Lemon, would you change Troy's grade?

Cases in the Ethics of Grading: Mr. Green and Tracked Classes

The following is my second in a series of hypothetical cases meant to raise questions about grading practices. I'd like to recognize Kenneth Strike and Jonas Soltis for their book "The Ethics of Teaching," which inspired the style and structure of this case. Enjoy and discuss!

Mr. Green is the sole history teacher at a small high school. Last year the school experimented with having an honors science class and feedback from teachers, parents, and students was generally positive. This year the school has decided to expand its selection of honors courses and Mr. Green will be responsible for teaching the school's first section of Honors World History. The class is designed for 10th and 11th graders, almost all of whom Mr. Green has already taught in lower-level courses.

Because of the school's size, there are only enough students to warrant two sections of World History: one honors and one regular. Mr. Green is responsible for selecting which students are to be placed in the honors section, and he overwhelmingly and sensibly chooses students who have earned grades of A or B in his previous courses. That naturally leaves almost entirely C or below students for the non-honors section of World History. Mr. Green is fully aware that the creation of Honors World History amounts to tracking, a controversial subject thought by some to be outdated and unethical. To ease his own concerns, Mr. Green ensures that both classes receive the same curriculum, the same textbooks and materials, and the same styles of instruction. To make the honors section worthy of its title, he skips the chapter review day before each test, thereby slightly speeding up the class, and makes the grading scale marginally more challenging.

The first semester of the two sections of World History seems to go smoothly. Mr. Green's plan to teach both classes using the same materials and methods appears to have avoided controversy. Having not heard any complaints from students, parents, or administrators, he proceeds into the second semester using the same class policies and procedures.

Three weeks into the second semester, the Mr. Green's principal, Mary Williams, checks the grades for both honors and regular World History. She is upset to find that while there are plenty of students earning As and Bs in the honors class, not a single student in the regular World History class has a grade above a C. Ms. Williams calls Mr. Green into her office, accuses him of grading unfairly, and demands to know why the regular World History class doesn't have "its own As and Bs." Mr. Green offers several reasons, including: A) Tests usually help students raise their grades, but it's early in the semester and they haven't taken a test yet, B) The regular World History class doesn't have students with a previous record of earning As and Bs, so the lower grades should be expected; and C) It would be unfair to arbitrarily give As and Bs to students in regular World History if they weren't mastering the content at levels similar to the honors students, and doing so would take away the incentive for students to take Honors World History.

Ms. Williams is not swayed and flatly tells him that he "needs to give higher grades." When Mr. Green asks, "Does that mean I should give higher grades without regard to ability or achievement?" the she only responds by repeating herself: "You need to give higher grades."

Questions:
  1. Is it reasonable for Ms. Williams to expect each class to have a distribution of grades A through F?
  2. Suppose you raised some of the grades in the regular World History class, believing the class does deserve a broader distribution of grades. Would it be hypocritical to do this unless you also lowered some of the grades in Honors World History to include more Ds and Fs?
  3. Did Mr. Green trap himself in this dilemma by trying to make the two sections so similar? In other words, do you think he would have been more comfortable assigning each class "its own As and Bs" if the two classes were drastically different by design?
  4. Schools use grades and GPA to measure and sort students. For honors vs. regular scenarios such as this, whose responsibility is it to ensure the sorting is done fairly? Is it Mr. Green's duty to sort the students across both sections, as he was attempting to do, or should the district handle that burden through policy? (Example: Districts often choose to weight honors courses on a 5-point scale to reflect their increased difficulty over regular courses.)
  5. If you were Mr. Green, would you raise the grades? If so, why? If not, why not?

Cases in the Ethics of Grading: Dr. Jones and Tara Hightower

The following is a hypothetical case meant to raise questions about grading practices. I'd like to recognize Kenneth Strike and Jonas Soltis for their book "The Ethics of Teaching," which inspired the style and structure of this case. Enjoy and discuss!

Dr. Susan Jones, an assistant professor in her first year at Central State University, is teaching a molecular biology course to a class of about twenty undergraduates, most in their second or third year of college. The content of the class comes easy to Dr. Jones, but showing up on-time at 8:00 three days a week is not. Most of the class is struggling with the material and disliking the class, so to make the class more enjoyable she occasionally interrupts lecture with some "fun" activities, such as watching cartoons or playing games. Dr. Jones's primary suggestion to those struggling in the class is for them to see her during her office hours for individualized help. She is very generous with her time and has helped many students make improved progress with the coursework.

Tara Hightower is an honors student at Central State and is one of Dr. Jones's struggling students. Tara is proud and independent, and somewhat stubbornly chooses to not see Dr. Jones for individualized help. Tara attends every class, is always on time, and studies both the text and her notes for hours each week in order to keep up with the material. This has always been a successful strategy for Tara in the past, and she resents the idea that she should have to see Dr. Jones individually, especially since some class time is already being wasted on fun, games, and Dr. Jones's tardiness. Tara's scholarship and status in the honors program is dependent on her maintaining a 3.5 GPA, and she received a notice at midterm that she had a D in the class. If Tara can not substantially improve this grade, she'll be given a warning by the honors program and risks losing her scholarship.

At the end of the semester, Dr. Jones asks Tara to come see her about her final exam and grade for the course. At the meeting, Dr. Jones explains to Tara that while she passed both the homework and the final exam, she did not perform up to expectations and should take the course again. To ensure Tara retakes the course, Dr. Jones assigns her a failing grade. Tara feels that receiving an F for passing work isn't fair, but agrees that her performance was sub-par and knows she needs to retake the course, regardless if her grade was a D or an F. After meeting with her advisor, Tara changed her schedule so she could retake the course with a different professor the next semester. Tara received an A on the retaken course and regained a positive standing in the honors program.

Questions:
  1. Can the extra hours Dr. Jones spends working with students individually make up for misused class time?
  2. Should Tara's refusal of any out-of-class individual help influence the grade she receives from Dr. Jones?
  3. Should the fact that Tara must retake the class, regardless of earning a D or F, matter to Dr. Jones?
  4. Teachers typically retain autonomy over their own grading practices. If Tara were to choose to protest the failing grade, who should have the power to change it? What are the ethical implications for Dr. Jones if her grades can be overturned?

Grading and Teacher Autonomy

I've known some teachers that probably wish they had the freedoms teachers enjoyed in the one-room schoolhouse days, but in the modern context of standards-based education that just isn't a reality. Standards, for better or worse, shape our curriculum, choice of texts, long and short-term unit/lesson planning, instruction, and seemingly every level of assessment. Could there possibly be a significant daily teacher practice not guided by standards?

Sure there is: grading.

There are no national or state standards to tell you what a "B" means, no standards to tell you if you should use a weighted grading system, and no standards to tell you if habitual tardiness to class should result in a grading penalty. I know of no set of classroom grading standards published by any major educational organization at any level. Teachers are left to figure out grading for themselves. As a math teacher I put extra pressure on myself to develop the best possible grading system, and over six years I tried all sorts of variations, none of them perfect.

My general philosophy was to grade on the mastery of the content described by the standards, and leave out as much extraneous stuff as possible. (Like deducting points for tardiness, for example.) Even on that task I know I failed, because the vast majority of student homework was graded for completion, not correctness. Still, I felt I had a principle for grading and I tried to stick to it. I actively resisted meaningless grade inflation and expected grades to usefully measure what students knew and could do. I felt it was reasonable in Colorado, where more than 60% of high school students score partially proficient or below on the CSAP, for Cs and lower scores to be an acceptable reality. Also, I hoped my students' grades would have a strong positive correlation to their CSAP scores. (Sadly, I never tested this.) Grades that failed to correlate would have indicated problems with my grading system and been a disservice to my students. I was miles away from having this vision be true for all students all the time, but it was a goal nonetheless.

I hope you can get a feeling for the amount of thought I invested in my grading, and realize that teachers all do this to some degree as a result of the autonomy we have over grading. The autonomy isn't total, however. If I had given 95% of my students As and Bs, I probably wouldn't have had much reason to continually re-examine my grading practices. In most schools, a teacher who gives that many high scores will certainly avoid negative attention, or perhaps even be praised for being such a good teacher. Low grades, on the other hand, attract plenty of attention from students, parents, and administrators. You can get many interesting suggestions from all involved, including grade curving (which means raising in this context, trust me), extra credit, and throwing out low test scores. Most of this is based on appeasing students and parents, not the development of intelligence, but it happens in schools large and small.

So, with no grading standards, to whom does a teacher surrender their grading autonomy and to what degree? The ethical quandaries can build extremely quickly, and several such cases will be topics here on this site in the near future. One of those cases will present a situation where some might argue the teacher got total autonomy and used it to be purposely unfair to a student. Another will present a teacher being pressured to change a grade for non-academic reasons. Hopefully these cases will challenge you to think about your own classroom, and good reasons to either defend or change your own grading practices.

Lastly, I want to pose a question about a possible influence on our grading systems. This is not a who, but a what: How does your gradebook itself influence your grading? Most schools use a system like Powerschool or Infinite Campus. (I've used Goedustar and Integrade Pro.) How do the capabilities of the program, including calculation methods and input limitations, affect the way you grade? Do you find technical limitations an acceptable influence on your grading practices? Also, how do you feel about parents essentially having real-time, 24/7 access to their child's grades? Is this an unreasonable or unhealthy demand? I'd love to hear your thoughts!

"Grades...Are a Fantasy"

Chances are, if you're a teacher, you are solely responsible for assigning grades to your students. Rarely is it a pleasant process. You may have taken an evaluation and measurement class as part of your teacher preparation program, but now you're in a real classroom, with real students (who have real parents) and you get to become each child's judge, jury, and in some cases, executioner. It's not one of the easier parts of teaching, because we as teachers pressure ourselves to make our grading as accurate and fair as possible.

But what is "fair"? Just as there is no perfect or even universally-agreed definition of the purpose of education, there is no common understanding on exactly what grades do and what a grade means. Every semester, just when I think I have my grading system figured out, I'll notice students with lower grades that I consider to have mastered content at higher level than some of their higher-scoring classmates. Do I let my professional judgment trump my grading system and change the grade? Why not? After all, I was the one who designed the grading scale in the first place. The grading system is flawed either way.

If grades can create such a quandary for the teacher, imagine the myriad of possible meanings they have for students. Some think grades are earned for simply turning in work. Others think test scores are the only critical factor. Still other students depend on grades to reflect their effort, absent any measured achievement. With so many interpretations, it's a wonder grades hold any meaning at all. I'll never forget the words of my mentor Bob Anderson at Florence High School. "Grades...are a fantasy." He was implying that students often knew so little about what their grade represented and what it took to raise it (especially at the end of a grading period) that grades might as well not be based in any reality at all. I always got this feeling when a student would ask, "If I do well on the test, will it raise my grade?" Not only did they not understand grading, but apparently they didn't understand how averages worked, either.

I've given grading a great deal of thought, probably due to pressure I put on myself as a mathematician to develop a perfectly fair, objective, and accurate grading system, and failing every single time. Consider this post as the first in a series dealing with multiple aspects of grading such as teacher autonomy, influences of grading software, and ethical questions presented from case studies.

As a final thought, I realize my transition from teacher back to student means I'm again receiving grades instead of assigning them. When I got my first graded paper back in my Nature of Mathematics Education class I looked at the letter grade like it was some sort of novelty. In fact, I was surprised it was there at all, and I really didn't care what it was. (Okay, maybe I would have cared if it had been lower.) What I really cared about was the comments left by my professor, and any hints on what I could have explained better. Apparently I gave up on the grading fantasy a long time ago, and I think I'll sleep better knowing so.