Participatory Evaluation for Researchers: Super Team Guide | Issue 8 of 12

Who's This For

You run workshops where participants produce something: a broader impacts draft, a mentoring plan, a data management statement, a partnership design document. Afterward, you need to assess quality. Maybe you score the work yourself. Maybe you hire an external evaluator. Maybe you collect everything and nobody ever scores it at all.

You know the assessment process could be more useful. Participants submit their work and wait. They get a score or a paragraph of feedback. They revise or they do not. The scoring happened to them. They were not part of it.

This issue is for you if you have ever felt that assessment is a missed learning opportunity, that the act of evaluating quality could itself build the skills your program is trying to develop. You are about to turn scoring from something done to participants into something done with and by participants.

The Partnership Moment

You have just finished a three-session workshop series on writing broader impacts sections for NSF proposals. Twenty-two participants (faculty from across your institution) have submitted drafts. Your evaluation plan says you will have two external reviewers score each draft on a rubric and report aggregate quality scores in your annual report.

The reviewers cost money. They are busy. They return scores three weeks later, a number for each criterion, a sentence or two of feedback. You compile the data. Average clarity score: 3.2 out of 5. Average specificity: 2.8. Average feasibility: 3.5. Average NSF alignment: 2.6. You report these numbers to your funder. The numbers tell you that participants struggled most with specificity and NSF alignment. They do not tell you why.

Meanwhile, your participants received their scores, glanced at the feedback, and went back to their departments. A few revised their drafts. Most filed the feedback and moved on. The assessment happened at them. They experienced it as a verdict, not as a learning event.

Now imagine a different version. Instead of sending drafts to external reviewers, you make the scoring itself a workshop activity. Participants score each other's drafts. They compare their ratings. They argue about what "specificity" means. They discover that they disagree about NSF alignment because they understand NSF's criteria differently. The calibration conversation is the intervention. The scoring process is the professional development.

And you get better data. Not just average scores, but variance data: where participants disagree most, which concepts remain unclear, which instructional gaps your workshop has not yet addressed.

Under the Surface

Traditional assessment in professional development programs treats participants as subjects and evaluators as authorities. The expert scores. The participant receives the score. The knowledge flows one direction. This model works when the goal is certification or ranking, but it actively undermines learning when the goal is skill development.

The problem is threefold.

First, external scoring misses the learning moment. The act of evaluating quality (deciding whether a draft is clear, specific, feasible, aligned) requires the same skills that writing a good draft requires. When an external evaluator does this work, participants never develop those evaluation skills. They remain dependent on experts to tell them whether their work is good.

Second, external scoring produces thin data. A score of 3.2 on clarity tells you the average quality level. It does not tell you why clarity is low. Is it because participants do not know what clarity means in this context? Is it because they know but cannot execute? Is it because the rubric definition of clarity does not match what the field actually values? External scores measure the outcome but not the mechanism.

Third, external scoring scales poorly. Two reviewers scoring twenty-two drafts is expensive and slow. If you want richer feedback (multiple perspectives per draft, discussion of disagreements, calibrated standards), external review becomes logistically impossible.

Collaborative scoring solves all three problems. Participants develop evaluation skills by practicing them. The calibration discussion reveals why scores differ, not just that they differ. And peer review generates three or four reviews per draft instead of one or two, at no additional cost.

The deeper principle is this. Assessment moves from something done TO participants into something done WITH and BY participants. The process itself becomes the intervention.