1 Purpose and use of a criterion profile
A criterion profile provides a structured description of an evaluation criterion, clarifying what it means, what evidence counts, and how performance should be interpreted. Rather than treating a criterion as a single label, it breaks the concept into understandable components so that different people can apply it in similar ways.
1.1 What a criterion profile measures
A criterion profile measures a specific quality or capability expressed through a criterion. It identifies the criterion’s intent (the aspect of performance being evaluated) and translates it into criteria-relevant language. By pairing performance levels with indicators or evidence types, the profile links abstract expectations to observable outcomes, such as statements in a response, features of an artifact, or behaviors demonstrated during a task.
1.2 Who uses it (assessors, learners, designers)
Criterion profiles are used by several roles across an assessment workflow. Assessor-facing descriptions guide interpretation during scoring. Learner-facing descriptions support understanding of expectations and how feedback relates to improvement. Designers use criterion profiles to define and refine assessment tasks, align them to learning outcomes, and ensure that scoring focuses on the intended dimension.
1.3 Benefits for consistency and fairness
Because a criterion profile specifies how to recognize different performance levels, it supports consistent scoring across evaluators and contexts. This reduces the likelihood that judgments vary due to personal interpretation of vague terms. It can also improve fairness by ensuring that the same evidence rules are applied to all candidates and that feedback reflects the criterion rather than unrelated impressions.
1.4 When to apply a criterion profile
A criterion profile is especially useful when multiple scorers are involved, when assessments are repeated across cohorts, or when feedback must be specific. It is also valuable for complex tasks where performance varies across dimensions and where evidence is not uniform. In practice, criterion profiles are often introduced at the design stage to set expectations before data collection begins.
2 Structure and components
A criterion profile is typically organized into a criterion statement, performance levels, evidence indicators, and scoring guidance. Together these components create a shared interpretation of the criterion’s meaning and the basis for decisions.
2.1 Criterion statement
The criterion statement defines the criterion’s identity and purpose in a concise manner. It serves as the anchor that all later elements explain and operationalize.
2.1.1 Scope and boundaries of the criterion
Scope and boundaries clarify what is included and excluded. This prevents evaluators from drifting into neighboring dimensions. For example, a profile for “argument clarity” may specify that organization is considered only as it affects clarity of claims, rather than as a separate structural quality.
2.1.2 Related skills or learning outcomes
Many criterion profiles explicitly reference the learning outcomes or skills the criterion supports. This connection helps ensure alignment between the assessment and the broader instructional goals, and it also helps learners understand why the criterion matters.
2.2 Performance levels
Performance levels describe how quality changes from lower to higher performance. They are commonly ordered and limited in number (often four to six) to keep interpretation manageable.
2.2.1 Descriptor writing principles
Descriptors should be written so they are interpretable without requiring insider knowledge. Effective descriptors use consistent grammatical structure and focus on the criterion-relevant dimension. They also avoid mixing unrelated qualities within a single level statement, which can blur what differentiates one level from another.
2.2.2 Progressive differentiation between levels
Levels should show progression that is meaningful for decision-making. Each step ideally reflects differences in capability, completeness, or quality of evidence rather than only differences in confidence or effort. Progressive differentiation also helps scorers distinguish “just meets” performance from “exceeds” performance in a way that is recognizable across cases.
2.3 Evidence and indicators
Indicators specify what counts as evidence of the criterion in a learner’s work or behavior. They translate abstract expectations into tangible signals.
2.3.1 Observable behaviors or artifacts
Evidence may include behaviors (e.g., explanation steps during a demonstration) or artifacts (e.g., written sections, design components, or data representations). The profile should state what to look for and where it can typically be found, so evaluators can quickly verify alignment.
2.3.2 Quality markers and common misconceptions
Quality markers describe characteristics that reliably indicate strong performance, such as precision, coherence, or appropriate use of conventions depending on the criterion. Common misconceptions list frequent misinterpretations or partial understandings that often lead to inconsistent scoring. Including these helps evaluators avoid rewarding “surface signals” that do not actually reflect the intended quality.
2.4 Scoring guidance
Scoring guidance translates the profile into a decision process: how evaluators select a performance level and how the scoring system handles variability.
2.4.1 Weighting and point allocation
If a profile contributes points within a rubric, it may specify weights or point ranges per level. Weighting guidance clarifies how strongly the criterion influences the overall score and prevents assumptions that all criteria contribute equally.
2.4.2 Sampling rules (if evidence varies)
When evidence is incomplete or comes from multiple sources (e.g., drafts and final products, or multiple task segments), sampling rules specify which parts should be considered. Sampling guidance improves transparency by stating whether evaluators rely on the single best piece, a representative sample, or evidence aggregated across submissions.
3 Designing a criterion profile
Design involves clarifying the criterion, calibrating levels, ensuring consistent interpretation, and addressing accessibility needs.
3.1 Define the criterion clearly
A clear definition reduces ambiguity and increases the likelihood that the profile will be applied as intended.
3.1.1 Operational definitions
Operational definitions describe the criterion in terms that can be checked through observation or evidence collection. This may involve describing the kind of reasoning expected, the required structure in an artifact, or the behavioral actions that demonstrate the targeted capability.
3.1.2 Avoiding overlap with other criteria
Criterion profiles should be separated by conceptual boundaries. Designers often review other criteria to identify shared content, then adjust wording or evidence indicators so that each criterion rewards what is unique to its dimension. This prevents “double counting” where the same feature influences multiple scores.
3.2 Calibrate level descriptors
Calibration ensures that performance levels are interpreted consistently and correspond to real differences in work quality.
3.2.1 Benchmark examples
Benchmark examples can be provided to illustrate what performance looks like at each level. These examples help scorers internalize the distinctions and reduce variance. When possible, benchmarks should reflect the diversity of contexts and demonstrate both typical and borderline cases.
3.2.2 Trial scoring and revision
Designers commonly conduct trial scoring using sample artifacts. The results highlight where scorers disagree or where levels are indistinct. Profiles are then revised to tighten descriptors, adjust evidence requirements, or redistribute differences across levels.
3.3 Ensure reliability of interpretation
Reliability focuses on whether scorers interpret and apply the profile in a stable way.
3.3.1 Evaluator training prompts
Training can include walkthroughs of the profile, guided practice with feedback, and checklists for evidence identification. Prompts may encourage evaluators to cite which indicators were present and which were absent, improving alignment with the criterion definition.
3.3.2 Consistency checks across scorers
Consistency checks include comparing scoring distributions, running moderation sessions, and reviewing discrepancies. When disagreement occurs, teams typically analyze whether the disagreement is due to ambiguity in descriptors, missing evidence, or differences in evidence interpretation rules.
3.4 Accessibility and inclusivity considerations
Criterion profiles should be usable by learners from diverse backgrounds and by scorers using different support tools.
3.4.1 Plain-language descriptors
Using clear, plain-language terms supports accessibility and reduces cognitive load. Descriptors should avoid discipline-specific jargon unless the assessment context requires it, and they should define terms when technical meaning is unavoidable.
3.4.2 Multiple means of demonstrating evidence
Inclusive design anticipates that learners may express the same underlying capability in different formats. Allowing equivalent evidence forms—such as alternative representations or multiple modalities—can ensure that the criterion measures the intended quality rather than a specific medium.
4 Implementation in assessment practice
Implementation translates the designed profile into day-to-day scoring and feedback, while maintaining coherence across time.
4.1 Using the profile during scoring
Scorers use the profile as a structured reference to assign performance levels and document rationale.
4.1.1 Decision flow for selecting a level
A common workflow is to (1) identify the evidence present, (2) match evidence to indicators and quality markers, (3) compare the evidence profile to performance level descriptors, and (4) select the closest level using the scoring guidance. Some systems include a brief justification step that anchors the decision to named indicators.
4.1.2 Handling borderline cases
Borderline cases occur when evidence partially matches descriptors from adjacent levels. Profiles handle these situations by specifying “minimum evidence” thresholds, clarifying how to treat missing elements, and indicating whether evaluators should prioritize the strongest indicator or require a balanced set of indicators.
4.2 Providing feedback using descriptors
Feedback should be closely tied to the criterion profile so that learners can interpret what to change.
4.2.1 Linking feedback to specific indicators
Effective feedback references the indicators and quality markers that were present or absent. Instead of general commentary, it points to the exact criterion dimension that needs improvement, such as adding missing reasoning steps, revising for clarity, or adjusting the use of conventions relevant to the criterion.
4.2.2 Actionable next steps for learners
Feedback becomes more useful when it includes suggestions that directly address the profile’s evidence gaps. These next steps align with the learning outcomes and help learners target behaviors or artifact features likely to move them to a higher performance level.
4.3 Managing rubric drift over time
Over time, scoring systems can shift due to changes in tasks, interpretations, or training practices. Criterion profile management mitigates this “drift.”
4.3.1 Version control and update cycles
Profiles typically have version identifiers and documented change histories. Updates often occur after reviewing scoring outcomes, moderation findings, or evidence that learners are being assessed under new task formats. Controlled revisions help preserve comparability across cohorts.
5 Examples and templates
Templates and examples illustrate common ways to present criterion profiles and how they vary by assessment type and intended use.
5.1 Short criterion profile template
A short template typically includes: a one-sentence criterion statement, a brief scope statement, four performance levels with short descriptors, and a short list of key evidence indicators. It is suited for assessments with limited time for scoring or where the criterion is narrow.
5.2 Multi-level (e.g., 4–6) template
A multi-level template adds more granularity, often using five levels to distinguish “approaching” and “exceeding” performance. It includes more detailed indicators and may specify example evidence types for each level to reduce scorer ambiguity.
5.3 Criterion profile for formative assessment
Formative profiles emphasize learning and coaching. They commonly use descriptors and indicators that identify the next improvement step, and they may include guidance for self-assessment or peer review based on the same evidence rules used by scorers.
5.4 Criterion profile for summative assessment
Summative profiles prioritize consistency and defensible scoring. They may include clearer minimum thresholds per level and explicit rules for decision-making when evidence is incomplete, because scores often carry evaluative consequences.
5.5 Example: writing-based criterion profile
In a writing context, the criterion profile might define “argument clarity” with scope boundaries (e.g., clarity of claim and reasoning, not grammar quality). Performance levels could range from unclear or unsupported claims to well-reasoned arguments that anticipate objections. Evidence indicators could include the presence of a stated claim, the logical linkage between reasons and conclusions, and the use of relevant examples; quality markers might highlight coherence and avoidance of contradictory statements. Scoring guidance may specify that the evaluator uses the latest draft or final submission, depending on the assessment design.
5.6 Example: project-based criterion profile
For a project, a criterion profile might address “design feasibility.” It could define feasibility as practical alignment between requirements, constraints, and proposed implementation steps. Performance levels could differentiate conceptual plausibility from fully reasoned implementation planning. Evidence indicators might include a stated set of constraints, a step-by-step plan, identification of risks, and evidence that proposed solutions can be executed with available resources. Sampling rules may require evaluation of multiple project artifacts, such as proposal, prototype documentation, and final presentation.
6 Common pitfalls and quality checks
Even well-designed profiles can fail if they are unclear or inconsistently applied. Quality checks help identify these risks early.
6.1 Vague or overly broad descriptors
Vagueness makes levels difficult to distinguish and invites idiosyncratic scoring. Overly broad descriptors may cause evaluators to reward related features that are not truly part of the criterion. Quality checks often involve asking scorers to interpret each descriptor and comparing their interpretations against the criterion definition.
6.2 Misaligned criteria and learning goals
If the criterion does not correspond to the intended learning outcomes, the assessment measures something else—often the easiest-to-observe feature rather than the targeted capability. Alignment review checks whether instructional objectives, tasks, and indicators all refer to the same underlying construct.
6.3 Unequal level spacing
Levels that differ by uneven “distance” can lead to scoring compression, where many responses cluster in one band, or to inflated penalties near thresholds. Quality checks may include analyzing score distributions and reviewing whether improvements between adjacent levels represent distinct performance changes.
6.4 Inconsistent evidence requirements
When evidence rules vary across levels without explanation, scorers may apply different standards. Inconsistent evidence requirements can also cause confusion in borderline cases. Teams typically confirm that indicators and thresholds are stated consistently and that the profile’s evidence list matches the intended criterion scope.
6.5 Over-reliance on a single indicator
Profiles sometimes include a dominant indicator that becomes a proxy for the entire criterion. This can distort measurement, particularly when learners demonstrate strength through alternate pathways. Quality checks may involve evaluating whether multiple indicators are necessary for high-level performance and whether the profile fairly represents the full construct.
7 Related concepts
Criterion profiles connect to several broader assessment ideas. Understanding these helps place criterion profiles within the larger evaluation landscape.
7.1 Rubrics vs. criterion profiles
A rubric typically provides a scoring system across criteria and levels, often with point values. A criterion profile focuses on describing a single criterion in depth—its meaning, evidence, and level interpretation—so that scoring can be applied consistently. In practice, criterion profiles can function as components within broader rubrics.
7.2 Analytic vs. holistic assessment
Analytic approaches evaluate separate criteria or dimensions, allowing more diagnostic feedback. Holistic approaches judge overall quality as a single integrated judgment. Criterion profiles support analytic assessment by specifying what each dimension looks like, though similar descriptor logic can also inform holistic judgments when used carefully.
7.3 Learning outcomes mapping
Learning outcomes mapping links assessment tasks and scoring criteria to specific learning goals. Criterion profiles contribute to this process by clarifying how each goal is reflected in evidence and performance levels, ensuring the assessment targets the intended learning.
7.4 Validity, reliability, and transparency (assessment basics)
- Validity concerns whether the criterion profile measures the intended construct. Clear scope, aligned indicators, and coherent evidence rules support validity.
- Reliability concerns consistent interpretation across scorers and contexts. Calibration, training prompts, and evidence-based decision flows strengthen reliability.
- Transparency concerns whether learners and evaluators can understand how judgments are made. Plain-language descriptors, explicit indicators, and feedback tied to the profile improve transparency.