A product demonstration answers one question: can the supplier show the system working?
An institutional pilot answers a harder one: does the complete process work for our assessments, people, policies, systems, and risk level?
Those are not the same test.
A useful pilot needs a baseline, representative material, named decision-makers, predefined measures, difficult cases, and permission to stop. Without those elements, the pilot becomes a guided tour that happens to use institutional data.
This 90-day AI grading pilot checklist is built for universities, school groups, examination boards, coaching networks, certification bodies, and corporate academies. It can be used with any supplier. The scorecard is deliberately vendor-neutral.
TL;DR: UNSW evaluates education technology in 3 stages: vendor assessment, a pre-pilot score sheet, and an ongoing pilot. Use the same discipline for AI grading. Over 90 days, define authority, establish a human baseline, test a representative sample, examine edge cases, verify integrations, and record a formal decision. A successful pilot can end with “not yet.”
Why run a 90-day pilot instead of a longer demonstration?
UNSW's technology evaluation framework uses three stages: initial vendor assessment, a pre-pilot score sheet, and an ongoing full pilot. Its full evaluation looks at pedagogical, business, and technical factors. That separation prevents a polished demonstration from standing in for institutional evidence. For related product context, use GradeLab's AI grading guide.
Ninety days is not a scientific constant. It is a workable management window.
It is long enough to set governance, build a representative test set, run parallel evaluation, test difficult cases, and make a documented decision. It is short enough to keep the scope bounded. An institution with one assessment cycle may need less time. A national board or regulated certification body may need more.
The duration matters less than the sequence.
A pilot should test one clearly defined use case, such as preliminary scoring support for short answers, on-screen review of handwritten scripts, or rubric-aligned feedback for a low-stakes assessment. Testing several unrelated use cases at once makes failure difficult to diagnose.
Write the decision before the configuration begins:
Book a demo and get 90-day trial, free migration, and locked-in 2026 pricing.
Join thousands of educators and institutions using GradeLab's AI-powered grading platform to save time, ensure accuracy, and provide instant feedback to students.
Save 100+ hours per month
AI-powered accuracy & consistency
Instant student feedback
Easy LMS integration
AI Grading Pilot Checklist: A 90-Day Scorecard - GradeLab
“We are testing whether this assisted workflow can meet our stated validity, review, privacy, operational, and integration requirements for this assessment under these conditions.”
That sentence stops the pilot from quietly expanding.
What must be decided before day one?
The NIST AI Risk Management Framework places governance across four functions: govern, map, measure, and manage. It calls for documented roles, executive responsibility, testing, monitoring, and feedback routes. A pilot that begins before authority is clear is already collecting weak evidence. Review GradeLab's privacy information during the product-level governance check.
Complete this pre-pilot record:
Decision
Required answer
Intended use
The exact task the system may assist
Excluded use
Tasks, subjects, decisions, or data outside scope
Assessment owner
Person accountable for validity and academic standards
Technical owner
Person accountable for configuration, access, and integration
Data owner
Person accountable for lawful processing, retention, and deletion
Final authority
Person or body that approves marks and results
Pilot population
Defined assessment, subjects, languages, and response types
Baseline
Current human workflow and measures used for comparison
Stop conditions
Events that pause or end the pilot
Appeal route
How affected people obtain explanation and human review
The stop conditions deserve attention. Examples include unauthorized data exposure, inability to reconstruct a score, a material difference affecting one group, loss of the review trail, or performance outside a threshold set before testing.
Do not copy another institution's thresholds without context. The stakes, scoring scale, human baseline, and legal duties may differ.
Seven safety gates
The pilot should not proceed to live use if any gate remains unresolved:
Legal and policy authority
Assessment validity
Human oversight and final authority
Privacy and data governance
Fairness and accessibility
Security and operational continuity
Explanation, recheck, and appeal
These are gates, not points. A strong usability score cannot compensate for unlawful data processing or the absence of an appeal route.
Days 1 to 10: Set scope, authority, and safety gates
The University of Utah's 2026 guidance contains ten sections and states that faculty retain full responsibility for grades. It requires human review, approved tools, de-identification, pilot testing, ongoing monitoring, transparency, and access to a full human review. GradeLab's frequently asked questions can support the product-specific review.
Local policy will differ, but the opening work is similar.
Days 1 to 3: Define the use case
Choose one workflow. Describe its beginning and end.
Bad scope: “Test automated grading across the institution.”
Better scope: “Test preliminary rubric-aligned scoring for five-mark, English-language biology responses in one first-year module, with every suggestion reviewed before release.”
The better version identifies the subject, response type, score value, language, cohort, and human decision point.
Days 4 to 6: Map the data
Record every data element the pilot will process:
Student identifiers
Answer content
Sensitive or reflective material
Rubrics and model answers
Human marks and comments
System suggestions and confidence indicators
Reviewer actions
Final marks
Technical and audit logs
Then record where each element is collected, transferred, stored, accessed, retained, and deleted. The applicable law may include FERPA, GDPR, India's DPDP framework, POPIA, or another national regime. Legal obligations depend on jurisdiction, role, contract, and use.
Days 7 to 10: Approve the plan
Ask the assessment owner, data or privacy lead, technical owner, and project sponsor to sign the same scope. If the pilot involves regulated qualifications, include the responsible compliance or quality function.
No signature, no pilot.
This may feel slow. It is much quicker than discovering halfway through testing that the sample cannot lawfully be used or that nobody has authority to accept the result.
Days 11 to 30: Build the baseline and test set
Ofqual's 2026 working paper says evidence must be specific to the qualification, candidate population, construct, and item type. It also says agreement with human marks alone is insufficient. The test set therefore needs to represent the real decision, not just produce a convenient accuracy figure. GradeLab's paper grading overview can help teams define handwritten response and scan conditions.
Measure the current process first
Record what happens without the new system:
Time from submission to approved result
Active marking and moderation time
Human-to-human agreement, where double marking exists
Recheck and appeal rate
Number and type of handling errors
Missing or unreadable submissions
Support workload
Marker training and standardization effort
Do not turn every baseline into a financial estimate. Some measures describe quality, control, or resilience rather than cost.
Select a representative sample
There is no universal sample size for every assessment pilot. A statistically defensible design depends on the expected variation, scoring scale, subgroup analysis, risk, and decision being made. Ask a psychometrician, statistician, or qualified assessment specialist where the stakes justify it.
At minimum, the sample should cover:
Low, middle, high, and borderline score bands
Common and uncommon valid answers
Typical and poor scan quality
Different question types in scope
Languages and scripts in scope
Relevant accommodations
Answers with corrections, marginal notes, diagrams, or alternative reasoning
Known difficult cases from previous cycles
Keep a separate challenge set. Do not use it while adjusting the configuration. Test it only after the workflow is stable. Otherwise, the team may tune the system to the test.
Freeze the comparison plan
Define the measures before seeing pilot results. Useful measures include:
Exact agreement: proportion of assisted and reference scores that match exactly
Adjacent agreement: proportion within a defined score distance
Mean absolute difference: average absolute gap between two scores
Score-band agreement: agreement on classifications or grade bands
Override rate: proportion of suggestions changed by a reviewer
Escalation rate: proportion routed to specialist review
Subgroup difference: variation in performance across relevant groups
Processing completion: proportion that completes the full workflow without manual repair
No single measure proves validity. A high exact-agreement rate can still hide systematic errors at grade boundaries.
Days 31 to 60: Run controlled parallel evaluation
NIST's Measure 2.3 says performance or assurance criteria should be demonstrated in conditions similar to deployment and documented. During days 31 to 60, run the proposed workflow beside the existing process. Do not let system suggestions influence the reference marks unless that influence is part of the design being tested.
Keep the comparison independent
Where practical, have qualified markers score the sample without seeing the system suggestion. Then compare:
Human reference score
System suggestion
Final reviewed score
Reason for any change
This separates model performance from the quality of the human-system workflow.
It also reveals automation bias. A reviewer may accept a poor suggestion because it arrives with polished feedback or a precise-looking confidence value. Conversely, a reviewer may reject a good suggestion simply because it came from a machine. Both effects matter.
Record disagreement by type
Do not collect only the size of the difference. Classify why it occurred:
Text recognition error
Rubric ambiguity
Missed valid alternative
Incorrect partial credit
Weak interpretation of evidence
Language or script issue
Diagram or mathematical notation issue
Reviewer mistake
Configuration problem
Unexplained variation
Patterns lead to decisions. An average does not.
Review the workflow, not only the score
Watch real users complete the task. Can they see the original response clearly? Can they find the rubric? Can they understand why an item was flagged? Can they override a suggestion without losing the original record? Can a moderator reconstruct what happened?
The pilot team should meet weekly during this phase. Record changes to prompts, rubrics, models, settings, and procedures. If the system changes materially, note which results were produced under which version.
Days 61 to 75: Test difficult cases and institutional fit
LTI Advantage currently includes three services: Assignment and Grade Services, Deep Linking, and Names and Role Provisioning Services. 1EdTech also distinguishes certified implementations from suppliers that merely say they conform. Integration claims need evidence just as scoring claims do. Use GradeLab's digital grading overview to frame workflow questions.
Test the full technical journey
Use a sandbox where possible. Test:
User identity and role mapping
Assignment creation or import
Submission transfer
Grade and feedback return
Corrected grades
Duplicate and missing records
Permission changes
Network interruption
Retry behavior
Export and reconciliation
Audit history
Check the exact limitations of each target system. For example, Google Classroom API documentation states that developers can set numerical grades, but rubric scores can be read and not set through the API as of April 2026. Compare those requirements with GradeLab's digital grading workflow.
An integration can work in a demonstration and still fail the institutional workflow.
Test privacy, security, and deletion
Verify contractual and technical claims. Ask the supplier to show:
Hosting and data locations
Access controls and administrator logs
Encryption arrangements
Subprocessors
Model training and data-use terms
Retention settings
Backup and deletion behavior
Incident notification process
Export on contract termination
Then test the controls available to the institution. A policy document is not the same as an observed result.
Test accessibility with real tasks
The W3C WCAG overview is the international starting point for web accessibility. A pilot should go beyond a checkbox. Test keyboard use, focus order, labels, contrast, zoom, screen-reader behavior, error messages, time limits, and the accessibility of uploaded documents. Record unresolved product questions for GradeLab's frequently asked questions or pilot discussion.
Include people who use assistive technology. Automated checks catch only part of the problem.
Run the failure drill
Choose one realistic failure. The LMS is unavailable. A batch is indexed incorrectly. The model version changes. A set of scripts cannot be read. A user exports the wrong cohort.
Can the team stop processing, identify affected records, recover safely, communicate the issue, and continue without losing the audit trail?
If not, the pilot has found something valuable before live deployment.
Days 76 to 90: Score the evidence and make a decision
NIST defines Manage as an ongoing function, not a final project stage. Days 76 to 90 should therefore produce two outputs: a decision about the pilot and an operating plan for whatever happens next. Adoption without monitoring is not a completed evaluation.
The 100-point pilot scorecard
The following weights are a GradeLab editorial framework. They are not a regulatory standard or validated psychometric instrument. Change them before the pilot if your institution's risk profile demands it.
Turnaround, active effort, capacity, error handling compared with baseline
Supplier evidence and support
5
Documentation, response quality, version notice, issue resolution
Total
100
Score each dimension from 0 to its assigned weight and attach the evidence. A number without notes is not a decision record.
Suggested decision bands
These bands are a starting point, not a universal rule:
80 to 100: Consider a limited rollout only if all seven safety gates pass.
65 to 79: Extend the pilot, narrow the use case, or correct identified gaps.
Below 65: Stop or redesign before further deployment.
A high total cannot override a failed gate. Likewise, a lower score may reflect a fixable integration problem rather than invalid scoring. Read the pattern, not just the total.
Produce a signed decision record
The record should state:
What was tested and excluded
Data and versions used
Results by measure and subgroup
Known limits
Incidents and unresolved issues
Changes made during the pilot
Safety-gate decisions
Scorecard result
Final decision and reasons
Conditions for any rollout
Monitoring and review schedule
Possible decisions include adopt, adopt for a narrower use, extend the pilot, redesign the process, test another supplier, or stop.
“Stop” is not a failed pilot. It is evidence doing its job.
Where can GradeLab fit into this pilot?
The University of Utah guidance requires a representative small-scale comparison before AI-assisted grading is used. GradeLab can be evaluated under the same principle: one approved use case, one representative sample, one human-reviewed workflow, and measures set by the institution.
GradeLab is built for assessment workflows that may include handwritten and digital responses. Depending on the verified setup, a pilot can examine intake, rubric-based assistance, review queues, overrides, analytics, and result handling. None of those capabilities removes the need for local validation.
The best demonstration is your difficult case, not ours.
Frequently asked questions about AI grading pilots
The NIST AI RMF contains four connected functions and treats monitoring as an ongoing responsibility. That makes a pilot the beginning of evidence collection, not a permanent approval for every future model, subject, or cohort.
How long should an AI grading pilot run?
There is no universal duration. This framework uses 90 days because it accommodates five phases from governance through decision. Match the schedule to the assessment cycle, risk, sample design, and approval process. Preserve the sequence even if the calendar changes.
How many responses should the pilot include?
No single number fits every assessment. Sample size depends on score variation, intended analysis, subgroups, item types, and stakes. Ofqual says evidence should be specific to the qualification, candidate population, construct, and item type. Seek statistical advice for high-stakes decisions.
What accuracy should an AI grading system achieve?
Do not accept one universal percentage. Define exact agreement, adjacent agreement, score difference, boundary performance, override rate, and subgroup behavior. Compare results with the existing human process and the institution's risk tolerance. NIST recommends testing under conditions similar to deployment.
Can pilot data contain student names?
Follow applicable law and institutional policy. The University of Utah requires full de-identification even for approved tools. Other jurisdictions may impose different duties. Use the minimum data required, record the lawful basis, and confirm contractual controls before transfer.
Does passing a pilot prove future performance?
No. NIST calls for ongoing monitoring, feedback, and risk management. A pilot supports a decision for the tested conditions. New subjects, languages, rubrics, models, integrations, or candidate populations may require added validation.
End with evidence, not enthusiasm
NIST's four functions provide the final test: govern, map, measure, and manage. The 90-day structure turns those ideas into a practical institutional sequence.
Decide who is accountable. Map the assessment and its data. Measure the complete workflow against a real baseline. Test what happens when the system is wrong. Then manage the limits after launch.
If you want to evaluate GradeLab, bring one assessment, one representative de-identified sample, and your own success criteria. We can use this scorecard to structure the conversation.