ClaytonLead at Scale · White Paper · June 2026

How Clayton Scores:
Rubric-Based AI Assessment Explained

A score without a rubric is just a number. Here's how Clayton turns a leader's standards into a structured, repeatable scoring system — and how that system translates into measurable improvement over time.

The TL;DR (Short Version)

Clayton doesn't score work the way a spell-checker does — by checking for rule violations. It scores the way a trained evaluator does: by reading the work against an explicit rubric, dimension by dimension, and assigning a defensible 1–5 rating to each.

The result is a score the leader can trust, a score the team member can learn from, and a dataset that shows whether things are improving. This paper explains the mechanics: how the rubric is built, how the AI applies it, how the final score is calculated, and how ROI is measured.

Step 1 — The Rubric

The scoring foundation: structured dimensions, each with explicit level descriptions.

Every Clayton is built around a rubric — a set of named dimensions that define what "good" looks like for a given deliverable. A strategy Clayton might score on Problem Definition, Differentiation, Measurable Objectives, and Stakeholder Clarity. A feedback Clayton might score on Specificity, Impact Framing, and Actionability.

Each dimension has five levels — not just a number, but a written description of what a 1 looks like versus a 3 versus a 5. This is the key difference between rubric-based scoring and prompt-based scoring: the AI isn't left to decide what "good" means. The leader has already decided, and the AI's job is to apply that judgment, not form its own.

Example Rubric Dimension: Measurable Objectives

From a Planning/OKRs Clayton

1
No measurable objectives

Goals are described in purely qualitative terms. No metrics, baselines, or thresholds are specified.

2
Vague metrics

Some numeric references exist but lack baselines, timeframes, or clear ownership.

3
Partially measurable

Most objectives have metrics, but one or more are still qualitative or lack a verifiable baseline.

4
Measurable with minor gaps

All key results are numeric and time-bound. Minor gaps in baseline data or ownership.

5
Fully objective and verifiable

Every key result has a numeric target, a current baseline, a timeframe, and an owner. No subjective language anywhere.

Each dimension has five explicit level definitions. The AI must quote evidence before scoring — vague impressions are not permitted.

Step 2 — How the AI Applies the Rubric

Evidence-first scoring: the AI must show its work.

When a team member submits a piece of work, Clayton reads the document against each rubric dimension in sequence. The process is deliberately constrained:

01
Quote before scoring

The AI must cite specific text from the submitted document as evidence for its score. If no relevant text exists, it must say so explicitly — and score accordingly. This prevents the model from hallucinating rationale for a score it has already decided to give.

02
Apply the level definition

The score must match the level description for that dimension — not the AI's general sense of quality. A "3" means the document meets the criteria described in the Level 3 definition, nothing more and nothing less.

03
Downward ambiguity resolution

When evidence is mixed or incomplete, the score goes down, not up. This is an explicit instruction that overrides the model's natural tendency toward generosity. A borderline 3/4 is resolved as a 3.

04
No invented dimensions

The AI scores exactly the dimensions defined in the rubric. It cannot add new categories, rename existing ones, or skip any. This ensures every review is structurally comparable.

Step 3 — Calculating the Final Score

From individual dimension ratings to a single comparable number.

Each rubric dimension is scored 1–5. The final score is computed by averaging the dimension scores, then normalizing to a 0–100 scale. This means a document that scores 3/5 on every dimension receives a 60/100 overall — a reliable midpoint that reflects consistent, developing work.

Score Calculation — Worked Example

Planning/OKRs submission scored across 5 dimensions

Problem/Opportunity Clarity
4/5
Measurable Objectives
2/5
Prioritisation & Focus
4/5
Stakeholder Alignment
3/5
Feasibility & Risk
2/5
Average: 3.0 / 5.0
60/100

Formula: (sum of dimension scores ÷ max possible score) × 100. Each dimension is weighted equally unless the Clayton creator has applied custom weights.

80–100
High-quality

Meets or exceeds the leader's standard. Ready to share or submit.

60–79
Developing

Solid foundation with clear gaps. Specific improvements will move the score.

0–59
Needs work

Structural issues. Focused revision on the lowest-scoring dimensions is the priority.

What Scores Look Like in Practice

Individual progress and team performance — two views of the same data.

The value of a score isn't the number itself — it's the pattern the numbers reveal over time. A single score tells a team member where they stand today. A series of scores tells them whether they're moving in the right direction, and at what rate.

Average Personal Skill Score Over Time

76/100+34 pts

38 scored workstreams across 4 Claytons · avg +1.7 pts/session

Progress over time
Jun 2Jun 3Jun 5Jun 6Jun 8Jun 9Jun 12Jun 15Jun 17Jun 20Jun 22Jun 24Jun 260255075100
  • Overall Average
  • Strategy Advisor
  • Feedback Giving Coach
  • Planning/OKRs Coach
  • Accountability Agreement Coach

Score History by Clayton

Every scored workstream — click a legend item to isolate one Clayton

7d30d90dAll
76
Team Avg Score
17
Total Scores
3
Active Members
Jun 5Jun 10Jun 15Jun 20Jun 250255075100

Leaders see every team member's score trajectory. Members see their own. Both views show the same underlying data — from different vantage points.

Measuring ROI and Efficiency

How Clayton translates scoring data into a financial return calculation.

Score improvement is the qualitative outcome. ROI is the financial one. Clayton tracks both — connecting the coaching activity directly to time saved and cost avoided.

ROI Calculation — How It Works

Illustrative example for a single deliverable submission

Member Draft Time
4h
self-reported
Leader Review Time
0.5h
vs. 2h without Clayton
Time Saved
1.5h
leader hours recovered
The Calculation
Expected leader hours (without Clayton)

Set by leader when creating the Clayton

2.0 hrs
Actual leader hours (with Clayton)

Entered by leader after reviewing the submission

0.5 hrs
Hours saved per submission
1.5 hrs
Leader hourly rate (blended)

Set by the organization or leader

$150/hr
Value recovered per submission
$225

At Scale: 10 team members, 1–2 submissions/month each

15
Submissions/month
22.5 hrs
Hours saved/month
$3,375
Value recovered/month
Member self-reporting

After each submission, members log how long the initial draft took and how long revisions took. This separates the time cost of creation from the time cost of rework — both of which Clayton influences.

Leader entry

Leaders log their actual review time per submission. The system compares this to the expected time set when the Clayton was created — the delta is the efficiency gain. Over a quarter, the pattern becomes statistically meaningful.

Draft count matters

Clayton tracks how many versions of a document were scored in a session. Fewer iterations to reach the standard = a more efficient coaching loop. Score improvement per iteration is the leading indicator of genuine skill development.

Aggregate dashboard

The ROI & Efficiency tab aggregates across all Claytons, all members, and all time periods. Leaders can see total hours saved, cost recovered, and which Claytons are delivering the most impact — by Clayton, by member, or by deliverable type.

Why Scoring Consistency Is the Product

The most common failure mode in AI-assisted coaching is inconsistency: a team member submits the same document twice and gets different scores, or two members submit similar work and receive scores that can't be compared. Once the team notices, trust collapses.

Rubric-based scoring addresses this at the architecture level. Because every dimension has an explicit definition, because the AI must quote evidence before scoring, and because ambiguity is always resolved downward, the variance is structural rather than random. Two reviewers — human or AI — reading the same document against the same rubric should arrive at the same score. That reproducibility is what makes the data useful, the feedback credible, and the ROI calculation defensible.

Scoring model
Rubric-locked
No freeform interpretation
Evidence requirement
Quote-first
AI cites text before scoring
Ambiguity rule
Score down
Conservative by default
Clayton

See the scoring in action

Start free and build your first rubric in under 10 minutes — or schedule a free consultation to discuss how scoring can work for your team's specific deliverables.

or email clayton@sagely.ltd

© 2026 Sagely Advisory LLC. All rights reserved.Lead at Scale.