Lead at Scale · White Paper · June 2026A score without a rubric is just a number. Here's how Clayton turns a leader's standards into a structured, repeatable scoring system — and how that system translates into measurable improvement over time.
Clayton doesn't score work the way a spell-checker does — by checking for rule violations. It scores the way a trained evaluator does: by reading the work against an explicit rubric, dimension by dimension, and assigning a defensible 1–5 rating to each.
The result is a score the leader can trust, a score the team member can learn from, and a dataset that shows whether things are improving. This paper explains the mechanics: how the rubric is built, how the AI applies it, how the final score is calculated, and how ROI is measured.
The scoring foundation: structured dimensions, each with explicit level descriptions.
Every Clayton is built around a rubric — a set of named dimensions that define what "good" looks like for a given deliverable. A strategy Clayton might score on Problem Definition, Differentiation, Measurable Objectives, and Stakeholder Clarity. A feedback Clayton might score on Specificity, Impact Framing, and Actionability.
Each dimension has five levels — not just a number, but a written description of what a 1 looks like versus a 3 versus a 5. This is the key difference between rubric-based scoring and prompt-based scoring: the AI isn't left to decide what "good" means. The leader has already decided, and the AI's job is to apply that judgment, not form its own.
From a Planning/OKRs Clayton
Goals are described in purely qualitative terms. No metrics, baselines, or thresholds are specified.
Some numeric references exist but lack baselines, timeframes, or clear ownership.
Most objectives have metrics, but one or more are still qualitative or lack a verifiable baseline.
All key results are numeric and time-bound. Minor gaps in baseline data or ownership.
Every key result has a numeric target, a current baseline, a timeframe, and an owner. No subjective language anywhere.
Each dimension has five explicit level definitions. The AI must quote evidence before scoring — vague impressions are not permitted.
Evidence-first scoring: the AI must show its work.
When a team member submits a piece of work, Clayton reads the document against each rubric dimension in sequence. The process is deliberately constrained:
The AI must cite specific text from the submitted document as evidence for its score. If no relevant text exists, it must say so explicitly — and score accordingly. This prevents the model from hallucinating rationale for a score it has already decided to give.
The score must match the level description for that dimension — not the AI's general sense of quality. A "3" means the document meets the criteria described in the Level 3 definition, nothing more and nothing less.
When evidence is mixed or incomplete, the score goes down, not up. This is an explicit instruction that overrides the model's natural tendency toward generosity. A borderline 3/4 is resolved as a 3.
The AI scores exactly the dimensions defined in the rubric. It cannot add new categories, rename existing ones, or skip any. This ensures every review is structurally comparable.
From individual dimension ratings to a single comparable number.
Each rubric dimension is scored 1–5. The final score is computed by averaging the dimension scores, then normalizing to a 0–100 scale. This means a document that scores 3/5 on every dimension receives a 60/100 overall — a reliable midpoint that reflects consistent, developing work.
Planning/OKRs submission scored across 5 dimensions
Formula: (sum of dimension scores ÷ max possible score) × 100. Each dimension is weighted equally unless the Clayton creator has applied custom weights.
Meets or exceeds the leader's standard. Ready to share or submit.
Solid foundation with clear gaps. Specific improvements will move the score.
Structural issues. Focused revision on the lowest-scoring dimensions is the priority.
Individual progress and team performance — two views of the same data.
The value of a score isn't the number itself — it's the pattern the numbers reveal over time. A single score tells a team member where they stand today. A series of scores tells them whether they're moving in the right direction, and at what rate.
Average Personal Skill Score Over Time
38 scored workstreams across 4 Claytons · avg +1.7 pts/session
Score History by Clayton
Every scored workstream — click a legend item to isolate one Clayton
Leaders see every team member's score trajectory. Members see their own. Both views show the same underlying data — from different vantage points.
How Clayton translates scoring data into a financial return calculation.
Score improvement is the qualitative outcome. ROI is the financial one. Clayton tracks both — connecting the coaching activity directly to time saved and cost avoided.
Illustrative example for a single deliverable submission
Set by leader when creating the Clayton
Entered by leader after reviewing the submission
Set by the organization or leader
At Scale: 10 team members, 1–2 submissions/month each
After each submission, members log how long the initial draft took and how long revisions took. This separates the time cost of creation from the time cost of rework — both of which Clayton influences.
Leaders log their actual review time per submission. The system compares this to the expected time set when the Clayton was created — the delta is the efficiency gain. Over a quarter, the pattern becomes statistically meaningful.
Clayton tracks how many versions of a document were scored in a session. Fewer iterations to reach the standard = a more efficient coaching loop. Score improvement per iteration is the leading indicator of genuine skill development.
The ROI & Efficiency tab aggregates across all Claytons, all members, and all time periods. Leaders can see total hours saved, cost recovered, and which Claytons are delivering the most impact — by Clayton, by member, or by deliverable type.
The most common failure mode in AI-assisted coaching is inconsistency: a team member submits the same document twice and gets different scores, or two members submit similar work and receive scores that can't be compared. Once the team notices, trust collapses.
Rubric-based scoring addresses this at the architecture level. Because every dimension has an explicit definition, because the AI must quote evidence before scoring, and because ambiguity is always resolved downward, the variance is structural rather than random. Two reviewers — human or AI — reading the same document against the same rubric should arrive at the same score. That reproducibility is what makes the data useful, the feedback credible, and the ROI calculation defensible.

See the scoring in action
Start free and build your first rubric in under 10 minutes — or schedule a free consultation to discuss how scoring can work for your team's specific deliverables.
or email clayton@sagely.ltd