ClaytonLead at Scale · White Paper · May 2026

Building An AI Coach Is Easy
(Except for the Hard Bits)

The concept takes an afternoon. The content curation takes weeks. The distribution requires infrastructure you didn't plan for. And getting the AI to consistently behave the way you want? That's a problem that never fully goes away.

The TL;DR (Short Version)

Anyone can build a passable AI coach in a weekend. Describe your standards in a system prompt, upload a few documents, share the link. Done — in the way that a house is "done" when the walls are framed and the roof is on.

What comes next is where the work actually lives: curating the knowledge that makes the coach credible, distributing it to a team in a way that's authorized and measurable, and — the part most builders underestimate — getting the underlying language model to follow your rules instead of its own.

This paper describes what we've learned building Clayton: the three layers of genuine difficulty that separate a promising prototype from a coaching system that works reliably at scale.

The Three Layers of Genuine Difficulty

Each layer compounds the one before it. You can skip none of them.

01

Content

Finding and curating the right knowledge

Harder than expected
Source identification: What documents, frameworks, books, or institutional knowledge actually encode your standards? Most leaders discover they've never written it down — it lives in their head, in casual Slack messages, in red-pen edits they've made a hundred times.
Curation: Not everything you find belongs in the context. Including too much creates noise; including too little leaves gaps. The curation process — deciding what's canonical and what's supplementary — is a judgment call that takes weeks, not hours.
Retrieval optimization: LLMs don't read documents the way humans do. Chunking matters. Embedding strategy matters. Whether you store as dense prose or structured bullet lists affects retrieval fidelity. A rubric that reads beautifully as a PDF may perform terribly when retrieved by an AI.
Version control: Standards evolve. The strategy framework you used two years ago may conflict with the one you're teaching now. Without versioning, your AI coach becomes a fossil — confidently applying outdated criteria.
02

Distribution

Getting it to the right people, consistently

Much harder than expected
Access control: Who should see which coach? A strategy Clayton built for your leadership team probably shouldn't be accessible to new hires. Role-based access, team boundaries, and guest permissions require infrastructure — not just a shared link.
Authorization and trust: Every user who accesses your AI coach is interacting with something that speaks in your voice and enforces your standards. That's a significant trust relationship. Managing who's authorized, what they can do, and what the AI is allowed to say on your behalf requires explicit governance.
Usage metrics: Is anyone actually using it? Are they engaging or just uploading once and ignoring the feedback? Without usage data, you can't distinguish a thriving coaching system from an abandoned one. Knowing who submitted what, when, and how their scores changed is how you measure whether the investment is working.
Onboarding and adoption: Even the best-designed AI coach fails if the team doesn't know how to use it. Adoption requires communication, trust-building, and ongoing reinforcement — the same change management challenges as any new tool, with the added sensitivity that some team members may feel surveilled.
03

Managing the LLM

Getting the AI to actually behave

Far harder than expected
Preventing hallucinations: LLMs confabulate. They invent rationale for scores they've already decided to give. They cite frameworks they've seen in training data instead of the framework you explicitly provided. Without strict evidence requirements — "quote the text before you score it" — a confident-sounding evaluation can be entirely fabricated.
Ensuring rubric compliance: Write a rubric with five dimensions. Tell the LLM to score exactly those five. Without enforcement, it will add a sixth ("Strategic Alignment"), rename one ("Competitive Advantage" for "Differentiation"), and explain why it had to. The model has its own opinions about what matters, and it will assert them unless you prohibit it explicitly — in every prompt, every time.
Preventing training-data override: You upload a document that says "this is NOT a vision statement." The LLM, having been trained on thousands of strategy documents that use "vision statement" as a heading, uses it anyway. The model's prior beliefs about strategy are strong. Your single document is competing against millions of tokens of training data — and losing.
Maintaining scoring consistency: The same document submitted on Monday and Wednesday can receive meaningfully different scores. Temperature, context, prompt ordering — all of these introduce variance. Achieving consistent scoring across sessions, users, and time requires elaborate prompt engineering, explicit scoring mandates, and continuous testing.
Avoiding sycophancy: LLMs want to be agreeable. Left unguarded, a well-written document gets a high score regardless of whether it meets your actual criteria. You have to explicitly mandate downward ambiguity resolution: when the evidence is ambiguous, the score must go down, not up. This is not how the model is trained to behave by default.

A Word on Usage Metrics

Most builders skip this until someone asks "is anyone actually using it?" — and the honest answer is "I have no idea."

Usage metrics aren't just a reporting convenience. They're the feedback loop that tells you whether the system is working. Who submitted? Did they act on the feedback? Did their scores improve over the next submission? How much time is the leader spending on review compared to before?

Without this data, you're running a coaching program with no outcome measurement — the equivalent of sending everyone to a training course and never asking whether anything changed. The data is also what justifies the investment: organizations that can show "we reduced leader review time by 6 hours/week" or "team strategy scores improved 22 points over one quarter" have a very different conversation with leadership than those who can only say "we built an AI thing and people seem to like it."

The Hardest Part: Managing the Model

A representative sample of the specific behaviors we had to actively engineer against.

The model invented a rubric dimension

What happened: We defined 5 scoring criteria. The LLM consistently added a 6th — "Strategic Alignment" — that wasn't in our rubric.

How we fixed it: Added explicit instruction: "Score ONLY on the exact dimensions provided. Do NOT invent new dimensions. Score exactly N items, no more, no less."

Weeks of testing
"Vision" instead of "Strategy Statement"

What happened: Our framework explicitly says "this is NOT a vision statement." The model called it a vision statement anyway — because it had seen thousands of strategy docs with that label.

How we fixed it: Added hard prohibition in both rubric dimension descriptions and deliverable-level system prompt: "NEVER label this 'Vision.' The reverse is WRONG."

Discovered in production
High score with critical rationale

What happened: The LLM would write "this metric lacks a baseline and is unverifiable" — then give the dimension a 4/5.

How we fixed it: Explicit mandate: "If your rationale names a deficiency, the score MUST reflect it. A critical rationale + high score is a contradiction — correct the score down."

Months to fully resolve
Scoring based on structure, not content

What happened: An empty table with headers scored higher than a paragraph with actual content, because the model treated formatting as evidence of quality.

How we fixed it: Added rule: "Score only specific, filled-in values. Column headers and table structure are NEVER evidence of quality. Placeholders score as absent."

Ongoing monitoring required
Applying the wrong scoring methodology

What happened: We uploaded a specific model for OKR evaluation. The model used SMART criteria instead — because it had more training data on SMART than on our custom model.

How we fixed it: Required evidence-first scoring: "Quote the text from the document that supports your score before you score it. If no such text exists, state what is missing."

Systematic prompt restructure
Subjective language in objective scoring

What happened: The model used "preferred," "best-practice," and "leading" in feedback and headlines — none of which can be verified and all of which weaken the coaching.

How we fixed it: Added explicit prohibition: "Strictly prohibit subjective superlatives. Enforce reliance on objective, observable nouns and verifiable claims."

Still catching edge cases

The Honest Reality

Fast to conceptualize

The idea of an AI coach that knows your standards is intuitive and compelling. You can sketch it in an afternoon.

Time-consuming to curate

Identifying, cleaning, structuring, and optimizing your knowledge base takes weeks of focused effort. Most of it can't be delegated.

Harder to distribute than you'd think

Access control, authorization, usage tracking, and adoption management turn a personal tool into an organizational system.

Consistently behaving is a nightmare

Getting an LLM to reliably apply your rubric — without inventing dimensions, overriding your documents, or hallucinating evidence — is a continuous, technical, ongoing battle.

Why a Specialized System Proves Its Value

You can build a functional AI coach in a weekend. The problems appear in month two, when a team member gets a score they can't explain, when the model labels your strategy section "Vision," when you realize you have no idea how often people are using it, or when a new rubric dimension you added is being selectively ignored.

A specialized system doesn't eliminate these problems — it provides the infrastructure to detect, address, and prevent them: a versioned content layer, a distribution and authorization layer, a measurement layer, and a prompt engineering and governance layer that's been battle-tested against real submissions from real teams.

Build time
A weekend
Prototype that mostly works
Production-ready
2–3 months
Content, access, governance
Consistently reliable
Ongoing
LLM behavior requires perpetual management

What This Means for You

If you're a technologist or an early adopter, building your own AI coach is a worthwhile experiment. The lessons are valuable, and for simple use cases — one user, one deliverable type, no formal measurement requirements — a home-built solution may be entirely sufficient.

If you're a leader trying to scale your standards across a team, the calculus changes. The infrastructure you'd need to build — content management, access control, usage tracking, prompt governance — is exactly the problem a specialized system is designed to solve.

The question isn't whether you can build it. You can. The question is whether building and maintaining it is the best use of the time that would otherwise go toward the thing the system is trying to support: actually developing your team.

Clayton

Ready to skip the hard parts?

Start free, or schedule a free consultation to see how Clayton's infrastructure handles the problems described in this paper — so you can focus on the coaching.

or email clayton@sagely.ltd

© 2026 Sagely Advisory LLC. All rights reserved.Lead at Scale.