Lead at Scale · White Paper · May 2026The concept takes an afternoon. The content curation takes weeks. The distribution requires infrastructure you didn't plan for. And getting the AI to consistently behave the way you want? That's a problem that never fully goes away.
Anyone can build a passable AI coach in a weekend. Describe your standards in a system prompt, upload a few documents, share the link. Done — in the way that a house is "done" when the walls are framed and the roof is on.
What comes next is where the work actually lives: curating the knowledge that makes the coach credible, distributing it to a team in a way that's authorized and measurable, and — the part most builders underestimate — getting the underlying language model to follow your rules instead of its own.
This paper describes what we've learned building Clayton: the three layers of genuine difficulty that separate a promising prototype from a coaching system that works reliably at scale.
Each layer compounds the one before it. You can skip none of them.
Finding and curating the right knowledge
Getting it to the right people, consistently
Getting the AI to actually behave
Most builders skip this until someone asks "is anyone actually using it?" — and the honest answer is "I have no idea."
Usage metrics aren't just a reporting convenience. They're the feedback loop that tells you whether the system is working. Who submitted? Did they act on the feedback? Did their scores improve over the next submission? How much time is the leader spending on review compared to before?
Without this data, you're running a coaching program with no outcome measurement — the equivalent of sending everyone to a training course and never asking whether anything changed. The data is also what justifies the investment: organizations that can show "we reduced leader review time by 6 hours/week" or "team strategy scores improved 22 points over one quarter" have a very different conversation with leadership than those who can only say "we built an AI thing and people seem to like it."
A representative sample of the specific behaviors we had to actively engineer against.
What happened: We defined 5 scoring criteria. The LLM consistently added a 6th — "Strategic Alignment" — that wasn't in our rubric.
How we fixed it: Added explicit instruction: "Score ONLY on the exact dimensions provided. Do NOT invent new dimensions. Score exactly N items, no more, no less."
What happened: Our framework explicitly says "this is NOT a vision statement." The model called it a vision statement anyway — because it had seen thousands of strategy docs with that label.
How we fixed it: Added hard prohibition in both rubric dimension descriptions and deliverable-level system prompt: "NEVER label this 'Vision.' The reverse is WRONG."
What happened: The LLM would write "this metric lacks a baseline and is unverifiable" — then give the dimension a 4/5.
How we fixed it: Explicit mandate: "If your rationale names a deficiency, the score MUST reflect it. A critical rationale + high score is a contradiction — correct the score down."
What happened: An empty table with headers scored higher than a paragraph with actual content, because the model treated formatting as evidence of quality.
How we fixed it: Added rule: "Score only specific, filled-in values. Column headers and table structure are NEVER evidence of quality. Placeholders score as absent."
What happened: We uploaded a specific model for OKR evaluation. The model used SMART criteria instead — because it had more training data on SMART than on our custom model.
How we fixed it: Required evidence-first scoring: "Quote the text from the document that supports your score before you score it. If no such text exists, state what is missing."
What happened: The model used "preferred," "best-practice," and "leading" in feedback and headlines — none of which can be verified and all of which weaken the coaching.
How we fixed it: Added explicit prohibition: "Strictly prohibit subjective superlatives. Enforce reliance on objective, observable nouns and verifiable claims."
The idea of an AI coach that knows your standards is intuitive and compelling. You can sketch it in an afternoon.
Identifying, cleaning, structuring, and optimizing your knowledge base takes weeks of focused effort. Most of it can't be delegated.
Access control, authorization, usage tracking, and adoption management turn a personal tool into an organizational system.
Getting an LLM to reliably apply your rubric — without inventing dimensions, overriding your documents, or hallucinating evidence — is a continuous, technical, ongoing battle.
You can build a functional AI coach in a weekend. The problems appear in month two, when a team member gets a score they can't explain, when the model labels your strategy section "Vision," when you realize you have no idea how often people are using it, or when a new rubric dimension you added is being selectively ignored.
A specialized system doesn't eliminate these problems — it provides the infrastructure to detect, address, and prevent them: a versioned content layer, a distribution and authorization layer, a measurement layer, and a prompt engineering and governance layer that's been battle-tested against real submissions from real teams.
If you're a technologist or an early adopter, building your own AI coach is a worthwhile experiment. The lessons are valuable, and for simple use cases — one user, one deliverable type, no formal measurement requirements — a home-built solution may be entirely sufficient.
If you're a leader trying to scale your standards across a team, the calculus changes. The infrastructure you'd need to build — content management, access control, usage tracking, prompt governance — is exactly the problem a specialized system is designed to solve.
The question isn't whether you can build it. You can. The question is whether building and maintaining it is the best use of the time that would otherwise go toward the thing the system is trying to support: actually developing your team.

Ready to skip the hard parts?
Start free, or schedule a free consultation to see how Clayton's infrastructure handles the problems described in this paper — so you can focus on the coaching.
or email clayton@sagely.ltd