Automate the work. Keep the judgment. Grow the people who have it.
As Syzygy builds internal tools that draft, score, cluster and recommend, subject matter experts are the source of every "good" those tools know. This framework makes sure their judgment is encoded faithfully, stays in charge where it matters, and keeps getting sharper instead of quietly atrophying.
Every tool we ship puts expertise at risk in two ways
The tool captures the what, loses the why
An SME's rule of thumb gets hardcoded into a prompt or a scoring rule. The exceptions, the context and the reasoning stay in their head. Six months later the rule is stale, nobody knows why it exists, and the tool is confidently wrong at scale.
- Symptom: outputs look right but fail on edge cases a senior would catch
- Symptom: "why does the tool do this?" has no owner
People stop exercising the skill the tool replaced
The ironies of automation: the more reliable a tool gets, the less practice people get, and the worse they are at catching it when it fails. Juniors never build the judgment seniors had, so the next generation of SMEs never forms.
- Symptom: rubber-stamp approvals, near-zero override rate
- Symptom: nobody can do the task by hand anymore
Three judgment tiers
Before a workflow is built, each decision inside it is assigned a tier. The tier decides how much the tool may do alone and what SME involvement is mandatory.
Tool decides, SME audits
Low stakes, reversible, well defined. The tool acts; SMEs review a sample.
- Missing alt text, broken internal links
- Meta length checks, schema validation
- Audit: 5% sample weekly
Tool proposes, SME disposes
Real judgment, but patterned. Every output needs an accept, edit or reject with a reason code.
- Title and meta rewrites
- Keyword clustering, internal link targets
- Content brief outlines
SME decides, tool supports
High stakes, ambiguous, or client-relationship critical. The tool gathers evidence; a human makes the call first.
- Migration and redirect strategy
- Priority of a client roadmap
- AI search positioning recommendations
Decision router
Describe a decision a tool will touch and score it. The router assigns a tier and lists the SME controls that come with it. Example values are pre-filled.
Required controls
The judgment loop
Judgment is not captured once. It moves through a loop that repeats for the life of every tool, and each pass trains both the tool and the people.
Mechanisms that preserve and train judgment
Decision records
Each non-trivial SME call gets a short record: situation, options considered, choice, reasoning, what would change the call. They become the elicitation library for future tools and onboarding reading for juniors.
Rubrics and golden sets
SMEs own a versioned rubric and 50 to 200 scored examples per tool. No tool or prompt change ships unless it holds its score on the golden set. The SME, not the builder, signs off.
Reason-coded overrides
Accept, edit or reject is never a bare click. A short reason code list (wrong intent, brand voice, client context, factual, strategic) makes overrides countable and teachable.
Blind calls
For Assist and Reserve decisions, a rotating share (start at 10%) is shown to the SME before the tool output. This prevents anchoring, keeps the skill in use, and gives a clean measure of human versus tool accuracy.
Disagreement reviews
When two SMEs, or an SME and the tool, disagree, the case goes to a 30 minute monthly review. The resolution updates the rubric. Disagreement is signal, not noise.
Shadow and explain
Juniors review tool outputs next to a senior's decisions and must predict the senior's call and explain it. Accuracy on these predictions is their promotion evidence toward Practitioner.
Judgment log demo
This is what an Assist-tier workflow looks like from the SME's seat. Sample outputs from a title-tag rewriting tool. Record a decision with a reason; the health stats update live. Try a blind call first to see how it changes the flow.
| Case | Call | Reason |
|---|---|---|
| No decisions yet. Log one to start the record. | ||
Healthy Assist tiers usually sit between 10% and 35% overrides. Near zero suggests rubber-stamping; above 40% means the rubric or the tool needs recalibration.
The SME ladder
Expertise is made visible and earned through evidence the tools already produce, so moving up is about demonstrated judgment rather than tenure.
Apprentice
Learns the domain through the tools, not around them.
- Shadow-and-explain on 40+ cases
- Reads decision records
- No solo sign-off
Practitioner
Owns Assist-tier calls on live work.
- 80%+ prediction match with seniors
- Reason-coded overrides
- Regular blind calls
Expert
Makes Reserve-tier calls and mentors.
- Blind-call accuracy beats the tool
- Writes decision records
- Resolves disagreements
Steward
Owns the judgment inside a tool.
- Owns the rubric and golden set
- Signs off tool releases
- Runs recalibration
Roles and operating cadence
| Role | Owns | Time cost |
|---|---|---|
| SME Steward | Rubric, golden set, release sign-off, recalibration for one tool | 2 to 3 hrs / week per tool |
| Tool Owner | Build, logging, evals, versioning; never changes the rubric alone | As scoped |
| Practitioners | Reason-coded decisions and blind calls on live work | Built into delivery |
| Apprentices | Shadow-and-explain sets, decision record reading | 1 hr / week |
| Organic Lead | Tier assignments, ladder promotions, portfolio of tools | 1 hr / month |
| Cadence | Ritual | Output |
|---|---|---|
| Weekly | Override triage (15 min): steward scans the log for new patterns | Flagged cases for review |
| Monthly | Disagreement review (30 min) and golden set rerun | Rubric vN+1, eval scores |
| Quarterly | Tier review: does each decision still belong in its tier? | Tier changes, retirements |
| Twice a year | Ladder calibration using blind-call and prediction data | Promotions, training plans |
Health metrics
| Metric | What it tells us | Healthy signal |
|---|---|---|
| Override rate | Whether SMEs are really reviewing | 10 to 35% on Assist |
| Blind-call gap | Human accuracy versus tool on the same cases | SMEs at or above tool |
| Golden set score | Whether the tool still matches encoded judgment | No drop between releases |
| Inter-SME agreement | Whether the rubric is clear enough to be shared | Rising quarter on quarter |
| Rubric freshness | Whether judgment is being maintained | Updated in last 60 days |
| Time to Practitioner | Whether tools are training people, not replacing them | Falling over time |
| Stewardship coverage | Whether every tool has a named owner of its judgment | 100% of live tools |
90-day rollout
Pilot one tool
- Pick one Assist-tier workflow (title and meta rewrites)
- Name a steward, run an elicitation session
- Write rubric v1 and a 50-case golden set
- Turn on reason-coded logging
Close the loop
- Start blind calls at 10%
- First disagreement review and rubric v2
- Enroll two apprentices in shadow-and-explain
- Baseline the health metrics
Scale the pattern
- Tier every decision in the next two tools before build
- Publish the ladder criteria
- Review results with the Organic Lead
- Decide what becomes standard practice