Operating model · internal tools, products and workflows

Automate the work. Keep the judgment. Grow the people who have it.

As Syzygy builds internal tools that draft, score, cluster and recommend, subject matter experts are the source of every "good" those tools know. This framework makes sure their judgment is encoded faithfully, stays in charge where it matters, and keeps getting sharper instead of quietly atrophying.

Prepared by Jake Labate, Senior SEO & AI Search Manager · Draft for discussion · September 2026

Why this exists

Every tool we ship puts expertise at risk in two ways

Risk 1 · Flattening

The tool captures the what, loses the why

An SME's rule of thumb gets hardcoded into a prompt or a scoring rule. The exceptions, the context and the reasoning stay in their head. Six months later the rule is stale, nobody knows why it exists, and the tool is confidently wrong at scale.

  • Symptom: outputs look right but fail on edge cases a senior would catch
  • Symptom: "why does the tool do this?" has no owner
Risk 2 · Atrophy

People stop exercising the skill the tool replaced

The ironies of automation: the more reliable a tool gets, the less practice people get, and the worse they are at catching it when it fails. Juniors never build the judgment seniors had, so the next generation of SMEs never forms.

  • Symptom: rubber-stamp approvals, near-zero override rate
  • Symptom: nobody can do the task by hand anymore
The principle: every tool is also a training instrument. Each time an SME accepts, edits or rejects an output, that decision should improve the tool and the team at the same time.
Where judgment lives

Three judgment tiers

Before a workflow is built, each decision inside it is assigned a tier. The tier decides how much the tool may do alone and what SME involvement is mandatory.

Automate

Tool decides, SME audits

Low stakes, reversible, well defined. The tool acts; SMEs review a sample.

  • Missing alt text, broken internal links
  • Meta length checks, schema validation
  • Audit: 5% sample weekly
Assist

Tool proposes, SME disposes

Real judgment, but patterned. Every output needs an accept, edit or reject with a reason code.

  • Title and meta rewrites
  • Keyword clustering, internal link targets
  • Content brief outlines
Reserve

SME decides, tool supports

High stakes, ambiguous, or client-relationship critical. The tool gathers evidence; a human makes the call first.

  • Migration and redirect strategy
  • Priority of a client roadmap
  • AI search positioning recommendations
Interactive

Decision router

Describe a decision a tool will touch and score it. The router assigns a tier and lists the SME controls that come with it. Example values are pre-filled.

Revenue, client trust, brand or legal exposure
Can it be rolled back cheaply once shipped?
Do two senior SMEs regularly disagree on the answer?
How often does the situation fall outside past patterns?
Is this a skill juniors must learn to become SMEs?
Recommended tier
Assist
AutomateAssistReserve

Required controls

    Lifecycle

    The judgment loop

    Judgment is not captured once. It moves through a loop that repeats for the life of every tool, and each pass trains both the tool and the people.

    1
    ElicitSMEs talk through 15 to 30 real past cases out loud, including the ones they got wrong. Capture the reasoning, not only the answer.
    2
    EncodeTurn reasoning into a written rubric and a golden set of scored examples. The rubric, not the prompt, is the source of truth.
    3
    DeployShip with a tier per decision, a named SME steward, and a version number tied to the rubric version.
    4
    ObserveLog every accept, edit and reject with a reason code. Edits are stored as diffs so the correction itself is data.
    5
    ChallengeBlind calls, disagreement reviews and red-team cases test both the tool and the SMEs against each other.
    6
    RecalibrateMonthly: update the rubric and golden set, rerun evals, and turn recurring overrides into training for juniors.
    Building blocks

    Mechanisms that preserve and train judgment

    Preserve

    Decision records

    Each non-trivial SME call gets a short record: situation, options considered, choice, reasoning, what would change the call. They become the elicitation library for future tools and onboarding reading for juniors.

    Preserve

    Rubrics and golden sets

    SMEs own a versioned rubric and 50 to 200 scored examples per tool. No tool or prompt change ships unless it holds its score on the golden set. The SME, not the builder, signs off.

    Preserve

    Reason-coded overrides

    Accept, edit or reject is never a bare click. A short reason code list (wrong intent, brand voice, client context, factual, strategic) makes overrides countable and teachable.

    Train

    Blind calls

    For Assist and Reserve decisions, a rotating share (start at 10%) is shown to the SME before the tool output. This prevents anchoring, keeps the skill in use, and gives a clean measure of human versus tool accuracy.

    Train

    Disagreement reviews

    When two SMEs, or an SME and the tool, disagree, the case goes to a 30 minute monthly review. The resolution updates the rubric. Disagreement is signal, not noise.

    Train

    Shadow and explain

    Juniors review tool outputs next to a senior's decisions and must predict the senior's call and explain it. Accuracy on these predictions is their promotion evidence toward Practitioner.

    Interactive

    Judgment log demo

    This is what an Assist-tier workflow looks like from the SME's seat. Sample outputs from a title-tag rewriting tool. Record a decision with a reason; the health stats update live. Try a blind call first to see how it changes the flow.

    Case 1 of 6
    Page:
    Current title:
    Tool proposal
    0Decisions logged
    0%Override rate (edit + reject)
    0Blind calls made
    None yetTop override reason
    CaseCallReason
    No decisions yet. Log one to start the record.

    Healthy Assist tiers usually sit between 10% and 35% overrides. Near zero suggests rubber-stamping; above 40% means the rubric or the tool needs recalibration.

    Growing SMEs

    The SME ladder

    Expertise is made visible and earned through evidence the tools already produce, so moving up is about demonstrated judgment rather than tenure.

    Level 1

    Apprentice

    Learns the domain through the tools, not around them.

    • Shadow-and-explain on 40+ cases
    • Reads decision records
    • No solo sign-off
    Level 2

    Practitioner

    Owns Assist-tier calls on live work.

    • 80%+ prediction match with seniors
    • Reason-coded overrides
    • Regular blind calls
    Level 3

    Expert

    Makes Reserve-tier calls and mentors.

    • Blind-call accuracy beats the tool
    • Writes decision records
    • Resolves disagreements
    Level 4

    Steward

    Owns the judgment inside a tool.

    • Owns the rubric and golden set
    • Signs off tool releases
    • Runs recalibration
    Who does what, and when

    Roles and operating cadence

    RoleOwnsTime cost
    SME StewardRubric, golden set, release sign-off, recalibration for one tool2 to 3 hrs / week per tool
    Tool OwnerBuild, logging, evals, versioning; never changes the rubric aloneAs scoped
    PractitionersReason-coded decisions and blind calls on live workBuilt into delivery
    ApprenticesShadow-and-explain sets, decision record reading1 hr / week
    Organic LeadTier assignments, ladder promotions, portfolio of tools1 hr / month
    CadenceRitualOutput
    WeeklyOverride triage (15 min): steward scans the log for new patternsFlagged cases for review
    MonthlyDisagreement review (30 min) and golden set rerunRubric vN+1, eval scores
    QuarterlyTier review: does each decision still belong in its tier?Tier changes, retirements
    Twice a yearLadder calibration using blind-call and prediction dataPromotions, training plans
    Is it working?

    Health metrics

    MetricWhat it tells usHealthy signal
    Override rateWhether SMEs are really reviewing10 to 35% on Assist
    Blind-call gapHuman accuracy versus tool on the same casesSMEs at or above tool
    Golden set scoreWhether the tool still matches encoded judgmentNo drop between releases
    Inter-SME agreementWhether the rubric is clear enough to be sharedRising quarter on quarter
    Rubric freshnessWhether judgment is being maintainedUpdated in last 60 days
    Time to PractitionerWhether tools are training people, not replacing themFalling over time
    Stewardship coverageWhether every tool has a named owner of its judgment100% of live tools
    Getting started

    90-day rollout

    Days 1 to 30

    Pilot one tool

    • Pick one Assist-tier workflow (title and meta rewrites)
    • Name a steward, run an elicitation session
    • Write rubric v1 and a 50-case golden set
    • Turn on reason-coded logging
    Days 31 to 60

    Close the loop

    • Start blind calls at 10%
    • First disagreement review and rubric v2
    • Enroll two apprentices in shadow-and-explain
    • Baseline the health metrics
    Days 61 to 90

    Scale the pattern

    • Tier every decision in the next two tools before build
    • Publish the ladder criteria
    • Review results with the Organic Lead
    • Decide what becomes standard practice