Jump to content

Developer Note: Abstract Instrument Checklist — What Cals to Make (Seed)

From The Sovereign Games (MoA Lab)

Welcome to the MoA–TSG Lab. The wiki is the bench. The work is Metrology of the Abstract. Adopt the tools or leave them on the rack — either way, the need doesn't wait.

  • Lab Note: A redlink is not a failure. It identifies Calibration Debt—work waiting to be measured, mapped, and calibrated.

CYCLE Calibration position

PositionDeveloper note · Cal backlog / instrument inventory
CycleSeed
StatusSelf-Assessed · Find a need and cal it

This page is a conceptual instrument under Permanent Beta. It declares a real calibration position, not a finished product waiting to ship. Checking continues; an edit is only required when evidence demands it. Stage: Seed to Fruit.

Feedback welcome — especially clarity, failure modes, and calibration gaps. Use discussion or Contribute.


Developer Note: Abstract Instrument Checklist — What Cals to Make

One line: Current inventory of abstract instrument types and registries worth building — need-based, non-exhaustive, extended when use exposes missing measurements.



Sovereign-Games-OG-Image.jpg

Meta

Developer Note: Abstract Instrument Checklist — What Cals to Make (Seed)

Type Development Note
Functional Layer Inventory / Planning
Application Layer MoA procedures · Registries · Cross-domain assist
Category Development Notes
Version 0.2
Maturity Seed
Last Calibration 2026-07-30
Status Permanent Beta
Description Operating inventory of abstract instrument types and collections to calibrate or capture. Round-robin process (not a roadmap): independent vs integrated paths, need-based priority, self-extension by unmet measurement need. Non-exhaustive.

Core Principles

  • Reality gets final vote
  • See the Game. Refuse the Game. Build Better.
  • Permanent Beta

Navigation

Related



Disclaimer

  • Non-exhaustive. Silence ≠ denial.
  • Not a roadmap. No claim to know every future instrument. Priority is need-based.
  • Calibrating an instrument type ≠ seizing any domain’s Domain SI.
  • Round-robin: challenge splits, merge/split rows, mark blockers, name gaps this list missed.
  • Standing of this note remains Self-Assessed until a published path says otherwise.

0. Posture: inventory process, not roadmap

Roadmaps assume the future is already known. This page assumes we almost certainly forgot something.

Operating loop (self-extension):

Need
  → Identify abstract instrument
  → Write procedure / card template
  → Deploy (shop → wall when fit)
  → Use under load
  → Discover missing instruments or failure modes
  → Expand inventory
  → Repeat

That is how physical metrology grows: new instruments appear because reality demanded them, not because a committee listed everything in advance.

MoA organizes around metrological discipline — recurring measurement needs, failure cost, and re-run pressure — not around personalities or ideologies.

Metrology extends itself by finding unmet measurement needs. This checklist is one surface for that discovery. Round-robin is how the map is pressured.

Instrument classes below (metric card, bundle, summary FIT, risk language, …) are treated as calibration targets, not brainstorm topics.


How to read the tables

Column Meaning
Independent cal Own procedure path and certificate object (card / RAM / CR-style) without waiting on a parent bundle
Integrated cal Run as module inside ADM / Deep Dive / metric stack / bundle / project gate — shared lock or shared job
Both Can stand alone or plug into a larger packet; prefer independent first when the string is load-bearing everywhere

Optional future row status (when a live registry exists): Exists · Needs revision · Missing · Candidate — same question labs ask: *what instruments are currently unsupported?*


A. Abstract instrument types (what to calibrate)

A1. Decision & performance language

ID Instrument type Independent cal? Integrated cal? Notes / first product shape
A1.1 Analysis metrics / KPIs Yes — metric card + joint N∧S Inside closed-loop, stage-gate, dashboard packs Priority spine; see Abstract Analysis Metrics Seeds
A1.2 Metric bundles + combine rules Yes — bundle card Assembly of member metric cards After ≥2 member cards exist
A1.3 Success / done criteria Both — reusable done-card template Project / MVP / release Deep Dive module Job = exit decision; often gate-integrated in practice
A1.4 Requirement senses (must / should / optional) Yes — sense-lock per keyword family Spec / contract ADM under locked text Child instruments per register
A1.5 Standing language (Self-Assessed / Confirmed / …) Yes — sense-lock card for standing terms Every cert, card, and map claim Prevents standing theater; used across MoA surfaces

A2. Risk, evidence, audit language

ID Instrument type Independent cal? Integrated cal? Notes / first product shape
A2.1 Risk language (low / acceptable / critical) Yes — sense + threshold card Risk register / safety case packet High Goodhart risk if targeted
A2.2 Evidence grades (support vs illustration vs anecdote) Yes — grade scale card Report / paper / claim ADM Pairs with summary FIT
A2.3 Audit findings language (severity labels) Yes Audit metric + HITL triggers Bind severity → action on card
A2.4 Adversarial / audit metrics (role) Yes as metric cards Bundle role = audit/constraint See assistance framework § audit
A2.5 Uncertainty-budget language (what “low residual” claims mean) Yes — thin sense-lock when used in public claims Metric cards, RAM, reports Optional until pilots abuse vague residual talk

A3. Claims, summaries, models

ID Instrument type Independent cal? Integrated cal? Notes / first product shape
A3.1 Summary / report claims Yes — summary FIT checklist as instrument LLM meta-analysis gate; exec abstract module Pin to metric IDs
A3.2 Model task labels (helpful, safe, on-policy) Yes — label cards AI eval / RLHF-style stacks Vendor COI mandatory
A3.3 Benchmark claims Yes — “score means…” card Leaderboard / model release packet Classic Goodhart surface
A3.4 Classifier / score-model dependency Fields on metric card (not always a separate religion) Integrated into any model-backed KPI Version pin + COI on A1.1 card

A4. Policy, interface, permission, change

ID Instrument type Independent cal? Integrated cal? Notes / first product shape
A4.1 Policy operative terms (eligibility, harm, materiality) Yes — ADM/RAM per term under lock Doctrine / statute assembly Deep Dive Hostile multi-sense likely
A4.2 Interface / compatibility contracts (semantic) Yes — contract sense card API / standards clause packet Neighbor ≠ substitute
A4.3 Consent / permission scopes Yes Data/governance bundles Scope bleed failure mode
A4.4 Drift / change notices (material change) Yes — what counts as regime shift Re-lock trigger library; metric promote path Travels across metric, policy, and RAM

A5. Lexical / definition stack (already in motion — keep on list)

ID Instrument type Independent cal? Integrated cal? Notes
A5.1 Dictionary / term mechanisms (TDP, CR) Yes Feeds ADM / RAM Active
A5.2 Multi-sense headwords Yes per child sense Interconnect map across children Split hostile vs benign polysemy only if pilots demand
A5.3 RAM L0–L5 / Deep Dive bodies Yes Page + report packet Active IA

B. Collections / registries (what to capture)

Capture layer — not a substitute for calibrating the instruments above.

ID Collection Independent build? Integrated with Capture contents
B1 Metric card registry Yes A1.1, promote path ID, version, standing, links
B2 Bundle registry Yes A1.2 Members, combine rule, version
B3 Failure-mode library Yes All metric/RAM work M/B/H/T/V codes + examples
B4 Goodhart watchlist Yes (keep separate from B3) M7; target-pressure metrics Metric ID, incentive, review date
B5 Regime-shift / re-lock log Yes A4.4; control stack Trigger, decision, version from→to
B6 Competing maps/cards index Yes Federation / self-cal Same Domain SI cite → N cards
B7 Summary FIT logs Yes A3.1 Claim text hash, pin result, date
B8 COI declarations registry Yes Vendor/author maps Who, lean, instrument IDs
B9 Claim-type index Yes MoA guidelines A–F Cert URL, claim type, method pin
B10 Term / definition RAM index Yes A5.* Map: / DeepDive: hubs

C. Suggested cal paths (integration vs independent)

C1. Prefer independent first

  • A1.1 Metric cards (template + 1–2 live pilots)
  • A5.* Term/RAM (already moving)
  • B1 Metric registry + B3 Failure-mode library (thin)
  • A3.1 Summary FIT checklist (short standalone instrument)
  • A1.5 Standing-language sense-lock (thin; used everywhere)

Why: Reusable across domains; unblocks bundles, gates, and honest cert claims.

C2. Prefer integrated (after independents exist)

  • A1.2 Bundles (needs member cards)
  • Joint N∧S + HITL + re-lock as modules on the metric procedure, not separate religions
  • A2.4 Audit role inside bundle cal
  • A3.4 Classifier dependency fields inside metric card cal

Why: Avoid procedure sprawl; same promote path.

C3. Parallel independent when hostile or multi-domain

  • A4.1 Policy operative terms (own ADM runs)
  • A3.2 / A3.3 Model labels & benchmark claims (AI eval domain)
  • A4.3 Consent scopes (governance domain)

Why: Different Domain SI; different failure politics; don’t wait on manufacturing KPI pilots.

C4. Round-robin prompts (check list)

  • [ ] Missing instrument type?
  • [ ] Row that should split?
  • [ ] Row that should merge?
  • [ ] Wrong independent/integrated flag?
  • [ ] Top 3 need-based pilots for the next work stretch?
  • [ ] Any item that is physical Domain SI creep (cut it)?
  • [ ] What instruments does MoA itself still lack for its own claims?

D. Direct “what cals to make” checklist (operators)

Instrument cals

  • [ ] Metric card procedure + template (A1.1)
  • [ ] Bundle card procedure + template (A1.2)
  • [ ] Done/success criteria card (A1.3)
  • [ ] Must/should/optional sense-lock (A1.4)
  • [ ] Standing-language sense-lock (A1.5)
  • [ ] Risk language card (A2.1)
  • [ ] Evidence grade scale (A2.2)
  • [ ] Audit severity language (A2.3)
  • [ ] Uncertainty-budget language card if residual claims go public (A2.5)
  • [ ] Summary FIT instrument (A3.1)
  • [ ] Model task label cards (A3.2)
  • [ ] Benchmark claim card (A3.3)
  • [ ] Policy term ADM/RAM pilots (A4.1)
  • [ ] Semantic compatibility / interface card (A4.2)
  • [ ] Consent scope card (A4.3)
  • [ ] Material change / regime-shift definition (A4.4)
  • [ ] Continue term/RAM/Deep Dive pipeline (A5.*)

Collection cals / stand-ups

  • [ ] B1 Metric registry
  • [ ] B2 Bundle registry
  • [ ] B3 Failure-mode library
  • [ ] B4 Goodhart watchlist
  • [ ] B5 Re-lock log
  • [ ] B6 Competing maps index
  • [ ] B7 Summary FIT log
  • [ ] B8 COI registry
  • [ ] B9 Claim-type index
  • [ ] B10 Term/RAM hub index

E. Priority seed (editable under friction)

Suggestion only — replace when need or round-robin says otherwise:

  1. A1.1 Metric card cal (procedure + 1 pilot with real metric ID)
  2. B1 + B3 Registry + failure-mode library (thin)
  3. A3.1 Summary FIT
  4. A1.5 Standing-language sense-lock (thin)
  5. A1.2 Bundle (once two metrics exist)
  6. A5 continue + one A4.1 policy term if a live need appears

Growth is driven by identified measurement gaps, not by inventing initiatives to fill a roadmap.


F. One-line doctrine

List the abstract instruments, calibrate them independently when they travel alone, integrate them when they only make sense inside a gate or bundle, capture versions in registries, and expand the inventory when use exposes a missing measurement — need-based, non-exhaustive, Domain SI left alone.


See also: Abstract Analysis Metrics — Seed Assistance Framework (General) · Abstract Analysis Metrics — Seed Assistance Framework (Metrology / NIST-oriented posture) · Developer Note: Domains, Not Initiatives · MoA Guidelines — Maps, Certificates, and Self-Calibration · Domain SI — We Calibrate, We Do Not Set · TDP / ADM / RAM notes · Permanent Beta