Skip to content

10. Specification & Conformance

The VMx behavior contract starts in spec/, not in any single language implementation.

10.1. Source Of Truth

  • Spec index: spec/README.md
  • Current spec version: spec/VERSION
  • Compatibility matrix: compatibility-matrix.md

10.2. What Lives In The Spec

  • 24 numbered chapters from 00-overview.md through 23-async-resource-vm.md
  • ADRs describing behavior and design decisions
  • shared JSON fixtures consumed by the language flavors
  • the cross-language conformance catalog in spec/12-conformance.md

10.3. Conformance Model

The current catalog contains:

  • 407 library IDs implemented by all five catalog-complete source flavors
  • 5 THEME-00x scenario IDs exercised by the flagship example apps
  • 412 total IDs in the published catalog

The source overview is here: spec/12-conformance.md.

Catalog completeness is a source statement: every library ID has a marked test in every flavor. Executed evidence, below, shows that each marked test ran and passed. The completed Rust parity ledger supplements both with focused member and edge-behavior evidence for Rust 0.27.0.

10.3.1. Ownership Assertions And Catalog Coverage

A conformance marker proves that a test is assigned to a catalog entry; its assertions must still detect a violation of that entry's behavior. Rust's COL-055 and COL-062 check caller-owned item lifecycle after every exercised mutation, including failed keyed preflight and empty operations. A retained caller Arc always has at least one strong reference, so strong_count >= 1 cannot prove that a collection released its item handles.

The Rust tests count known membership, snapshot, lookup and returned-value handles separately from VM lifecycle hooks. They assert no collection-driven construct, destruct or dispose hook, then exactly one explicit caller disposal and one final native release per item. Retained Rust change histories contain metadata rather than item payloads. A temporary mutation that leaks the items inside both clear implementations fails both tests while the unmodified implementations pass. This strengthens existing IDs; the catalog remains 403 library IDs. See Rust ownership-test conventions.

10.3.2. Catalog, Execution, And Behavioral Evidence

Three separate statements describe a flavor's conformance:

  1. Catalog completeness. Every library ID has a marker on a test in the flavor's sources. tools/check-conformance-coverage.py checks this without running anything.
  2. Executed evidence. A runner discovered, executed, and passed at least one test for every library ID. tools/check-conformance-execution.py reads the runner's own report and checks this.
  3. Behavioral strength. A test's assertions would catch a violation of the behavior its ID names. Reviewers judge this; no tool proves it.

Neither of the first two implies the third, and none of them is member-level parity between flavors.

The execution check maps each executed case to its IDs. Python records each conformance marker as a JUnit property, and TypeScript uses a describe title that is exactly the ID. C#, Swift, and Rust match the reported class and function to the test that carries the source marker, with no name-only fallback, so a same-named test elsewhere is never credited. For each ID:

  • the ID fails if any executed case for it failed;
  • the ID passes if none failed and at least one passed;
  • the ID is skipped if every case for it was skipped;
  • the ID is missing if no executed case carries it.

Only a passing ID qualifies. Parameterized variants and duplicate reports each count, so a failing variant fails the ID and a skipped variant is ignored. A report that cannot be parsed, is truncated, disagrees with its own totals, or records a collection or build failure is rejected rather than read as success.

Each flavor workflow uploads a conformance-evidence-<flavor>-<configuration> artifact with the runner report and a JSON file that maps every library ID to its executed cases and outcomes, stamped with the commit, toolchain, and configuration:

Flavor Report Evidenced cells
C# TRX from dotnet test --logger trx, all frameworks Linux and macOS
Python JUnit XML from pytest --junitxml every OS and Python version
TypeScript JUnit XML from Vitest's junit reporter Linux and macOS, every Node line
Swift xUnit XML from swift test --xunit-output both Xcode cells
Rust libtest output from cargo test Linux and macOS

Windows cells run the same tests, but the C#, TypeScript, and Rust jobs have no reliable python3 there, so they publish no separate evidence. SwiftPM's xUnit report does not mark XCTest skips. The Swift conformance suite uses no XCTSkip, so a skipped Swift case cannot pass for executed.

10.4. How The Repo Enforces It

  • Each language flavor carries a conformance suite under its own tree.
  • tools/check-conformance-coverage.py enforces full library coverage across C#, Python, TypeScript, Swift, and Rust.
  • tools/check-conformance-execution.py requires a passing executed case for every library ID in each flavor workflow.
  • The examples workflows enforce the separate flagship scenario contract. The THEME-00x scenario IDs and two shared scenarios run in all four flagship suites: a Notes workspace lifecycle (examples/notes-showcase-scenario.json) and the five THEME scenarios in order (examples/notes-showcase-theme-scenario.json). The shared scenarios compare semantic outcomes, not just test names, and add no catalog IDs.

Coverage floors, catalog markers, and assertion strength are separate evidence. A marker assigns a test to a catalog ID. An assertion shows that the test would catch a violation. A coverage floor shows how much library code the suites run. No one of them implies the others. See Coverage Floors.

10.5. Consumer Adapter Suites

VMx TypeScript 3.21.0 introduced the optional @thekaveh/vmx/conformance entry point. It validates and executes consumer-owned operation/assertion suites without making those suites part of the normative VMx behavior catalog. The root @thekaveh/vmx entry does not export this tooling or load its Ajv validator.

The canonical schema is spec/schemas/consumer-conformance-v1.schema.json. It is a non-normative, independently versioned adapter contract governed by ADR-0102.

10.5.1. Discovery And Gap Table

VMx's fixtures, the conformance catalog, and DayDreams' consumer files have different responsibilities. Adapter v1 standardizes only their executable boundary.

Concern VMx fixtures/catalog DayDreams YAML/JSON Adapter v1 decision
Version $schema-version on four JSON fixtures YAML version; JSON fixtures unversioned Require $schema-version: "1.0.0" on executable suites
Identity Fixture-specific id/name; catalog headings hold normative IDs JSON cases hold AVM/GVM/WVM IDs Require a unique stable case id; IDs remain consumer-owned
Description Free text on some scenarios YAML and JSON descriptions Optional suite/case descriptions
Setup Fixture-specific roots such as states, transforms, or scene paths initialRoute, entries/scene fixtures, streaming setup Opaque JSON suite/case fixture values passed to the factory
Operations via, mutation tuples, producer sends, predicate/task flags YAML operations; JSON when.op with bespoke payloads Ordered invoke steps with a name and JSON argument array
State Fixture-specific expected fields Bespoke then keys assert-state with RFC 6901 JSON Pointer and structural JSON equality
Messages Message-ordering arrays and conformance prose Encoded trace strings such as PropertyChangedMessage:model@appVm Exact ordered JSON records consumed by assert-messages
Async No shared execution shape No generic async contract Await factory, every invoke, and dispose before continuing
Errors Lifecycle legal; test code owns exceptions Bespoke harness branches Validation/execution errors retain suite, case, step, and JSON paths
Teardown Each flavor test owns cleanup Each test calls dispose() Runner calls adapter dispose exactly once after factory success
YAML model fields None vm, model, types, children, state, dependencies, derived, lifecycle, conformance Discovery input only; not adapter-schema fields
Code generation None Swift generation described as a future goal Out of scope

10.5.2. Artifact Field Inventory

The discovery used the following concrete fields. DayDreams fields are pinned to consumer commit 8d314dd; they are evidence, not VMx-owned names.

Artifact Fields observed Adapter status
VMx command-truthtable.json $schema-version; cases[] with id, predicate, task, trigger_emits, can_execute, execute_invokes_task, can_execute_changed_fires Shipped adapter keeps each row as case.fixture, invokes evaluate, and asserts the three result fields
VMx derived-properties.json $schema-version, transforms; scenarios[] with name, sources_initial, transform, mutations, expected_values Could map mutations to invoke and values to state assertions; no adapter is shipped in v1
VMx lifecycle-transitions.json $schema-version, states, initial_state, terminal_states, notes; transitions[] with from, via, legal, to_intermediate, to_final Requires a factory operation per transition and explicit legal/error normalization; no adapter is shipped in v1
VMx message-ordering.json $schema-version; scenarios[] with id, description, producer-send variants, subscriber count, unsubscribe flag, and expected-observed variants Can map producer actions to invoke and expected arrays to assert-messages; no adapter is shipped in v1
VMx conformance catalog Prefix-to-chapter table; stable XXX-NNN heading IDs; Given/When/Then prose; source-chapter conformance ranges Normative metadata remains Markdown and marker-tool input; adapter case IDs do not become catalog IDs
DayDreams app.vm.yaml vm, version, description, model, types, operations, derived, lifecycle, conformance YAML remains descriptive; only its linked JSON was translated in the pilot
DayDreams gallery.vm.yaml vm, version, description, model, children, state, types, dependencies, operations, lifecycle, conformance Not parsed or standardized
DayDreams world.vm.yaml vm, version, description, model, children, registry, heightfields, types, operations, scene_bridge, lifecycle, conformance Not parsed or standardized
DayDreams AppVM JSON initialRoute; cases[] with id, description, bespoke when, and bespoke then fields AVM-001/002 were translated to suite fixture plus ordered invoke/state/message steps in the no-push pilot
DayDreams GalleryVM JSON entriesFixture; cases[] with id, description, when.op, when.args, and domain-specific then fields Factory/fixture adaptation is feasible but not included
DayDreams WorldVM JSON sceneFixture, cases, streamingCases; cases use id, optional description, assert/equals or domain-specific world, when, and then Multiple setup modes and domain serializers remain consumer-owned; not included

10.5.3. Suite Shape

Unknown structural fields fail validation. Fixtures, invocation arguments, expected state, and normalized messages accept arbitrary JSON values.

{
  "$schema-version": "1.0.0",
  "suite": "app-vm",
  "fixture": { "initialRoute": "gallery" },
  "cases": [
    {
      "id": "AVM-001",
      "steps": [
        { "kind": "invoke", "operation": "navigate", "args": ["world"] },
        {
          "kind": "assert-state",
          "path": "/model/route",
          "equals": "world"
        },
        {
          "kind": "assert-messages",
          "equals": [
            {
              "type": "PropertyChangedMessage",
              "propertyName": "model",
              "senderName": "appVm"
            }
          ]
        }
      ]
    }
  ]
}

assert-state uses RFC 6901 decoding and distinguishes a missing path from a present null. Object key order is irrelevant; array and message order are exact. assert-messages drains the adapter once for each assertion.

10.5.4. Factory And Runner

The runner owns sequencing and teardown; the consumer owns construction, operation dispatch, JSON snapshots, and message normalization.

import {
  parseConsumerConformance,
  runConsumerConformance,
  type ConsumerConformanceFactory,
} from "@thekaveh/vmx/conformance";

const factory: ConsumerConformanceFactory = ({ caseFixture }) => {
  const vm = createAppVm(caseFixture);
  return {
    invoke: async (operation, args) => invokeApp(vm, operation, args),
    snapshot: () => snapshotApp(vm),
    drainMessages: () => messageRecorder.drain(),
    dispose: () => vm.dispose(),
  };
};

const suite = parseConsumerConformance(input);
const report = await runConsumerConformance(suite, factory);

runConsumerConformanceCase integrates one case with any test framework. runConsumerConformance executes all cases sequentially and returns passed and failed results. The runner imports no Vitest or Jest API.

Factory, operation, state, message, and disposal failures include actionable instance paths. Disposal runs exactly once in finally; if execution and disposal both fail, the execution error stays primary and retains the teardown cause.

adaptCommandTruthTableFixture demonstrates adapting every unchanged VMx command fixture row. The existing CMD-007 test and all five-flavor conformance coverage remain in place.

10.5.5. Non-Goals And Follow-Up Gate

Adapter v1 does not parse YAML, define a VMx viewmodel dialect, generate tests, or generate Swift. A Swift or code-generation proposal needs a separate ADR, successful use by two independent consumers, a native Swift factory without TypeScript assumptions, and a maintained domain-type mapping that makes generation safer than manual implementation.

10.6. Test Marker Grammar

The coverage checker recognizes one intentional marker form per flavor:

Flavor Marker
C# [Trait("Conformance", "LIFE-001")]
Python @pytest.mark.conformance("LIFE-001")
TypeScript describe("LIFE-001", ...)
Swift // LIFE-001 — description or /// LIFE-001 — description, attached to a test function
Rust /// LIFE-001 — description attached through #[test] to a test function

For Rust, the ID must be the first token after /// and an em dash must follow it on the same line. Only doc-comment and attribute lines may separate the marker from #[test] and its function. Ordinary comments, file summaries, unattached markers, and markers in block-commented tests are ignored. Duplicate markers are one set-based coverage claim. In required mode, a missing Rust ID is listed under MISSING and fails the coverage command.

In required mode, a marker for an ID the catalog does not define is listed under ORPHAN and fails the coverage command for that flavor. It claims coverage of nothing, and usually means a renamed or retired ID whose real test lost its marker. Without --require, the report still lists it for information.

10.7. Practical Reading Path

  1. Read spec/README.md for chapter ownership and release history.
  2. Read the primitive pages on this site for a faster conceptual map.
  3. Use the flavor README when you need package details or host-specific examples.
  4. Use the parity matrix and conformance catalog when you need proof rather than overview.

10.8. Controlled Concurrency Interleavings

Concurrency regressions must acknowledge the boundary that creates the tested ordering. A worker-start signal only proves that a worker began; elapsed time only proves that a deadline passed. Neither proves that VMx entered a specific wait. The selected lifecycle and notification regressions instead acknowledge the exact per-instance wait, then inspect the protected state while that wait is active. The Swift command regression uses the synchronous executing notification to cancel inside the admission window, before the body task's cancellation handle is installed.

The test owns every participant and error channel. It releases gates in cleanup, joins workers, awaits task results, and reports callback failures to the coordinator. The forbidden outcomes are premature completion before the acknowledged boundary is released, missing terminal delivery, a swallowed worker or subscriber error, and cleanup that leaves work running. Diagnostic timeouts fail a test; they never count as evidence that an operation was correctly blocked.

Negative controls mutate the actual production boundary exercised by the test, run in a subprocess group with a watchdog, and restore the original source bytes in finally. A valid control either reaches the selected assertion or is terminated as a demonstrated deadlock mutant. Compilation failure, zero matched tests, an unrelated assertion, crash, or watchdog expiry cannot be reported as the selected assertion proof. CI repeats the restored Linux checks normally and with one CPU from the runner's allowed affinity set. The Swift lane records five normal selected passes, five failures from removing admission-cancellation forwarding, and a rebuilt restored pass.

Order test infrastructure as carefully as runtime participants. A package build that rewrites imported fixtures or schemas must finish before workers read them. TypeScript uses global setup and an awaited rerun hook for this boundary; generated outputs are excluded from watch triggers and explicitly invalidated after rebuilding. A filesystem error alone does not identify the process holding a lock: trace the actual writer and reader before assigning a cause.

These techniques answer different questions:

  • controlled thread interleavings force and acknowledge a particular runtime boundary;
  • stress tests sample many schedules and remain useful supplementary evidence;
  • event-loop drains allow already-queued continuations to run but do not prove that a foreign thread reached a boundary;
  • virtual schedulers deterministically advance modeled time and do not establish native-thread scheduling behavior.

The controls establish the named paths, not exhaustive freedom from races. Linux CPU affinity does not represent macOS scheduling, and an external watchdog contains a stuck process without proving in-process cleanup. The maintained scope and follow-up dispositions are recorded in the Concurrency Test Audit.