10. Specification & Conformance¶
The VMx behavior contract starts in spec/, not in any single language
implementation.
10.1. Source Of Truth¶
- Spec index: spec/README.md
- Current spec version: spec/VERSION
- Compatibility matrix: compatibility-matrix.md
10.2. What Lives In The Spec¶
- 24 numbered chapters from
00-overview.mdthrough23-async-resource-vm.md - ADRs describing behavior and design decisions
- shared JSON fixtures consumed by the language flavors
- the cross-language conformance catalog in
spec/12-conformance.md
10.3. Conformance Model¶
The current catalog contains:
- 407 library IDs implemented by all five catalog-complete source flavors
- 5
THEME-00xscenario IDs exercised by the flagship example apps - 412 total IDs in the published catalog
The source overview is here: spec/12-conformance.md.
Catalog completeness is a source statement: every library ID has a marked test in every flavor. Executed evidence, below, shows that each marked test ran and passed. The completed Rust parity ledger supplements both with focused member and edge-behavior evidence for Rust 0.27.0.
10.3.1. Ownership Assertions And Catalog Coverage¶
A conformance marker proves that a test is assigned to a catalog entry; its
assertions must still detect a violation of that entry's behavior. Rust's
COL-055 and COL-062 check caller-owned item lifecycle after every exercised
mutation, including failed keyed preflight and empty operations. A retained
caller Arc always has at least one strong reference, so strong_count >= 1
cannot prove that a collection released its item handles.
The Rust tests count known membership, snapshot, lookup and returned-value
handles separately from VM lifecycle hooks. They assert no collection-driven
construct, destruct or dispose hook, then exactly one explicit caller disposal
and one final native release per item. Retained Rust change histories contain
metadata rather than item payloads. A temporary mutation that leaks the items
inside both clear implementations fails both tests while the unmodified
implementations pass. This strengthens existing IDs; the catalog remains 403
library IDs. See Rust ownership-test conventions.
10.3.2. Catalog, Execution, And Behavioral Evidence¶
Three separate statements describe a flavor's conformance:
- Catalog completeness. Every library ID has a marker on a test in the
flavor's sources.
tools/check-conformance-coverage.pychecks this without running anything. - Executed evidence. A runner discovered, executed, and passed at least one
test for every library ID.
tools/check-conformance-execution.pyreads the runner's own report and checks this. - Behavioral strength. A test's assertions would catch a violation of the behavior its ID names. Reviewers judge this; no tool proves it.
Neither of the first two implies the third, and none of them is member-level parity between flavors.
The execution check maps each executed case to its IDs. Python records each
conformance marker as a JUnit property, and TypeScript uses a describe
title that is exactly the ID. C#, Swift, and Rust match the reported class and
function to the test that carries the source marker, with no name-only
fallback, so a same-named test elsewhere is never credited. For each ID:
- the ID fails if any executed case for it failed;
- the ID passes if none failed and at least one passed;
- the ID is skipped if every case for it was skipped;
- the ID is missing if no executed case carries it.
Only a passing ID qualifies. Parameterized variants and duplicate reports each count, so a failing variant fails the ID and a skipped variant is ignored. A report that cannot be parsed, is truncated, disagrees with its own totals, or records a collection or build failure is rejected rather than read as success.
Each flavor workflow uploads a conformance-evidence-<flavor>-<configuration>
artifact with the runner report and a JSON file that maps every library ID to
its executed cases and outcomes, stamped with the commit, toolchain, and
configuration:
| Flavor | Report | Evidenced cells |
|---|---|---|
| C# | TRX from dotnet test --logger trx, all frameworks |
Linux and macOS |
| Python | JUnit XML from pytest --junitxml |
every OS and Python version |
| TypeScript | JUnit XML from Vitest's junit reporter |
Linux and macOS, every Node line |
| Swift | xUnit XML from swift test --xunit-output |
both Xcode cells |
| Rust | libtest output from cargo test |
Linux and macOS |
Windows cells run the same tests, but the C#, TypeScript, and Rust jobs have no
reliable python3 there, so they publish no separate evidence. SwiftPM's xUnit
report does not mark XCTest skips. The Swift conformance suite uses no
XCTSkip, so a skipped Swift case cannot pass for executed.
10.4. How The Repo Enforces It¶
- Each language flavor carries a conformance suite under its own tree.
tools/check-conformance-coverage.pyenforces full library coverage across C#, Python, TypeScript, Swift, and Rust.tools/check-conformance-execution.pyrequires a passing executed case for every library ID in each flavor workflow.- The examples workflows enforce the separate flagship scenario contract. The
THEME-00xscenario IDs and two shared scenarios run in all four flagship suites: a Notes workspace lifecycle (examples/notes-showcase-scenario.json) and the five THEME scenarios in order (examples/notes-showcase-theme-scenario.json). The shared scenarios compare semantic outcomes, not just test names, and add no catalog IDs.
Coverage floors, catalog markers, and assertion strength are separate evidence. A marker assigns a test to a catalog ID. An assertion shows that the test would catch a violation. A coverage floor shows how much library code the suites run. No one of them implies the others. See Coverage Floors.
10.5. Consumer Adapter Suites¶
VMx TypeScript 3.21.0 introduced the optional
@thekaveh/vmx/conformance entry point. It validates and executes
consumer-owned operation/assertion suites without making those suites part of
the normative VMx behavior catalog. The root @thekaveh/vmx entry does not
export this tooling or load its Ajv validator.
The canonical schema is
spec/schemas/consumer-conformance-v1.schema.json. It is a non-normative,
independently versioned adapter contract governed by ADR-0102.
10.5.1. Discovery And Gap Table¶
VMx's fixtures, the conformance catalog, and DayDreams' consumer files have different responsibilities. Adapter v1 standardizes only their executable boundary.
| Concern | VMx fixtures/catalog | DayDreams YAML/JSON | Adapter v1 decision |
|---|---|---|---|
| Version | $schema-version on four JSON fixtures |
YAML version; JSON fixtures unversioned |
Require $schema-version: "1.0.0" on executable suites |
| Identity | Fixture-specific id/name; catalog headings hold normative IDs |
JSON cases hold AVM/GVM/WVM IDs | Require a unique stable case id; IDs remain consumer-owned |
| Description | Free text on some scenarios | YAML and JSON descriptions | Optional suite/case descriptions |
| Setup | Fixture-specific roots such as states, transforms, or scene paths | initialRoute, entries/scene fixtures, streaming setup |
Opaque JSON suite/case fixture values passed to the factory |
| Operations | via, mutation tuples, producer sends, predicate/task flags |
YAML operations; JSON when.op with bespoke payloads |
Ordered invoke steps with a name and JSON argument array |
| State | Fixture-specific expected fields | Bespoke then keys |
assert-state with RFC 6901 JSON Pointer and structural JSON equality |
| Messages | Message-ordering arrays and conformance prose | Encoded trace strings such as PropertyChangedMessage:model@appVm |
Exact ordered JSON records consumed by assert-messages |
| Async | No shared execution shape | No generic async contract | Await factory, every invoke, and dispose before continuing |
| Errors | Lifecycle legal; test code owns exceptions |
Bespoke harness branches | Validation/execution errors retain suite, case, step, and JSON paths |
| Teardown | Each flavor test owns cleanup | Each test calls dispose() |
Runner calls adapter dispose exactly once after factory success |
| YAML model fields | None | vm, model, types, children, state, dependencies, derived, lifecycle, conformance |
Discovery input only; not adapter-schema fields |
| Code generation | None | Swift generation described as a future goal | Out of scope |
10.5.2. Artifact Field Inventory¶
The discovery used the following concrete fields. DayDreams fields are pinned
to consumer commit 8d314dd; they are evidence, not VMx-owned names.
| Artifact | Fields observed | Adapter status |
|---|---|---|
VMx command-truthtable.json |
$schema-version; cases[] with id, predicate, task, trigger_emits, can_execute, execute_invokes_task, can_execute_changed_fires |
Shipped adapter keeps each row as case.fixture, invokes evaluate, and asserts the three result fields |
VMx derived-properties.json |
$schema-version, transforms; scenarios[] with name, sources_initial, transform, mutations, expected_values |
Could map mutations to invoke and values to state assertions; no adapter is shipped in v1 |
VMx lifecycle-transitions.json |
$schema-version, states, initial_state, terminal_states, notes; transitions[] with from, via, legal, to_intermediate, to_final |
Requires a factory operation per transition and explicit legal/error normalization; no adapter is shipped in v1 |
VMx message-ordering.json |
$schema-version; scenarios[] with id, description, producer-send variants, subscriber count, unsubscribe flag, and expected-observed variants |
Can map producer actions to invoke and expected arrays to assert-messages; no adapter is shipped in v1 |
| VMx conformance catalog | Prefix-to-chapter table; stable XXX-NNN heading IDs; Given/When/Then prose; source-chapter conformance ranges |
Normative metadata remains Markdown and marker-tool input; adapter case IDs do not become catalog IDs |
DayDreams app.vm.yaml |
vm, version, description, model, types, operations, derived, lifecycle, conformance |
YAML remains descriptive; only its linked JSON was translated in the pilot |
DayDreams gallery.vm.yaml |
vm, version, description, model, children, state, types, dependencies, operations, lifecycle, conformance |
Not parsed or standardized |
DayDreams world.vm.yaml |
vm, version, description, model, children, registry, heightfields, types, operations, scene_bridge, lifecycle, conformance |
Not parsed or standardized |
| DayDreams AppVM JSON | initialRoute; cases[] with id, description, bespoke when, and bespoke then fields |
AVM-001/002 were translated to suite fixture plus ordered invoke/state/message steps in the no-push pilot |
| DayDreams GalleryVM JSON | entriesFixture; cases[] with id, description, when.op, when.args, and domain-specific then fields |
Factory/fixture adaptation is feasible but not included |
| DayDreams WorldVM JSON | sceneFixture, cases, streamingCases; cases use id, optional description, assert/equals or domain-specific world, when, and then |
Multiple setup modes and domain serializers remain consumer-owned; not included |
10.5.3. Suite Shape¶
Unknown structural fields fail validation. Fixtures, invocation arguments, expected state, and normalized messages accept arbitrary JSON values.
{
"$schema-version": "1.0.0",
"suite": "app-vm",
"fixture": { "initialRoute": "gallery" },
"cases": [
{
"id": "AVM-001",
"steps": [
{ "kind": "invoke", "operation": "navigate", "args": ["world"] },
{
"kind": "assert-state",
"path": "/model/route",
"equals": "world"
},
{
"kind": "assert-messages",
"equals": [
{
"type": "PropertyChangedMessage",
"propertyName": "model",
"senderName": "appVm"
}
]
}
]
}
]
}
assert-state uses RFC 6901 decoding and distinguishes a missing path from a
present null. Object key order is irrelevant; array and message order are
exact. assert-messages drains the adapter once for each assertion.
10.5.4. Factory And Runner¶
The runner owns sequencing and teardown; the consumer owns construction, operation dispatch, JSON snapshots, and message normalization.
import {
parseConsumerConformance,
runConsumerConformance,
type ConsumerConformanceFactory,
} from "@thekaveh/vmx/conformance";
const factory: ConsumerConformanceFactory = ({ caseFixture }) => {
const vm = createAppVm(caseFixture);
return {
invoke: async (operation, args) => invokeApp(vm, operation, args),
snapshot: () => snapshotApp(vm),
drainMessages: () => messageRecorder.drain(),
dispose: () => vm.dispose(),
};
};
const suite = parseConsumerConformance(input);
const report = await runConsumerConformance(suite, factory);
runConsumerConformanceCase integrates one case with any test framework.
runConsumerConformance executes all cases sequentially and returns passed and
failed results. The runner imports no Vitest or Jest API.
Factory, operation, state, message, and disposal failures include actionable
instance paths. Disposal runs exactly once in finally; if execution and
disposal both fail, the execution error stays primary and retains the teardown
cause.
adaptCommandTruthTableFixture demonstrates adapting every unchanged VMx
command fixture row. The existing CMD-007 test and all five-flavor conformance
coverage remain in place.
10.5.5. Non-Goals And Follow-Up Gate¶
Adapter v1 does not parse YAML, define a VMx viewmodel dialect, generate tests, or generate Swift. A Swift or code-generation proposal needs a separate ADR, successful use by two independent consumers, a native Swift factory without TypeScript assumptions, and a maintained domain-type mapping that makes generation safer than manual implementation.
10.6. Test Marker Grammar¶
The coverage checker recognizes one intentional marker form per flavor:
| Flavor | Marker |
|---|---|
| C# | [Trait("Conformance", "LIFE-001")] |
| Python | @pytest.mark.conformance("LIFE-001") |
| TypeScript | describe("LIFE-001", ...) |
| Swift | // LIFE-001 — description or /// LIFE-001 — description, attached to a test function |
| Rust | /// LIFE-001 — description attached through #[test] to a test function |
For Rust, the ID must be the first token after /// and an em dash must follow
it on the same line. Only doc-comment and attribute lines may separate the
marker from #[test] and its function. Ordinary comments, file summaries,
unattached markers, and markers in block-commented tests are ignored. Duplicate
markers are one set-based coverage claim. In required mode, a missing Rust ID is
listed under MISSING and fails the coverage command.
In required mode, a marker for an ID the catalog does not define is listed under
ORPHAN and fails the coverage command for that flavor. It claims coverage of
nothing, and usually means a renamed or retired ID whose real test lost its
marker. Without --require, the report still lists it for information.
10.7. Practical Reading Path¶
- Read
spec/README.mdfor chapter ownership and release history. - Read the primitive pages on this site for a faster conceptual map.
- Use the flavor README when you need package details or host-specific examples.
- Use the parity matrix and conformance catalog when you need proof rather than overview.
10.8. Controlled Concurrency Interleavings¶
Concurrency regressions must acknowledge the boundary that creates the tested ordering. A worker-start signal only proves that a worker began; elapsed time only proves that a deadline passed. Neither proves that VMx entered a specific wait. The selected lifecycle and notification regressions instead acknowledge the exact per-instance wait, then inspect the protected state while that wait is active. The Swift command regression uses the synchronous executing notification to cancel inside the admission window, before the body task's cancellation handle is installed.
The test owns every participant and error channel. It releases gates in cleanup, joins workers, awaits task results, and reports callback failures to the coordinator. The forbidden outcomes are premature completion before the acknowledged boundary is released, missing terminal delivery, a swallowed worker or subscriber error, and cleanup that leaves work running. Diagnostic timeouts fail a test; they never count as evidence that an operation was correctly blocked.
Negative controls mutate the actual production boundary exercised by the test,
run in a subprocess group with a watchdog, and restore the original source
bytes in finally. A valid control either reaches the selected assertion or is
terminated as a demonstrated deadlock mutant. Compilation failure, zero matched
tests, an unrelated assertion, crash, or watchdog expiry cannot be reported as
the selected assertion proof. CI repeats the restored Linux checks normally and
with one CPU from the runner's allowed affinity set. The Swift lane records five
normal selected passes, five failures from removing admission-cancellation
forwarding, and a rebuilt restored pass.
Order test infrastructure as carefully as runtime participants. A package build that rewrites imported fixtures or schemas must finish before workers read them. TypeScript uses global setup and an awaited rerun hook for this boundary; generated outputs are excluded from watch triggers and explicitly invalidated after rebuilding. A filesystem error alone does not identify the process holding a lock: trace the actual writer and reader before assigning a cause.
These techniques answer different questions:
- controlled thread interleavings force and acknowledge a particular runtime boundary;
- stress tests sample many schedules and remain useful supplementary evidence;
- event-loop drains allow already-queued continuations to run but do not prove that a foreign thread reached a boundary;
- virtual schedulers deterministically advance modeled time and do not establish native-thread scheduling behavior.
The controls establish the named paths, not exhaustive freedom from races. Linux CPU affinity does not represent macOS scheduling, and an external watchdog contains a stuck process without proving in-process cleanup. The maintained scope and follow-up dispositions are recorded in the Concurrency Test Audit.