Testing
Running Tests
npm test # Run unit + conformance suites once
npm run test:watch # Watch mode — re-runs on file changes
npm run test:coverage # Run with v8 coverage report + threshold gateIntegration tests are not part of npm test — they run separately against
a live runtime (see Integration Tests below).
Test Structure
tests/
├── unit/
│ ├── projections/ # Per-mode projection state machines (5 files) + accepted-only-contract.test.ts + message-id-dedup.test.ts
│ ├── sessions/ # Per-mode session helpers (5 files) + session-id-validation.test.ts
│ ├── agent/ # Dispatcher, participant, strategies, transports, runner, cancel-callback
│ ├── helpers/
│ │ └── grpc-stub.ts # stubUnary() — shared gRPC stubbing helper (not a test file)
│ ├── auth.test.ts # Auth factory, identity guard, metadata
│ ├── base-session.test.ts # BaseSession extension point + BaseProjection outcome table
│ ├── client.test.ts # TLS guard, sender-identity enforcement, public exports
│ ├── commitment-hash-frozen-fields.test.ts # Runs real tsc over mutated CommitmentPayload copies
│ ├── client-stream.test.ts # MacpStream data path (openStream, responses, read, close)
│ ├── client-unary.test.ts # Full unary RPC surface + metadata/deadline dispatch matrix
│ ├── conformance-guard.test.ts # duplicateAcceptedBallots() over synthetic input
│ ├── envelope.test.ts # Envelope builder functions
│ ├── errors.test.ts # Error class hierarchy
│ ├── fixture-drift-gate.test.ts # Drives the real `make verify-fixtures`/`sync-fixtures`/`verify-parity`/`sync-parity` recipes
│ ├── logging.test.ts # Structured logger + configureLogging()
│ ├── policy.test.ts # Policy builders
│ ├── proto-registry.test.ts # Protobuf encode/decode roundtrips
│ ├── public-api.test.ts # Public runtime-surface snapshot guard vs. public-api-snapshot.json
│ ├── public-api-snapshot.json # Committed snapshot (not a test file)
│ ├── retry.test.ts # Retry policy + backoff
│ ├── validation.test.ts # Runtime-adjacent payload validation
│ └── watchers.test.ts # Registry/roots/signal/policy/session-lifecycle watchers
├── conformance/
│ ├── conformance.test.ts # Fixture-driven projection replay harness + anomaly/duplicate guards
│ ├── duplicate-ballots.ts # duplicateAcceptedBallots() — non-test module, imported by both the harness and conformance-guard.test.ts
│ ├── schema.json # Fixture schema (shared with the spec repo)
│ └── *_happy_path.json / *_reject_paths.json / … # Per-mode fixtures + ext.multi_round.v1
├── vectors/
│ ├── cmt-hash.test.ts # Spec-vector runner replaying RFC-MACP-0013's canonical vectors
│ └── cmt-hash/ # Vendored vectors + SOURCE.md (provenance, gated by verify-fixtures)
├── parity/
│ ├── contract.test.ts # Cross-SDK parity contract: asserts runtime values against contract.json
│ ├── contract.json # Vendored copy (provenance, gated by verify-parity)
│ └── SOURCE.md # Provenance note
├── commitment-hash.test.ts # RFC-MACP-0013 commitment hash: determinism, JCS, D3
└── integration/
├── README.md # Runtime setup + bearer envs
└── runtime.test.ts # Full-surface tests against a live runtimeCoverage Gates
vitest.config.ts enforces v8 coverage thresholds over src/** (pure
type/barrel files are excluded so they don't skew the function percentage):
| Metric | Floor | Measured (2026-08-31 suite) |
|---|---|---|
| Lines | 93 | 95.65 |
| Branches | 83 | 85.28 |
| Functions | 90 | 92.65 |
| Statements | 92 | 94.13 |
The convention: floors are the current measured value minus 2 percentage
points, and they are raised when new tests land. CI gates on these via
npm run test:coverage — if coverage drops below a floor, the run fails.
The json/json-summary reporters feed the sticky PR coverage comment in CI.
The v3 → v4 measurement break
These numbers are lower than the floors this table carried before 2026-08
(94 / 88 / 90 / 94) even though no test was removed. Vitest 4 rewrote the v8
coverage provider to remap through a rolldown AST instead of v8-to-istanbul,
and dropped the ignoreEmptyLines option along with it. The new mapping counts
branches the old one missed — optional chaining, default parameters, logical
short-circuits — so the same suite measures stricter.
The practical consequence: v3-era and v4-era percentages are not comparable.
A number from a pre-2026-08 run, a PR coverage comment, or an old PROGRESS.md
entry is on a different ruler. Do not read the drop as a regression, and do not
try to recover the old figures by widening exclude.
Client Transport Tests
tests/unit/client-unary.test.ts and tests/unit/client-stream.test.ts
exercise MacpClient and MacpStream without any network, sharing the
stubUnary helper from tests/unit/helpers/grpc-stub.ts.
stubUnary
stubUnary(client, rpcName, response, options?) replaces one unary RPC on
the private gRPC client behind a real MacpClient and records every call.
MacpClient.unary() dispatches on four argument shapes depending on whether
auth metadata and a deadline are present — (req, cb), (req, metadata, cb),
(req, {deadline}, cb), (req, metadata, {deadline}, cb). The callback is
always the last function-typed argument, so one helper covers all four;
everything between the request and the callback is recorded in
calls[i].extras for assertions. Pass { fail: true } with an
Error-like { code, details, message } response to make the RPC fail.
import { stubUnary } from '../helpers/grpc-stub';
const calls = stubUnary(client, 'GetSession', {
metadata: { sessionId: 's-1', state: 'SESSION_STATE_OPEN' },
});
await client.getSession('s-1');
expect(calls).toHaveLength(1);
expect(calls[0].request).toEqual({ sessionId: 's-1' });
// calls[0].extras holds the metadata / {deadline} positional args, if anyclient-unary.test.ts covers the full unary surface (initialize, send,
discovery/registry RPCs, policy RPCs, sendSignal/sendProgress, watch-stream
factories) plus the metadata/deadline dispatch matrix itself.
Stream data path
client-stream.test.ts pins down MacpStream semantics:
- Envelope unwrap — an envelope frame is read from
chunk.envelope.chunk.responseis the proto3oneofarm name (a string, e.g.'envelope'/'error'), not a nested object — verified by a real encode/ decode roundtrip via therealStreamResponse()helper, not a hand-built literal. - Inline errors — an application-level
chunk.errorinvokesonInlineErrorcallbacks and logs onewarn, while the stream stays open. read()timeouts —read(timeoutMs)throwsMacpTimeoutErrorwhen no envelope arrives in time.STREAM_ENDsemantics — end-of-stream is sticky:read()returnsnulland returnsnullagain on the next call (the sentinel is re-pushed), andresponses()observes the same end.
Session and Base-Session Tests
tests/unit/sessions/ has one file per mode (decision, proposal, task,
handoff, quorum). Each covers, with client.send mocked:
- Projection roundtrip — each helper method applies its envelope to the
projection on
ack.ok === trueand does not whensendthrowsMacpAckError. - cancel/suspend/resume delegation —
session.cancel()/suspend()/resume()delegate toclient.cancelSession/suspendSession/resumeSessionwith the session's id. - Resolved NACK not applied — when
sendresolves with{ ok: false, ... }(e.g.raiseOnNack: falseflows), the ack is returned to the caller but nothing is applied to the projection.
tests/unit/base-session.test.ts covers the BaseSession extension point
(the recommended base class for custom/extension modes): input validation,
commit() feeding the projection, and a branch table for
BaseProjection.isPositiveOutcome (undefined without a commitment; defaults
applied per commitment shape).
Participant Tests
tests/unit/agent/participant.test.ts exercises the agent framework's run
loop end to end (with a fake transport):
- Action delegation —
ctx.actions.evaluate/vote/propose/commit/senddelegate to the mode session with the participant's id assender. - Run-loop terminal path — reaching a terminal phase fires the
onTerminalhandler and exitsrun(). - Cancel-callback wiring — a
cancelCallbackconfig starts the HTTP server onrun()and tears it down onstop(). - Kickoff defaults — an initiator
kickoffProposal defaultsproposalIdto<sessionId>-kickoffandoptionto"decide", and accepts the snake_caseproposal_idspelling.
The file's makeIncomingMessage helper mints a distinct messageId per
call. Since applyEnvelope (src/projections/base.ts) dedups by
message_id, reusing one hardcoded id across multiple
messages in the same test would silently drop every message after the first
from the real projection instance the session applies to — invisible to
tests that only assert handler calls, since processMessage dispatches
handlers after applying to the projection. 'two distinct envelopes both land on the projection (not silently deduped)' is the regression test for
this: it asserts projection state (evaluations.length), not just handler
invocation counts.
Writing Projection Tests
Projections are pure state machines — they accept envelopes and update internal state. This makes them ideal for unit testing without any I/O:
import { describe, it, expect, beforeEach } from 'vitest';
import { DecisionProjection } from '../../../src/projections/decision';
import { ProtoRegistry } from '../../../src/proto-registry';
import { buildEnvelope } from '../../../src/envelope';
import { MODE_DECISION } from '../../../src/constants';
// Create a real ProtoRegistry — tests the full protobuf round-trip
const registry = new ProtoRegistry();
// Helper to build envelopes with encoded payloads
function makeEnvelope(
messageType: string,
payload: Record<string, unknown>,
sender = 'agent-a',
) {
return buildEnvelope({
mode: MODE_DECISION,
messageType,
sessionId: 'test-session',
sender,
payload: registry.encodeKnownPayload(MODE_DECISION, messageType, payload),
});
}
describe('DecisionProjection', () => {
let projection: DecisionProjection;
beforeEach(() => {
projection = new DecisionProjection();
});
it('tracks proposals and transitions phase', () => {
projection.applyEnvelope(
makeEnvelope('Proposal', { proposalId: 'p1', option: 'A' }),
registry,
);
expect(projection.proposals.size).toBe(1);
expect(projection.phase).toBe('Evaluation');
});
it('computes vote totals correctly', () => {
projection.applyEnvelope(
makeEnvelope('Proposal', { proposalId: 'p1', option: 'A' }),
registry,
);
projection.applyEnvelope(
makeEnvelope('Vote', { proposalId: 'p1', vote: 'approve' }, 'alice'),
registry,
);
projection.applyEnvelope(
makeEnvelope('Vote', { proposalId: 'p1', vote: 'approve' }, 'bob'),
registry,
);
expect(projection.voteTotals()).toEqual({ p1: 2 });
});
});Key Testing Pattern
Tests use a real ProtoRegistry instance that loads actual .proto files. This means:
- Payloads are encoded to protobuf wire format then decoded back
- Field name casing (camelCase ↔ snake_case) is exercised
- Missing or extra fields are caught
- Protobuf default values are handled correctly
What to Test
For each projection:
- State transitions: Verify
phasechanges at the right time - Record population: Check that maps/arrays are updated correctly
- Query helpers: Test convenience methods with edge cases
- Commitment handling: Verify terminal state
- Mode isolation: Envelopes for other modes are ignored
- Redelivery idempotence: applying the same
message_idtwice must not double-append totranscriptor to any accumulate-on-apply site (seetests/unit/projections/message-id-dedup.test.ts, and Projections API › Redelivery) - Anomalies surface:
anomaliesstarts empty andhasAnomaliesstartsfalseon every projection; onlyDecisionProjection(duplicateVote) andQuorumProjection(duplicate ballot) populate it (see Projections API › Anomalies)
Testing redelivery (message_id dedup)
tests/unit/projections/message-id-dedup.test.ts covers applyEnvelope's
RFC-MACP-0006 §3.2 redelivery idempotence across all six entry points (the
five built-in mode projections plus a third-party BaseProjection subclass).
The pattern to copy for a new accumulate-on-apply site:
it('redelivery does not duplicate the <site>', () => {
const projection = new SomeProjection();
const envelope = buildEnvelope({
mode: MODE_SOME_MODE,
messageType: 'SomeMessageType',
sessionId: 'test-session',
sender: 'sender',
messageId: 'm-reused', // explicit, REUSED — never buildEnvelope's auto-minted id
payload: registry.encodeKnownPayload(MODE_SOME_MODE, 'SomeMessageType', { ... }),
});
projection.applyEnvelope(envelope, registry);
projection.applyEnvelope(envelope, registry); // redelivery
expect(projection.someAccumulateSite).toHaveLength(1);
});Two traps this file guards against, worth remembering when adding coverage elsewhere:
- Always pass an explicit, reused
messageId.buildEnvelopeauto-mints a fresh id per call when one isn't given, so two calls without an explicit shared id exercise the cardinality path, not the dedup path — the test would pass while testing nothing. - Empty
messageId('') must never be deduped — the guard is gated onif (envelope.messageId). A test asserting the debug log line must useconfigureLogging({ level: 'debug', sink }); the SDK's default level (warn) suppressesdebugentirely, so a test left at the default level passes vacuously regardless of whether the log call fires.
Testing the anomalies surface
tests/unit/projections/anomalies.test.ts covers the anomalies surface:
types, the anomalies field, the hasAnomalies getter, and
BaseProjection.recordAnomaly. All five built-in mode projections extend
BaseProjection (issue #91), and two of them — DecisionProjection and
QuorumProjection — call the inherited recordAnomaly directly at their
duplicate-detection call sites (see Projections API ›
Anomalies); the other three never record
an anomaly today. The tests additionally exercise recordAnomaly through a
synthetic third-party BaseProjection subclass, independent of either
built-in mode's own call site.
Two things worth copying when Phases 4-5 add real detection (Decision Vote,
Quorum ballots), or when testing a custom ext-mode's own anomaly detection:
- Test the array and the
logger.warncall as separate assertions in separate tests, not combined in one test with the array assertion first. If an assertion earlier in a test throws, later assertions in the same test body never run — so a mutation that breaks only the log call can look like it also "breaks" an array assertion that was never actually re-verified, and vice versa. Splitting them into siblingit()s (mirrored inanomalies.test.ts: "... (array half only)" / "... (sink half only)") makes each half's non-vacuity independently provable: deleting thepushfails only the array-half tests, deletinglogger.warnfails only the sink-half tests. - Construct a fresh instance per test and assert non-sharing explicitly.
anomalies(liketranscript) is areadonlyinstance field initialized in a property initializer, butreadonlydoes not prevent two instances from accidentally sharing the same array if a future refactor moves the initializer to a shared location (e.g. a static or prototype default). Asserta.anomalies !== b.anomaliesacross two instances, not just that lengths differ.
Conformance Tests
tests/conformance/conformance.test.ts replays the shared spec fixtures
(*_happy_path.json, *_reject_paths.json, …) through the projections:
each fixture's accepted message prefix is encoded via a real
ProtoRegistry, applied to the mode's projection, and the resulting phase,
resolution, and mode state are asserted against the fixture's expectations.
The harness's acceptedMessages filter in
tests/conformance/conformance.test.ts — fixture.messages.filter((m) => m.expect === 'accept') — is not an incidental implementation detail — it is
this harness upholding
applyEnvelope's accepted-only input contract (see Projections API ›
Input contract). applyEnvelope
cannot verify that contract itself (Envelope carries no acceptance marker),
so every caller, including this harness, is responsible for filtering before
replay. tests/unit/projections/accepted-only-contract.test.ts demonstrates
what happens when a caller skips that filtering step.
The fixture-sync contract — what a conformant SDK must replay and how fixtures and protos are kept in lockstep with the spec — is canonical in the spec's SDK Parity › Sync Mechanisms and SDK Parity › Conformance Test Suite.
Harness behaviours worth knowing:
ext.multi_round.v1fixtures replay too. The extension mode has no mode-specific projection, so its fixtures run through a transcript-onlyBaseProjectionsubclass — they are no longer silently skipped.- Unmapped modes fail loudly. A newly synced fixture whose
modehas no projection mapping fails the suite instead of being skipped, so fixture drift is caught at sync time. - Reject-path fixture contract. A dedicated test asserts that every
expect: "reject"message carries a canonical, non-emptyexpected_error_code(one of the NACK codes insrc/constants.ts) and a resolvablepayload_type. - Runtime NACK codes are out of scope here. This in-process harness only
replays the accepted prefix; whether the runtime actually returns each
expected_error_codeis asserted by the macp-runtime conformance oracle (see the runtime testing guide). The suite carries explicitit.skipmarkers documenting that split rather than pretending to cover it. - Zero anomalies, zero anomaly warnings, on every canonical fixture.
After the transcript-length assertion, each replay also asserts
projection.anomaliesis empty and that nologger.warn('projection anomaly', ...)fired during that replay (an injected sink, reset per fixture, makes the log channel observable). The array and the log are independent channels — see Projections API › Anomalies — and both must be silent for a conforming transcript. A non-empty result here means either a fixture regressed to carrying a duplicate accepted Vote/ballot (the next bullet should already have caught that) or a projection is over-flagging. - No duplicate accepted Vote or ballot in any fixture. A sibling
describe(conformance: no duplicate accepted vote or ballot) runsduplicateAcceptedBallots(tests/conformance/duplicate-ballots.ts) over every fixture's messages and asserts[]. The predicate is scoped toexpect === 'accept'— a rejected duplicate is exactly the missing upstream fixture this SDK has requested from the spec repo (multiagentcoordinationprotocol#84; adecision_reject_paths.json/quorum_reject_paths.jsoncase with"expect": "reject"/"expected_error_code": "INVALID_ENVELOPE"for a duplicate Vote/ballot), and this guard must welcome that fixture landing, not treat it as a violation. The predicate lives in its own non-test module, not inconformance.test.ts, sotests/unit/conformance-guard.test.tscan unit-test it against synthetic input without re-registeringconformance.test.ts's three module-scopedescribeblocks. See that module's docblock for why the ballot arm must be gated on the message's ownpayload_typerather than onmessage_typename alone —Rejectis defined identically-named in both Proposal and Quorum mode, and both are live in the canonical corpus. - Fixtures are canonical — never hand-edit
tests/conformance/*.json. If either guard above ever fails against a real fixture, the fix belongs upstream in the spec repo;make verify-fixturesdiffs this SDK's copies against the canonical source bidirectionally, so a local edit would just fail that gate asDRIFTinstead of fixing anything.
Fixture Drift Gate
make verify-fixtures diffs this SDK's vendored fixtures against the spec repo's
canonical copies and fails if either has drifted; make sync-fixtures refreshes
the vendored copies from a local spec checkout. Both targets cover two fixture
sets in one invocation:
tests/conformance/*.jsonagainst$(SPEC_CONFORMANCE_DIR)/*.json(flat).tests/vectors/cmt-hash/*.jsonagainst$(SPEC_CONFORMANCE_DIR)/cmt-hash/*.json(the RFC-MACP-0013 commitment-hash vectors — seetests/vectors/cmt-hash/SOURCE.mdfor that directory's provenance).
verify-fixtures reports every problem across both sets before exiting — a
canonical file that differs from (or is missing from) the vendored copy prints
DRIFT:, and a vendored file with no canonical source prints EXTRA:. CI runs
this gate on every PR via .github/workflows/conformance-fixtures.yml, so a
spec-side fixture or vector change that isn't synced here fails the build.
sync-fixtures copies but never deletes: if a canonical file is renamed or
removed upstream, the sync leaves the orphaned vendored file in place and
verify-fixtures keeps reporting it as EXTRA: until it's removed by hand.
make verify-fixtures # uses the default sibling-checkout path
make verify-fixtures SPEC_CONFORMANCE_DIR=/path/to/schemas/conformance
make sync-fixtures SPEC_CONFORMANCE_DIR=/path/to/schemas/conformanceParity Contract Gate
A second, single-file zero-drift gate covers the cross-SDK parity contract:
schemas/parity/contract.json in the spec repo, a small non-normative
manifest pinning values that already agree — by convention, not by a shared
source of truth — across macp-runtime, macp-sdk-python, and this SDK
(protocol version, standard/extension mode ids, defaults, error codes, retry
policy, the ProjectionAnomaly field set/order, the commitment-hash
accept/reject behavior, and the Contribute payload byte vectors). See
tests/parity/SOURCE.md for provenance and
tests/parity/contract.test.ts for
the assertions against this SDK's actual runtime values.
make verify-parity/make sync-parity mirror verify-fixtures/sync-fixtures,
scoped to the one vendored file (tests/parity/contract.json) instead of a
fixture directory:
make verify-parity # uses the default sibling-checkout path
make verify-parity SPEC_PARITY_DIR=/path/to/schemas/parity
make sync-parity SPEC_PARITY_DIR=/path/to/schemas/parityCI runs both gates in the same job (.github/workflows/conformance-fixtures.yml),
reusing one spec-repo checkout for both. contract.json is explicitly
non-normative — see its own $comment field and the spec repo's
schemas/parity/README.md — so neither this guide nor any other doc in this
repo should cite contract.json itself as a source of truth; cite the named
RFC/registry/proto file instead.
Proto Registry Tests
Test encode/decode roundtrips for every message type across all modes:
it('decision/Proposal roundtrip', () => {
const payload = { proposalId: 'p1', option: 'deploy', rationale: 'ready' };
const encoded = registry.encodeKnownPayload(MODE_DECISION, 'Proposal', payload);
const decoded = registry.decodeKnownPayload(MODE_DECISION, 'Proposal', encoded);
expect(decoded).toHaveProperty('proposalId', 'p1');
expect(decoded).toHaveProperty('option', 'deploy');
});Integration Tests
Integration tests drive the full SDK against a live MACP runtime in Docker.
They are local-only — excluded from npm test (and from CI) via the
vitest config split, and run with their own config
(vitest.integration.config.ts). See
tests/integration/README.md for the
full harness (runtime setup, env vars, and the optional direct-agent-auth
block gated on MACP_TEST_BEARER_ALICE / MACP_TEST_BEARER_BOB).
Common flow:
docker build -t macp-runtime ../macp-runtime/
docker run -d --name macp-runtime-test -p 50051:50051 \
-e MACP_BIND_ADDR=0.0.0.0:50051 -e MACP_ALLOW_INSECURE=1 \
-e MACP_MEMORY_ONLY=1 macp-runtime
npm run test:integration
docker rm -f macp-runtime-testWriting Integration Tests
import { describe, it, expect } from 'vitest';
import { Auth, MacpClient, DecisionSession } from '../../src';
const address = process.env.MACP_RUNTIME_ADDRESS ?? 'localhost:50051';
describe('MacpClient integration', () => {
it('initializes successfully', async () => {
const client = new MacpClient({
address,
secure: false,
allowInsecure: true, // local runtime uses MACP_ALLOW_INSECURE=1
auth: Auth.devAgent('test-agent'),
});
const init = await client.initialize();
expect(init.selectedProtocolVersion).toBe('1.0');
client.close();
});
});