Deployment Guide

This guide covers everything you need to run the MACP Runtime in production: configuration, storage backends, crash recovery, monitoring, and container deployment. For protocol-level deployment topologies and security requirements, see the protocol deployment and protocol security documentation.

Production checklist

Before exposing the runtime to production traffic, ensure these four items are configured:

  1. TLS certificates -- Set MACP_TLS_CERT_PATH and MACP_TLS_KEY_PATH to valid PEM files. The runtime refuses to start without TLS unless MACP_ALLOW_INSECURE=1 is set.

  2. Authentication -- Configure at least one of the resolvers. For opaque bearer tokens, create a tokens.json mapping tokens to agent identities and set MACP_AUTH_TOKENS_FILE. For JWT bearer tokens, set MACP_AUTH_ISSUER together with a JWKS source (MACP_AUTH_JWKS_JSON inline or MACP_AUTH_JWKS_URL fetched + cached). Both can be configured at once -- JWT-shaped tokens are routed to the JWT resolver and opaque tokens to the static resolver. See the Getting Started guide for the token format and JWT claim layout.

  3. Data directory -- Ensure MACP_DATA_DIR points to a directory with write permissions. This is where session logs and snapshots are stored.

  4. Bind address -- Set MACP_BIND_ADDR to the desired listen address. The default 127.0.0.1:50051 only accepts local connections.

Upgrading into registration-time policy validation

This release tightens what the governance policy registry accepts, what the Quorum mode will bind, how the Decision evaluator reads a weighted round, when a decline may be finalized over a passing vote, and how a percentage Quorum threshold is computed. Ten changes are operationally visible -- seven tightenings, one relaxation (item 7, which cannot affect an existing deployment) and two corrections that lower a bar (items 9 and 10) -- and item 6 is the only one that can change how an already-stored session replays -- read it first if you have any persisted Decision session at all. Item 6 carries two independent predicates and only one of them involves weighted: the other reaches any Decision policy that sets commitment.allow_decline_over_approval: true, majority policies included. Read this section before upgrading any deployment that sets MACP_POLICIES_DIR, or that has persisted sessions bound to a policy with a Quorum threshold, a weighted voting.algorithm, or commitment.allow_decline_over_approval: true. CHANGELOG.md is generated from commit subjects and does not carry this detail.

1. An invalid policy file now refuses startup

Both routes into the registry -- the RegisterPolicy RPC and the MACP_POLICIES_DIR preload -- gained value-domain and conditional checks; the enforced set is listed in Policy. A file an earlier release accepted may now be out of domain: a fractional or zero Quorum threshold.value, threshold.type: "weighted", an unknown voting.algorithm or voting.quorum.type, a weighted algorithm with an empty weights map, a supermajority threshold at or below 0.5 (including one that omits threshold entirely and relies on the 0.5 default), or a wildcard ("*") policy carrying a Quorum threshold that was previously validated against the Decision schema alone and therefore never checked. Loading stops at the first rejection and startup aborts -- the preload error is propagated, not logged and skipped.

That is deliberate fail-closed behaviour, and it is why MACP_POLICIES_DRY_RUN=1 exists. Run the dry run with the new binary before you upgrade:

MACP_POLICIES_DRY_RUN=1 MACP_POLICIES_DIR=/etc/macp/policies macp-runtime

It reports every *.json file by name as OK <path> (<policy_id>) or REJECTED <path>: <reason>, prints a checked/rejected count, and exits 0 if the directory would load or 1 if anything in it would be rejected. It binds no port, opens no storage, and replays nothing, so it needs neither TLS nor MACP_ALLOW_INSECURE=1. Its output is written to stdout/stderr directly rather than through tracing, so no RUST_LOG filter can suppress the report. Two things to know: the startup environment-configuration check still runs ahead of it, so an unrelated malformed variable aborts before the report is produced; and a readable directory containing no *.json exits 0 with an explicit WARNING, because a mis-pointed MACP_POLICIES_DIR otherwise looks identical to a clean run.

2. Do not unblock startup by deleting the rejected file

When a policy file blocks startup, the natural fix is to delete it. Correct the file instead. Deleting it does let the runtime boot, but persisted sessions bound to that policy_version are then replayed with the policy unresolved: replay resolves the version best-effort and leaves policy_definition empty when it cannot, and commitment enforcement treats an absent policy definition as "no policy to enforce" and returns early. Every in-flight session governed by the deleted policy therefore loses its governance silently -- commitments the policy would have denied are accepted, with no error and no log line tying it back to the deletion. The same applies to UnregisterPolicy on a policy that live sessions are still bound to.

Note the interaction with the next item: because an unresolved policy makes the Quorum mode fall back to the ApprovalRequest's own required_approvals -- a value the mode already constrains to 1..=participants -- deleting the policy also makes the replay failure below disappear. The two symptoms clear together, and the reason they clear is that the governance bar is no longer being applied.

3. A persisted Quorum session with an out-of-domain policy threshold no longer replays

RFC-MACP-0011 §5 rule 6 makes a policy threshold replace the ApprovalRequest's required_approvals, but nothing previously held the replacement to the same 1..=participants domain the runtime enforces on the field it replaces. The Quorum mode now refuses an ApprovalRequest whose effective threshold falls outside that domain, and replay dispatches the same code -- so such a session fails to replay. It is skipped with a warning, or is fatal at startup under MACP_STRICT_RECOVERY=1.

Detect it from the logs. The mode emits, at WARN:

quorum policy threshold is outside 1..=participants; refusing the ApprovalRequest
  session_id=... policy_id=... effective_threshold=... participants=...

naming the session, the bound policy, the computed threshold and the declared participant count -- everything needed to identify which policy to correct. Recovery follows it with failed to replay session; skipping carrying the same session_id.

What it takes to reach this. Not a legacy policy -- an ApprovalRequest this runtime accepted before the guard above existed. Registration is no substitute for the guard, because registration has no participant count to bound the threshold against: {"type": "n_of_m", "value": 66} passes every check in Policy under the new binary and still trips the guard on a three-participant session. Two classes of threshold reach it, and they are not equally benign:

  • An out-of-domain numeric threshold -- n_of_m or count above the declared participant count. The positive outcome was unreachable from the first message, and before this release commitment_ready carried no counted > 0 guard, so approvals + remaining < required held with no ballot cast and the coordinator could seal a binding quorum.rejected with zero approvals (issue #145). That is the condition RFC-MACP-0011 §5 rule 4a reads as grounds for a decline, reached without a vote. Such a session really was broken: the only outcome it could ever have sealed was a decline nobody cast a ballot for.
  • threshold.type: "weighted", or any unrecognised type. This class was working, and it stops replaying. The old shared fallback arm read any unrecognised type as a raw approval count (_ => rules.threshold.value as u32), so {"type": "weighted", "value": 2} on three participants was a perfectly satisfiable bar of two approvals, and sessions under it sealed legitimate positive commitments. The type now resolves to Unsatisfiable, the ApprovalRequest is refused, and the session no longer loads. Do not read the warning as a report of a session that was already dead.

The second class survives the upgrade through a checkpoint, not through the registry. A checkpoint serializes the resolved policy_definition inline and try_replay_from_checkpoint restores it verbatim without consulting the registry, so an old weighted definition is still live even though neither RegisterPolicy nor the MACP_POLICIES_DIR preload would accept it again. A session with no checkpoint re-resolves its policy_version against the live registry during full replay, and there a weighted policy file aborts startup at item 1 before recovery ever runs.

Recovery, for both classes: correct the threshold, do not delete the policy. Restate a weighted or unrecognised type as n_of_m, keeping the same value. QuorumThreshold::effective treats n_of_m as a raw approval count, which is exactly what the old fallback arm did, so the bar that session enforced is preserved. (If the old value was fractional, registration now refuses it; the old arm truncated, so its floor is the faithful integer.) For a numeric threshold above the participant count, bring it into 1..=participants -- no value reproduces that session's old behaviour, because its old behaviour was the zero-approval decline. Then restart: the append-only log is untouched, so the session was not loaded rather than lost, and it replays normally. Deleting the policy also clears the warning, but for the reason item 2 gives -- the governance bar stops being applied at all.

4. A negative weighted total now fails the Decision round

A weighted round whose cast weights sum below zero fails the round instead of computing a ratio over a negative denominator. In the approve direction this is a tightening: a round that previously reported Passed through an inverted ratio >= threshold comparison is now denied. In the decline direction it is not a tightening -- on that same round a negative commitment moves from denied to allowed, because a decline over Passed was refused while a decline over Failed is permitted once the universal reject-floor is satisfied. The case is reachable only from a directly-constructed PolicyDefinition, since registration already refuses negative weights. A weighted total of exactly zero is unchanged.

5. voting.threshold: 0.0 and zero voting.weights entries are no longer accepted

Spec #99 moved two Decision bounds in decision-rules.schema.json from inclusive to exclusive at zero -- voting.threshold to exclusiveMinimum: 0, and voting.weights.additionalProperties to exclusiveMinimum: 0 with minProperties: 1 on the map. This runtime mirrors both, and adds the schema's majority arm: a majority threshold below 0.5 is refused, where supermajority continues to require one strictly above 0.5. The asymmetry is deliberate -- the reserved policy.std.majority profile sets exactly 0.5.

A policy file an earlier release accepted may now be refused, and because the MACP_POLICIES_DIR preload aborts startup at the first rejection, a deployment carrying any of these on disk will fail to start:

  • voting.threshold: 0.0 (it made an all-REJECT round return Passed under both majority and weighted)
  • a voting.weights entry of 0.0, or a supplied but empty voting.weights: {}
  • a majority voting.threshold below 0.5

Run the dry run with the new binary before you upgrade -- it is the same pre-upgrade check item 1 describes:

MACP_POLICIES_DRY_RUN=1 MACP_POLICIES_DIR=/etc/macp/policies macp-runtime

Correcting a zero weight is not a matter of picking a small positive number. The weights map is the weighted electorate: a participant who should carry no voting weight is expressed by omission from the map, never by an explicit 0. Remove the entry rather than nudging it above zero. A map that would be left empty means no weighted electorate at all, which the weighted algorithm cannot express -- choose a different algorithm.

This item affects admission only; no stored session's replay changes, because a descriptor carrying any of these values evaluated the same before and after. Sessions already bound to such a descriptor through a checkpoint keep it, exactly as item 3 describes for the Quorum case.

6. The weights map is now the weighted electorate, and this one can break a stored session

Under voting.algorithm: "weighted", a declared participant absent from voting.weights used to weigh 1.0. It now weighs 0 and is non-decisive: its ballot contributes to neither side of the weighted ratio, does not enter the decisive tally, and does not satisfy the decline guard of RFC-MACP-0007 §6.2. (It still counts toward the voting.quorum participation floor -- that carve-out is explicit in RFC-MACP-0012 §4.1 and is unchanged.) The weights map is the electorate; an observer is expressed by omission, which is also why item 5 refuses an explicit 0.

This item is the only place in this release where a stored session's replay can change, and unlike item 4 it is not confined to a hand-built descriptor. The electorate rule is keyed on nothing -- RFC-MACP-0012 §4.1 makes it "normative for every schema version", so it reaches stored schema_version: 1 and 2 descriptors as well as version 3.

This item carries two independent changes, and the second one needs no weights map at all. Alongside the electorate rule above, the decline guard of RFC-MACP-0007 §6.2 is now applied on a Passed round as well as on Failed and NoVotes -- §6.2 says "the guard applies across all three voting results" and had said so before spec #99, while this runtime applied it in only one. So commitment.allow_decline_over_approval now waives the approval result and not the guard: a decline over a round with no decisive reject is denied whatever the knob says. That reaches ordinary majority policies that have never had a voting.weights key. Audit for both predicates below; neither subsumes the other.

Blast radius -- two independent predicates, either one sufficient.

Predicate A -- the weighted electorate. A stored session matches when both hold:

  1. the session is bound to a Decision policy whose voting.algorithm is weighted, and
  2. at least one accepted Vote was cast by a participant that does not appear as a key in that policy's voting.weights map, with a vote other than ABSTAIN.

A does not depend on the committed outcome's direction, on schema_version, or on allow_decline_over_approval.

Predicate B -- the decline guard on a passing round. A stored session matches when all three hold:

  1. the session is bound to a Decision policy that sets commitment.allow_decline_over_approval: true, and
  2. its accepted history contains a negative Commitment (outcome_positive: false), and
  3. the round that commitment sealed was Passed with zero decisive rejects -- in the ordinary case an all-APPROVE tally, since neither an abstention nor a missing ballot is a rejection.

B needs no weighted algorithm and no weights map. The minimal instance is fully registerable and trivially reachable: {"voting": {"algorithm": "majority", "threshold": 0.5}, "commitment": {"allow_decline_over_approval": true}} with three APPROVEs and a negative Commitment was Allow before this release and is Deny after.

Neither predicate is complete on its own, and A is not the safe one to check. An operator who audits only for weighted policies, finds none, and concludes the deployment is unaffected can be wrong -- B reaches ordinary majority, supermajority, unanimous and plurality policies. RFC-MACP-0012 §8's "Bounded exception -- weight-0 decisiveness" describes something narrower than either predicate (a decline over a Passed tally under allow_decline_over_approval: true whose only reject came from a weight-0 participant), because that is the only configuration the spec's earlier text had defined an outcome for. Two reasons that framing must not be read as this runtime's exposure. For A: this runtime had defined an outcome for every unlisted-voter configuration -- it defaulted the weight to 1.0 -- so A is wider than §8's corner in both directions. For B: §8 does not cover it at all. §8 is about weight-0 decisiveness, and B has no weights; B is a pre-existing conformance bug being fixed, not a semantics change §8 sanctioned, which is precisely why the bounded exception does not extend to it.

Worked example -- the positive direction, flipping from accepted to denied. This is the common shape, not an edge case. Weights {"agent://a": 1.0}, participants a, b, c; b and c cast APPROVE, a casts REJECT; voting.threshold: 0.5.

  • Before. Every voter weighed 1.0, so the total was 3.0, the approve share 2.0 / 3.0 = 0.667 >= 0.5, the voting result Passed, and a positive Commitment was accepted into history.
  • After. The total decisive weight is 1.0 (only a is in the electorate, and a rejected), the approve share is 0.0, the result is Failed, and the commitment is denied.

That session is authorable today with an ordinary, schema-valid, registered weighted policy, simply by omitting two participants from weights -- the majority-approves-but-the-weighted-voter-dissents shape, which is the most natural reason to choose weighted in the first place.

The failure mode is not "it replays differently" -- an affected session does not load. Replay dispatches the same commitment path as acceptance, so a Commitment the new rule denies becomes a POLICY_DENIED error out of replay_session rather than a different outcome. On the default recovery path the session is skipped with one WARN:

failed to replay session; skipping
  session_id=... error=...

and the session disappears from the registry on restart, silently apart from that line. Under MACP_STRICT_RECOVERY=1 the same error is fatal and the runtime refuses to start. Those are the two symptoms item 3 describes for the Quorum threshold guard, and the shapes are a pair -- read them together. One difference matters: validate_replay_consistency never runs for these sessions, because replay errors out before the consistency comparison is reached, so that check cannot be relied on to surface this.

Pre-upgrade audit query. Find affected sessions before upgrading. Run both predicates -- a deployment can match B without owning a single weighted policy:

  • For A: any Decision session whose bound policy has voting.algorithm: "weighted" and whose accepted Vote messages include a sender that does not appear as a key in that policy's voting.weights map, with a vote other than ABSTAIN.
  • For B: any Decision session whose bound policy sets commitment.allow_decline_over_approval: true and whose accepted history contains a negative Commitment sealed over a round carrying zero accepted REJECT ballots.

Sessions matching either may fail to replay after the upgrade.

How to run it -- the surfaces that carry each fact. Everything both predicates name is readable on the old binary, before you upgrade, over the ordinary RPC surface:

  • The session list and its mode. ListSessions enumerates current session metadata; page it with page_size and the returned next_page_token until the token comes back empty. GetSession returns one session's SessionMetadata, which carries its mode and the policy_version it bound.
  • The policy's rules. GetPolicy on that policy_version returns the PolicyDefinition; read voting.algorithm, the key set of voting.weights, and commitment.allow_decline_over_approval out of its rules. ListPolicies with a mode filter narrows the sweep to Decision policies first, which is usually the cheaper order -- a deployment with no matching policy needs no per-session pass at all.
  • The ballots and the committed outcome. StreamSession replays a session's accepted history: send a frame carrying subscribe_session_id with after_sequence: 0 and it emits the accepted envelopes in order, which is where the Vote senders and their values are, and where the Commitment's outcome_positive is. Note the one gap: a session whose log has been compacted answers with session history before ordinal N was compacted instead of the early entries, so treat a compaction response as "not audited" rather than "clean".
  • Offline, with the runtime down. The same accepted history is on disk under the file backend at <MACP_DATA_DIR>/sessions/<session_id>/log.jsonl, one JSON entry per line, filterable with jq without starting the runtime. (A SessionResume entry's banked_ms field changed meaning in this release: it now records the remaining TTL banked at suspend rather than the pause's duration -- entries written before this change carry the older quantity under the same field name, with no discriminator between the two.)

Contrast with item 4, which shipped ungated for a reason that does not transfer. Item 4's negative-weighted-total change was ungated because it was "reachable only from a directly-constructed PolicyDefinition, since registration already refuses negative weights". This one is reachable from an ordinary registered policy: a weighted descriptor that simply omits a declared participant from weights passes every admission check, before and after. So do not read item 6 as another item 4.

Why the change is nonetheless right: the old reading let a participant the policy author had explicitly given no voting weight cast ballots that moved the outcome -- in §8's corner, the unilateral power to convert an approving weighted electorate's result into a decline; more commonly, as above, the power to carry a positive round the electorate had rejected. §8 accepts the resulting stored-replay break in writing. The rule itself is in Policy.

Recovery for A. There is no configuration that restores the old reading -- the rule is ungated by design. Correct the policy instead: add to voting.weights, with the weight they should carry, the participants who were always meant to vote, and leave genuine observers omitted. A registered policy is re-resolved from the live registry on full replay, so correcting it there is enough for a session with no checkpoint; a session that carries the old descriptor in a checkpoint keeps that descriptor verbatim, exactly as item 3 describes for the Quorum case.

Be honest about the case the correction does not cover. If the flipped ballots really were cast by intended observers, no weight map makes that session replay: the commitment in its history was authorized by voters the policy had given no weight, which is precisely the outcome §8 calls unsound. The append-only log is untouched either way -- an affected session was not loaded rather than lost -- so the history remains available for audit while you decide.

Recovery for B: there is none, and that is the honest answer. No policy edit makes a B session replay. Clearing commitment.allow_decline_over_approval denies the decline earlier rather than later; setting it changes nothing, because it is the guard and not the knob that refuses. The reject the guard requires was never cast, so there is no descriptor under which that history is authorized -- RFC-MACP-0007 §6.2 held the same before spec #99, which is what makes this a bug fix rather than a semantics change. As with A the append-only log is untouched, so such a session is not loaded rather than lost and its history stays available for audit; but plan for it to stay unloaded, and under MACP_STRICT_RECOVERY=1 plan to clear it before the upgrade rather than after.

7. A Decision SessionStart may now bind an empty participants list

This one is a relaxation and cannot affect an existing deployment. SessionStart for macp.mode.decision.v1 is accepted with participants: [], where earlier releases refused it; every other standards-track mode still requires a non-empty roster. Nothing that used to be accepted is now refused, so no stored session's replay changes and no policy file needs correcting -- and no stored session can already contain an empty roster, because one could not have been accepted.

The resulting session is inert, which is worth knowing before an operator reads one in ListSessions and expects it to progress: with no declared participants, Proposal, Evaluation, Objection and Vote are all refused with FORBIDDEN (the initiator's included), so no proposal can exist and no Commitment can be sealed. Such a session can only expire or be cancelled, and the roster cannot be added to afterwards. See API for the full SessionStart contract and Modes for why Decision is the one carve-out.

8. A Quorum threshold.value of 0 no longer registers

RFC-MACP-0012 1.2.0-draft moved quorum threshold.value from minimum: 0 to exclusiveMinimum: 0 -- the quorum-side twin of the Decision-side floor this release already tightened (item 5). A zero approval bar is trivially satisfied, so a policy that reads as restrictive approved everything. {"threshold": {"type": "n_of_m", "value": 0}} and its percentage equivalent are now refused at registration, which means a MACP_POLICIES_DIR file carrying one aborts startup -- see items 1 and 2.

The floor is keyed on the value key being present. A rules object that omits threshold entirely, or supplies {"threshold": {}} or {"threshold": {"type": "percentage"}}, still registers and still leaves the rule inert, so the ApprovalRequest's own required_approvals stands. Nothing about the built-in policy.default wildcard changes.

9. A percentage Quorum threshold can now resolve one approval lower

RFC-MACP-0012 §4.2 promoted the percentage ceiling rule into normative text and pinned its arithmetic: the effective bar is ceil(value x declared_participant_count / 100), computed with exact integer arithmetic, and implementations "MUST NOT use floating-point division". This runtime already ceiled, but divided by 100 first, and that division is inexact in binary64 for most integer percentages -- the error survived into the ceiling and produced a bar one approval too high. value: 28 over 25 participants required 8 approvals where the rule gives 7; value: 7 over 100 gave 8 instead of 7; value: 68 over 75 gave 52 instead of 51. Thirteen (value, participants) pairs diverge within value 1-100 and participants 1-100, and the divergence grows more common at larger rosters.

Who this reaches. Only sessions bound to a policy whose threshold.type is percentage, and only those whose (value, participant count) pair is one of the affected ones. For them the bar drops by one, so a positive commitment that this runtime used to deny is now allowed at one fewer approval. No stored session's accepted history changes: a commitment the old bar denied was rejected, and rejected messages never enter accepted history (RFC-MACP-0001 §8.3), so replay is unaffected. What changes is the live bar and anything computed from it -- QuorumMode::effective_threshold_for_session, the 1..=participants domain guard on an ApprovalRequest, and the readiness the coordinator polls. RFC-MACP-0012 §8's completion note covers this explicitly: the earlier text stated no rounding direction, so it determined no outcome at a non-integral product, and a runtime that had privately chosen one is the only thing this can perturb.

The denominator is also pinned: it is the participant count declared at SessionStart and does not shrink as ballots, abstentions included, are cast. This runtime already read it that way.

10. An objection-authorized decline is no longer denied by require_vote_quorum or the evaluation gate

RFC-MACP-0007 §6.2 exempts an objection-authorized decline -- a negative Commitment under objection_handling.critical_objection_action: "finalize_decline" with a standing critical Objection -- from the voting tri-state and the decline guard. The guard is a two-conjunct conjunction whose second conjunct is commitment.require_vote_quorum, and §6.2 did not say whether the waiver reached it. This runtime raised that as spec issue #117 and, pending the ruling, kept the quorum gate and the evaluation.* prerequisites applying. Spec PR #126 ruled the waiver covers the guard whole: the quorum condition legitimizes an outcome deriving its authority from the voting result, an objection-authorized decline derives none, so a runtime "MUST NOT deny it for an unmet voting quorum", and evaluation.required_before_voting / evaluation.minimum_confidence "are prerequisites of the same voting pipeline and likewise MUST NOT be applied".

Who this reaches. Only sessions bound to a Decision policy that sets critical_objection_action: "finalize_decline" and at least one of commitment.require_vote_quorum: true or evaluation.required_before_voting: true, and only once a critical objection is standing. For them a negative Commitment that this runtime used to deny with POLICY_DENIED is now allowed. Nothing moves from allowed to denied, and the positive direction is untouched -- a positive commitment under the same veto is still denied, and still reports the unmet quorum and evaluation prerequisite among its reasons.

No stored session's replay changes. The commitment the old gates denied was rejected, and rejected messages never enter accepted history (RFC-MACP-0001 §8.3), so no stored history can contain one. §6.2 states this explicitly for this rule. What changes is live acceptance: a finalize_decline session that was previously reachable only by TTL expiry or an initiator CancelSession can now record a committed negative outcome, so an operator who was working around the strand -- leaving require_vote_quorum false, or soliciting a throwaway ABSTAIN to clear the participation floor -- can drop the workaround. See Policy for the full rule and tests/conformance/decision_finalize_decline_quorum_waiver.json for the canonical discriminator.

Upgrading into suspension-interval recording

1. Expect one-time suspension_intervals mismatch replay warnings on the first boot

A session now records each completed (suspended_at, resumed_at) pair, so that time spent Suspended can be subtracted from the Handoff implicit-accept timeout (RFC-MACP-0010 §5.1). On the first boot after upgrading, startup recovery emits one warning per persisted session that was ever suspended and resumed:

WARN replay/snapshot suspension_intervals mismatch
  session_id=... replayed_suspension_cycles=2 snapshot_suspension_cycles=0

It is benign, and it does not recur. The warning is the startup consistency check comparing two sources that necessarily disagree exactly once: the snapshot on disk was written by the previous release, which had no such field, so it deserializes as empty; replay rebuilds the pairs correctly from the SessionSuspend/SessionResume entries in the append-only log, which were always there. Replay is the authority -- the recovered session is the correct one, and it is what the runtime serves. Recovery then re-saves each replayed session, so the next boot's snapshot carries the intervals and the comparison agrees. Nothing is lost and no action is required.

Two consequences worth knowing:

  • The warning is advisory only. It increments recovery_replay_mismatches, which is read zero-vs-nonzero, so a nonzero count on this one boot is expected and does not indicate a determinism bug. MACP_STRICT_RECOVERY does not turn these into startup failures -- it governs recovery errors, not consistency warnings.
  • A session whose recovery goes through a mid-session checkpoint written before this release replays with its pre-checkpoint pauses missing for good: the checkpoint fast path replays only the entries after the checkpoint, and a legacy checkpoint carries no intervals. That is deliberately the safe direction -- a short interval list can only make a newly computed implicit-accept deadline land earlier, never later, and it never changes a deadline already recorded in accepted history. There is no setting that avoids this for a checkpoint already on disk: MACP_CHECKPOINT_INTERVAL gates only whether new checkpoints are written, while the recovery fast path keys on a Checkpoint entry being present in the log and never consults the variable. Leaving it at 0 (the default) prevents future legacy-shaped gaps; it does not undo an existing one. The condition is also self-limiting -- every pause recorded from this release forward is in the log after the old checkpoint, so it replays normally.

Environment variables

VariableDefaultDescription
MACP_BIND_ADDR127.0.0.1:50051gRPC listen address
MACP_TLS_CERT_PATH--TLS certificate PEM (required unless insecure)
MACP_TLS_KEY_PATH--TLS private key PEM (required unless insecure)
MACP_AUTH_TOKENS_FILE--Path to bearer token configuration file
MACP_AUTH_TOKENS_JSON--Inline bearer token config as JSON string
MACP_DATA_DIR.macp-dataDirectory for session persistence
MACP_STORAGE_BACKENDfileBackend: file, rocksdb, redis
MACP_ROCKSDB_PATH.macp-data/rocksdbRocksDB database path
MACP_REDIS_URLredis://127.0.0.1:6379Redis connection URL
MACP_MEMORY_ONLYoffSet to 1 to disable persistence entirely
MACP_ALLOW_INSECUREoffAllow plaintext connections (development only)
MACP_AUTH_ISSUER--JWT resolver expected iss claim (enables JWT auth)
MACP_AUTH_AUDIENCEmacp-runtimeJWT resolver expected aud claim
MACP_AUTH_JWKS_JSON--Inline JWKS document (JSON) for JWT validation
MACP_AUTH_JWKS_URL--JWKS endpoint URL (fetched + cached)
MACP_AUTH_JWKS_TTL_SECS300JWKS cache TTL when fetched from URL
MACP_AUTH_JWT_ALGSRS256,ES256Comma-separated JWT algorithm allowlist (HS256 requires explicit opt-in)
MACP_MAX_PAYLOAD_BYTES1048576Maximum envelope payload size in bytes
MACP_SESSION_START_LIMIT_PER_MINUTE60Per-sender session creation rate limit
MACP_MESSAGE_LIMIT_PER_MINUTE600Per-sender message rate limit
MACP_LIST_SESSIONS_DEFAULT_PAGE_SIZE100ListSessions page size when the request sends page_size = 0
MACP_LIST_SESSIONS_MAX_PAGE_SIZE1000Hard cap a requested ListSessions page_size is clamped to
MACP_METRICS_ADDR-- (off)Prometheus text endpoint bind address, e.g. 127.0.0.1:9464 (src/main.rs:584)
MACP_CONCURRENCY_LIMIT_PER_CONNECTION64tonic per-connection concurrency limit (src/main.rs:456)
MACP_MAX_CONCURRENT_STREAMS128HTTP/2 max concurrent streams (src/main.rs:460)
MACP_REQUEST_TIMEOUT_SECS30Per-request timeout (src/main.rs:464)
MACP_SHUTDOWN_DRAIN_SECS10Graceful-shutdown drain deadline (src/main.rs:524)
MACP_CHECKPOINT_INTERVAL0 (disabled)Log entries between checkpoints
MACP_CLEANUP_INTERVAL_SECS60Background maintenance interval in seconds: TTL expiry, memory eviction, disk GC, and eager observation of mode-computed deadlines (the handoff implicit accept, RFC-MACP-0010 §5.1(2))
MACP_SESSION_RETENTION_SECS3600Age (from session start) at which terminal sessions are evicted from memory; their durable data is kept
MACP_SESSION_DISK_RETENTION_SECS0 (keep forever)Age (from session start) at which terminal sessions' durable data is deleted; 0 disables disk GC entirely
MACP_STRICT_RECOVERYoffSet to 1 to fail on any recovery error
MACP_POLICIES_DIR--Directory of governance policy JSON files preloaded at startup; a file that fails validation aborts startup, and the wire registry becomes read-only
MACP_POLICIES_DRY_RUNoffSet to 1 to validate MACP_POLICIES_DIR and exit 0/1 without starting the server
MACP_POLICY_SCHEMAS_DIR--Development and CI only; the server never reads it. Path to the spec repository's schemas/json/policy directory, used by the enum_lists_match_the_canonical_schemas parity test -- see the warning below
RUST_LOGinfoLog level filter

MACP_POLICY_SCHEMAS_DIR is listed here because it is otherwise documented nowhere, and a contributor changing a registration mirror needs it. It is read only by macp-policy's parity unit test, which asserts the hand-written value-domain mirrors in crates/macp-policy/src/registry.rs still match the canonical schemas. Two warnings:

  • Point it at a clean git archive export of the spec commit CI reads, never at a sibling working tree. CI checks the spec repo out at SPEC_REV (.github/workflows/ci.yml), so a sibling checkout that is dirty, or ahead of or behind that pin, produces parity failures that do not exist in CI -- and it can move under you mid-session. Export first: git -C <spec-repo> archive <sha> schemas/ | tar -x -C <tmpdir>, then point the variable at <tmpdir>/schemas/json/policy.
  • A set-but-missing directory panics by design. Setting the variable asserts the canonical schemas are available, so the test refuses to skip silently. Unset it to fall back to a sibling checkout, or to skip the parity check entirely when no checkout exists.

Governance policy files

Validate a policies directory before you roll it out: MACP_POLICIES_DRY_RUN=1 MACP_POLICIES_DIR=/etc/macp/policies macp-runtime reports every file by name and exits 0/1 without starting the server. See Policy.

When a rejected policy file blocks startup, correct the file — do not delete it. Deleting it lets the runtime boot but silently voids governance for every in-flight session bound to that policy_version: the policy resolves to nothing on replay and commitment enforcement then treats the session as having no policy at all. UnregisterPolicy on a policy live sessions are still bound to does the same. The full mechanism, and the four other operational changes in this release, are in Upgrading into registration-time policy validation.

Storage backends

The runtime supports four storage configurations, selected via MACP_STORAGE_BACKEND:

File backend (default) stores each session in its own directory under MACP_DATA_DIR/sessions/<session_id>/. An append-only log.jsonl records every accepted message, and a session.json snapshot is written on each state change. Writes use an atomic tmp-file-then-rename pattern to prevent partial-write corruption.

RocksDB backend uses an embedded key-value store for higher throughput. Enable it by building with the rocksdb-backend Cargo feature and setting MACP_STORAGE_BACKEND=rocksdb. The database path defaults to MACP_ROCKSDB_PATH.

Redis backend stores session data in a remote Redis instance. Enable it with the redis-backend feature and set MACP_STORAGE_BACKEND=redis with MACP_REDIS_URL pointing to your Redis instance.

Durability matrix

The runtime acknowledges a message only after the log append "commit point". What that acknowledgement guarantees differs by backend:

BackendAcked ⇒ survives process crashAcked ⇒ survives host power lossNotes
fileyesyeslog appends fsync before ack; snapshots are tmp-fsync-rename atomic
rocksdbyesyeslog appends sync the WAL before ack; session snapshots are async (recovered via replay)
redisyes (Redis process survives)noRPUSH acks in Redis memory; no WAIT/AOF barrier. Cache-tier / single-writer only — the runtime logs a warning at startup

All backends skip individually corrupt log entries on load (with a warning) rather than failing the whole session; MACP_STRICT_RECOVERY=1 makes recovery-level errors fatal.

Single-writer: state authority lives in the runtime process's memory. Two runtimes must never share one data directory or Redis instance.

Observation-surface authorization

GetSession is participant/observer-scoped, while ListSessions and WatchSessions return metadata for all sessions to any authenticated identity (RFC-0006 permits this shape). Deployments with confidentiality requirements between agent groups should front these RPCs with a proxy or restrict which identities may call them. WatchSignals requires authentication; ListModes/GetManifest/WatchModeRegistry/WatchRoots are open discovery surfaces by design. ListSessions is now paged, and the decision not to sign its opaque page_token rests on exactly the unfiltered property described here -- a forged cursor can only reposition a caller within a listing it may already read in full, so if per-caller filtering is ever added to ListSessions, that no-signature decision must be re-analyzed first.

Memory-only mode disables persistence entirely. Set MACP_MEMORY_ONLY=1 for testing or ephemeral workloads. All session data is lost when the process exits.

Authentication

The runtime applies a pluggable resolver chain assembled at startup:

  1. JWT bearer (active when MACP_AUTH_ISSUER is set) -- validates signature, issuer, audience, and expiration against a JWKS. Default algorithm allowlist: RS256, ES256; HS256 (shared-secret) requires explicit opt-in via MACP_AUTH_JWT_ALGS=HS256. The sub claim becomes the sender; an optional macp_scopes claim carries capability flags (allowed_modes, can_start_sessions, max_open_sessions, can_manage_mode_registry, is_observer).
  2. Static bearer (active when MACP_AUTH_TOKENS_FILE or MACP_AUTH_TOKENS_JSON is set) -- looks up opaque tokens in a preloaded identity map. Accepts Authorization: Bearer <token> or the alternate x-macp-token: <token> header.
  3. Dev-mode fallback -- activates only when neither JWT nor static bearer is configured. Any Authorization: Bearer <value> header authenticates the caller as sender <value> with full capabilities. Intended strictly for local development.

If a credential matches a resolver but fails verification (expired JWT, unknown static token), the request is rejected with UNAUTHENTICATED -- the chain does not fall through to a later resolver.

Crash recovery

When persistence is enabled, the runtime rebuilds all sessions from their append-only logs on startup. This process is fully automatic:

  • Each session's log.jsonl is replayed through the mode engine to reconstruct the session state.
  • If a checkpoint exists, replay starts from the checkpoint and only processes subsequent entries.
  • Temporary files (.tmp suffixes) left by interrupted atomic writes are cleaned up.
  • The number of recovered sessions is logged at startup.

Log append failures are treated as fatal: the runtime rejects the message rather than acknowledging it without a durable record. This ensures the log is always the authoritative source of truth.

If MACP_STRICT_RECOVERY=1 is set, the runtime exits on any recovery error. Without it, individual session recovery failures are logged as warnings and the remaining sessions are loaded normally.

Monitoring

The runtime provides operational visibility through several mechanisms:

Logging -- All significant events are logged to stderr: session creation, resolution, expiration, recovery results, persistence failures, and rate limit hits. Set RUST_LOG to debug for detailed request-level logging.

TTL enforcement -- Sessions are expired both lazily (on next access) and proactively by a background task running every MACP_CLEANUP_INTERVAL_SECS. This ensures expired sessions are cleaned up even if no new messages arrive.

Mode deadlines -- The same background task observes deadlines a mode computes, on the same interval. The one in the standards-track modes today is the handoff implicit accept (RFC-MACP-0010 §5.1(2)): an outstanding HandoffOffer past its acceptance.implicit_accept_timeout_ms is accepted by the runtime, which appends the HandoffAccept to accepted history and publishes it to StreamSession subscribers. MACP_CLEANUP_INTERVAL_SECS is therefore the observation-latency bound on that acceptance -- but only the latency: the recorded timestamp_unix_ms is the computed deadline whichever path emits it, so changing the interval never changes permanent history. The deadline is also observed on demand, ahead of the next session-scoped message, so no message is ever evaluated against a stale offer regardless of the interval. The sweep runs after TTL expiry within a pass: a session whose TTL and implicit-accept deadline both lapsed unobserved expires rather than accepting, matching the on-demand path's precedence.

Session eviction -- Terminal sessions (resolved, expired, or cancelled) are evicted from memory once their age exceeds MACP_SESSION_RETENTION_SECS (default one hour), measured from session start rather than from when they became terminal. This is on by default and bounds memory usage. Their data remains on disk and can be replayed if needed. Deleting that durable data is a separate, opt-in step governed by MACP_SESSION_DISK_RETENTION_SECS, which defaults to 0 -- disk GC does not run at all unless you set it.

Log compaction -- When a session reaches a terminal state, the runtime automatically compacts its log into a single checkpoint entry. This reduces storage footprint for completed sessions.

gRPC reflection (dev/debug only) -- Build with the reflection Cargo feature (cargo build --features reflection) to enable the standard gRPC Server Reflection API, so tools like grpcurl/grpcui can call RPCs without a local .proto/.protoset file (e.g. grpcurl -plaintext localhost:50051 list). Off by default and never enabled in the published Docker image -- it exposes the full service/message schema to any client that can reach the port, so treat it the same as MACP_ALLOW_INSECURE: fine for a local instance or an ephemeral CI container, not for a production deployment. The runtime logs a warning at startup when it's active. Without this feature, point grpcurl -proto/-protoset at the published proto instead (@multiagentcoordinationprotocol/proto on npm, macp-proto on crates.io).

Container deployment

Use the repository's Dockerfile (multi-stage, non-root macp user) to build your own image, or pull the image this repo publishes to GitHub Container Registry: ghcr.io/multiagentcoordinationprotocol/macp-runtime. The image does not enable dev mode: the runtime refuses to start without configured authentication and TLS unless you explicitly opt in for local development:

# Production: configure auth + TLS. Pin an immutable :X.Y.Z tag (see below) --
# substitute the version you're deploying.
docker run -p 50051:50051 \
  -e MACP_AUTH_TOKENS_FILE=/etc/macp/tokens.json \
  -e MACP_TLS_CERT_PATH=/etc/macp/tls.crt -e MACP_TLS_KEY_PATH=/etc/macp/tls.key \
  -v ./secrets:/etc/macp ghcr.io/multiagentcoordinationprotocol/macp-runtime:X.Y.Z

# Local development ONLY: any bearer token becomes a fully-privileged identity
docker run -p 50051:50051 -e MACP_ALLOW_INSECURE=1 ghcr.io/multiagentcoordinationprotocol/macp-runtime:X.Y.Z

When deploying in containers:

  • Mount a persistent volume at MACP_DATA_DIR so session logs survive container restarts.
  • Expose port 50051 (or the port configured via MACP_BIND_ADDR).
  • Provide TLS certificates and auth tokens via mounted secrets.
  • Set MACP_BIND_ADDR=0.0.0.0:50051 to accept connections from outside the container.

Published image tags

Every image is multi-arch (linux/amd64, linux/arm64) with SLSA provenance and an SBOM attached, built by .github/workflows/docker.yml. Every push to main (a "branch build") publishes three of these together, always pointing at the same digest; a release publishes the other two, also always sharing a digest with each other but never with a branch build's:

TagMutabilityMeaning
:X.Y.Z (e.g. 0.8.1)ImmutableExactly one release. Pin this for production.
:X.Y (e.g. 0.8)Moves within the minorAlways the newest X.Y.Z patch published so far.
:latestMoves on every push to mainThe tip of main — not the newest release. Do not use latest in production; it is not a release channel.
:mainMoves on every push to mainSynonymous with :latest today (both are branch-build tags on the one push: branches: [main] trigger) — documented separately since nothing in docker.yml guarantees they stay synonymous if a second branch trigger is ever added.
:<sha> (e.g. :bb45705)ImmutableOne image per commit pushed to main (branch builds only — a release build never gets a SHA tag; see below).

A release commit is main's tip at the moment it's cut, but main moves on immediately afterward as later commits land — so by the time you read this, latest/main/:<sha> are very likely already ahead of every semver tag. Never assume latest matches the newest release.

A :X.Y.Z tag and its release commit's :<sha> tag do not share a digest, even though they're built from the same source tree: the release build and the branch build for that commit run as two separate, concurrent CI jobs, and each stamps its own org.opencontainers.image.created timestamp. Same source, two images — don't expect docker inspect to report identical digests across them.

Every image's org.opencontainers.image.revision label names the exact commit it was built from — this is reliable even for a manually backfilled release image (a workflow_dispatch naming an older tag), where org.opencontainers.image.revision still correctly reports the tag's commit. The one residual gap on a manual backfill: buildkit's SLSA provenance attestation records the runner's own GITHUB_SHA (the dispatch commit, i.e. whatever main was at dispatch time), not the backfilled tag's commit — only the revision label is authoritative there. Every image published by the normal, automatic release path (cutting a release via release-plz) has correct provenance with no such gap.

Do not pin the macp-runtime-v* git tag format as an image tag — GHCR image tags drop the macp-runtime-v prefix (macp-runtime-v0.8.1 → image tag 0.8.1). docker.yml's resolve step strips it deliberately, to emit a bare semver value for docker/metadata-action's {{version}}/ {{major}}.{{minor}} patterns — the prefix is redundant once the image already lives in the macp-runtime repository path, not a Docker tag syntax restriction (macp-runtime-v0.8.1 would itself be a perfectly valid tag string).

Development tools

For development and CI, these additional tools are useful:

  • cargo-tarpaulin for coverage reporting
  • cargo-audit for dependency security auditing
  • buf for protocol buffer linting and management