Deployment Guide
This guide covers everything you need to run the MACP Runtime in production: configuration, storage backends, crash recovery, monitoring, and container deployment. For protocol-level deployment topologies and security requirements, see the protocol deployment and protocol security documentation.
Production checklist
Before exposing the runtime to production traffic, ensure these four items are configured:
-
TLS certificates -- Set
MACP_TLS_CERT_PATHandMACP_TLS_KEY_PATHto valid PEM files. The runtime refuses to start without TLS unlessMACP_ALLOW_INSECURE=1is set. -
Authentication -- Configure at least one of the resolvers. For opaque bearer tokens, create a
tokens.jsonmapping tokens to agent identities and setMACP_AUTH_TOKENS_FILE. For JWT bearer tokens, setMACP_AUTH_ISSUERtogether with a JWKS source (MACP_AUTH_JWKS_JSONinline orMACP_AUTH_JWKS_URLfetched + cached). Both can be configured at once -- JWT-shaped tokens are routed to the JWT resolver and opaque tokens to the static resolver. See the Getting Started guide for the token format and JWT claim layout. -
Data directory -- Ensure
MACP_DATA_DIRpoints to a directory with write permissions. This is where session logs and snapshots are stored. -
Bind address -- Set
MACP_BIND_ADDRto the desired listen address. The default127.0.0.1:50051only accepts local connections.
Upgrading into registration-time policy validation
This release tightens what the governance policy registry accepts, what the Quorum mode will bind, how the Decision evaluator reads a weighted round, when a decline may be finalized over a passing vote, and how a percentage Quorum threshold is computed. Ten changes are operationally visible -- seven tightenings, one relaxation (item 7, which cannot affect an existing deployment) and two corrections that lower a bar (items 9 and 10) -- and item 6 is the only one that can change how an already-stored session replays -- read it first if you have any persisted Decision session at all. Item 6 carries two independent predicates and only one of them involves weighted: the other reaches any Decision policy that sets commitment.allow_decline_over_approval: true, majority policies included. Read this section before upgrading any deployment that sets MACP_POLICIES_DIR, or that has persisted sessions bound to a policy with a Quorum threshold, a weighted voting.algorithm, or commitment.allow_decline_over_approval: true. CHANGELOG.md is generated from commit subjects and does not carry this detail.
1. An invalid policy file now refuses startup
Both routes into the registry -- the RegisterPolicy RPC and the MACP_POLICIES_DIR preload -- gained value-domain and conditional checks; the enforced set is listed in Policy. A file an earlier release accepted may now be out of domain: a fractional or zero Quorum threshold.value, threshold.type: "weighted", an unknown voting.algorithm or voting.quorum.type, a weighted algorithm with an empty weights map, a supermajority threshold at or below 0.5 (including one that omits threshold entirely and relies on the 0.5 default), or a wildcard ("*") policy carrying a Quorum threshold that was previously validated against the Decision schema alone and therefore never checked. Loading stops at the first rejection and startup aborts -- the preload error is propagated, not logged and skipped.
That is deliberate fail-closed behaviour, and it is why MACP_POLICIES_DRY_RUN=1 exists. Run the dry run with the new binary before you upgrade:
MACP_POLICIES_DRY_RUN=1 MACP_POLICIES_DIR=/etc/macp/policies macp-runtimeIt reports every *.json file by name as OK <path> (<policy_id>) or REJECTED <path>: <reason>, prints a checked/rejected count, and exits 0 if the directory would load or 1 if anything in it would be rejected. It binds no port, opens no storage, and replays nothing, so it needs neither TLS nor MACP_ALLOW_INSECURE=1. Its output is written to stdout/stderr directly rather than through tracing, so no RUST_LOG filter can suppress the report. Two things to know: the startup environment-configuration check still runs ahead of it, so an unrelated malformed variable aborts before the report is produced; and a readable directory containing no *.json exits 0 with an explicit WARNING, because a mis-pointed MACP_POLICIES_DIR otherwise looks identical to a clean run.
2. Do not unblock startup by deleting the rejected file
When a policy file blocks startup, the natural fix is to delete it. Correct the file instead. Deleting it does let the runtime boot, but persisted sessions bound to that policy_version are then replayed with the policy unresolved: replay resolves the version best-effort and leaves policy_definition empty when it cannot, and commitment enforcement treats an absent policy definition as "no policy to enforce" and returns early. Every in-flight session governed by the deleted policy therefore loses its governance silently -- commitments the policy would have denied are accepted, with no error and no log line tying it back to the deletion. The same applies to UnregisterPolicy on a policy that live sessions are still bound to.
Note the interaction with the next item: because an unresolved policy makes the Quorum mode fall back to the ApprovalRequest's own required_approvals -- a value the mode already constrains to 1..=participants -- deleting the policy also makes the replay failure below disappear. The two symptoms clear together, and the reason they clear is that the governance bar is no longer being applied.
3. A persisted Quorum session with an out-of-domain policy threshold no longer replays
RFC-MACP-0011 §5 rule 6 makes a policy threshold replace the ApprovalRequest's required_approvals, but nothing previously held the replacement to the same 1..=participants domain the runtime enforces on the field it replaces. The Quorum mode now refuses an ApprovalRequest whose effective threshold falls outside that domain, and replay dispatches the same code -- so such a session fails to replay. It is skipped with a warning, or is fatal at startup under MACP_STRICT_RECOVERY=1.
Detect it from the logs. The mode emits, at WARN:
quorum policy threshold is outside 1..=participants; refusing the ApprovalRequest
session_id=... policy_id=... effective_threshold=... participants=...naming the session, the bound policy, the computed threshold and the declared participant count -- everything needed to identify which policy to correct. Recovery follows it with failed to replay session; skipping carrying the same session_id.
What it takes to reach this. Not a legacy policy -- an ApprovalRequest this runtime accepted before the guard above existed. Registration is no substitute for the guard, because registration has no participant count to bound the threshold against: {"type": "n_of_m", "value": 66} passes every check in Policy under the new binary and still trips the guard on a three-participant session. Two classes of threshold reach it, and they are not equally benign:
- An out-of-domain numeric threshold --
n_of_morcountabove the declared participant count. The positive outcome was unreachable from the first message, and before this releasecommitment_readycarried nocounted > 0guard, soapprovals + remaining < requiredheld with no ballot cast and the coordinator could seal a bindingquorum.rejectedwith zero approvals (issue #145). That is the condition RFC-MACP-0011 §5 rule 4a reads as grounds for a decline, reached without a vote. Such a session really was broken: the only outcome it could ever have sealed was a decline nobody cast a ballot for. threshold.type: "weighted", or any unrecognised type. This class was working, and it stops replaying. The old shared fallback arm read any unrecognised type as a raw approval count (_ => rules.threshold.value as u32), so{"type": "weighted", "value": 2}on three participants was a perfectly satisfiable bar of two approvals, and sessions under it sealed legitimate positive commitments. The type now resolves toUnsatisfiable, theApprovalRequestis refused, and the session no longer loads. Do not read the warning as a report of a session that was already dead.
The second class survives the upgrade through a checkpoint, not through the registry. A checkpoint serializes the resolved policy_definition inline and try_replay_from_checkpoint restores it verbatim without consulting the registry, so an old weighted definition is still live even though neither RegisterPolicy nor the MACP_POLICIES_DIR preload would accept it again. A session with no checkpoint re-resolves its policy_version against the live registry during full replay, and there a weighted policy file aborts startup at item 1 before recovery ever runs.
Recovery, for both classes: correct the threshold, do not delete the policy. Restate a weighted or unrecognised type as n_of_m, keeping the same value. QuorumThreshold::effective treats n_of_m as a raw approval count, which is exactly what the old fallback arm did, so the bar that session enforced is preserved. (If the old value was fractional, registration now refuses it; the old arm truncated, so its floor is the faithful integer.) For a numeric threshold above the participant count, bring it into 1..=participants -- no value reproduces that session's old behaviour, because its old behaviour was the zero-approval decline. Then restart: the append-only log is untouched, so the session was not loaded rather than lost, and it replays normally. Deleting the policy also clears the warning, but for the reason item 2 gives -- the governance bar stops being applied at all.
4. A negative weighted total now fails the Decision round
A weighted round whose cast weights sum below zero fails the round instead of computing a ratio over a negative denominator. In the approve direction this is a tightening: a round that previously reported Passed through an inverted ratio >= threshold comparison is now denied. In the decline direction it is not a tightening -- on that same round a negative commitment moves from denied to allowed, because a decline over Passed was refused while a decline over Failed is permitted once the universal reject-floor is satisfied. The case is reachable only from a directly-constructed PolicyDefinition, since registration already refuses negative weights. A weighted total of exactly zero is unchanged.
5. voting.threshold: 0.0 and zero voting.weights entries are no longer accepted
Spec #99 moved two Decision bounds in decision-rules.schema.json from inclusive to exclusive at zero -- voting.threshold to exclusiveMinimum: 0, and voting.weights.additionalProperties to exclusiveMinimum: 0 with minProperties: 1 on the map. This runtime mirrors both, and adds the schema's majority arm: a majority threshold below 0.5 is refused, where supermajority continues to require one strictly above 0.5. The asymmetry is deliberate -- the reserved policy.std.majority profile sets exactly 0.5.
A policy file an earlier release accepted may now be refused, and because the MACP_POLICIES_DIR preload aborts startup at the first rejection, a deployment carrying any of these on disk will fail to start:
voting.threshold: 0.0(it made an all-REJECTround returnPassedunder bothmajorityandweighted)- a
voting.weightsentry of0.0, or a supplied but emptyvoting.weights: {} - a
majorityvoting.thresholdbelow0.5
Run the dry run with the new binary before you upgrade -- it is the same pre-upgrade check item 1 describes:
MACP_POLICIES_DRY_RUN=1 MACP_POLICIES_DIR=/etc/macp/policies macp-runtimeCorrecting a zero weight is not a matter of picking a small positive number. The weights map is the weighted electorate: a participant who should carry no voting weight is expressed by omission from the map, never by an explicit 0. Remove the entry rather than nudging it above zero. A map that would be left empty means no weighted electorate at all, which the weighted algorithm cannot express -- choose a different algorithm.
This item affects admission only; no stored session's replay changes, because a descriptor carrying any of these values evaluated the same before and after. Sessions already bound to such a descriptor through a checkpoint keep it, exactly as item 3 describes for the Quorum case.
6. The weights map is now the weighted electorate, and this one can break a stored session
Under voting.algorithm: "weighted", a declared participant absent from voting.weights used to weigh 1.0. It now weighs 0 and is non-decisive: its ballot contributes to neither side of the weighted ratio, does not enter the decisive tally, and does not satisfy the decline guard of RFC-MACP-0007 §6.2. (It still counts toward the voting.quorum participation floor -- that carve-out is explicit in RFC-MACP-0012 §4.1 and is unchanged.) The weights map is the electorate; an observer is expressed by omission, which is also why item 5 refuses an explicit 0.
This item is the only place in this release where a stored session's replay can change, and unlike item 4 it is not confined to a hand-built descriptor. The electorate rule is keyed on nothing -- RFC-MACP-0012 §4.1 makes it "normative for every schema version", so it reaches stored schema_version: 1 and 2 descriptors as well as version 3.
This item carries two independent changes, and the second one needs no weights map at all. Alongside the electorate rule above, the decline guard of RFC-MACP-0007 §6.2 is now applied on a Passed round as well as on Failed and NoVotes -- §6.2 says "the guard applies across all three voting results" and had said so before spec #99, while this runtime applied it in only one. So commitment.allow_decline_over_approval now waives the approval result and not the guard: a decline over a round with no decisive reject is denied whatever the knob says. That reaches ordinary majority policies that have never had a voting.weights key. Audit for both predicates below; neither subsumes the other.
Blast radius -- two independent predicates, either one sufficient.
Predicate A -- the weighted electorate. A stored session matches when both hold:
- the session is bound to a Decision policy whose
voting.algorithmisweighted, and - at least one accepted
Votewas cast by a participant that does not appear as a key in that policy'svoting.weightsmap, with a vote other thanABSTAIN.
A does not depend on the committed outcome's direction, on schema_version, or on allow_decline_over_approval.
Predicate B -- the decline guard on a passing round. A stored session matches when all three hold:
- the session is bound to a Decision policy that sets
commitment.allow_decline_over_approval: true, and - its accepted history contains a negative
Commitment(outcome_positive: false), and - the round that commitment sealed was
Passedwith zero decisive rejects -- in the ordinary case an all-APPROVEtally, since neither an abstention nor a missing ballot is a rejection.
B needs no weighted algorithm and no weights map. The minimal instance is fully registerable and trivially reachable: {"voting": {"algorithm": "majority", "threshold": 0.5}, "commitment": {"allow_decline_over_approval": true}} with three APPROVEs and a negative Commitment was Allow before this release and is Deny after.
Neither predicate is complete on its own, and A is not the safe one to check. An operator who audits only for weighted policies, finds none, and concludes the deployment is unaffected can be wrong -- B reaches ordinary majority, supermajority, unanimous and plurality policies. RFC-MACP-0012 §8's "Bounded exception -- weight-0 decisiveness" describes something narrower than either predicate (a decline over a Passed tally under allow_decline_over_approval: true whose only reject came from a weight-0 participant), because that is the only configuration the spec's earlier text had defined an outcome for. Two reasons that framing must not be read as this runtime's exposure. For A: this runtime had defined an outcome for every unlisted-voter configuration -- it defaulted the weight to 1.0 -- so A is wider than §8's corner in both directions. For B: §8 does not cover it at all. §8 is about weight-0 decisiveness, and B has no weights; B is a pre-existing conformance bug being fixed, not a semantics change §8 sanctioned, which is precisely why the bounded exception does not extend to it.
Worked example -- the positive direction, flipping from accepted to denied. This is the common shape, not an edge case. Weights {"agent://a": 1.0}, participants a, b, c; b and c cast APPROVE, a casts REJECT; voting.threshold: 0.5.
- Before. Every voter weighed
1.0, so the total was3.0, the approve share2.0 / 3.0 = 0.667 >= 0.5, the voting resultPassed, and a positiveCommitmentwas accepted into history. - After. The total decisive weight is
1.0(onlyais in the electorate, andarejected), the approve share is0.0, the result isFailed, and the commitment is denied.
That session is authorable today with an ordinary, schema-valid, registered weighted policy, simply by omitting two participants from weights -- the majority-approves-but-the-weighted-voter-dissents shape, which is the most natural reason to choose weighted in the first place.
The failure mode is not "it replays differently" -- an affected session does not load. Replay dispatches the same commitment path as acceptance, so a Commitment the new rule denies becomes a POLICY_DENIED error out of replay_session rather than a different outcome. On the default recovery path the session is skipped with one WARN:
failed to replay session; skipping
session_id=... error=...and the session disappears from the registry on restart, silently apart from that line. Under MACP_STRICT_RECOVERY=1 the same error is fatal and the runtime refuses to start. Those are the two symptoms item 3 describes for the Quorum threshold guard, and the shapes are a pair -- read them together. One difference matters: validate_replay_consistency never runs for these sessions, because replay errors out before the consistency comparison is reached, so that check cannot be relied on to surface this.
Pre-upgrade audit query. Find affected sessions before upgrading. Run both predicates -- a deployment can match B without owning a single weighted policy:
- For A: any Decision session whose bound policy has
voting.algorithm: "weighted"and whose acceptedVotemessages include a sender that does not appear as a key in that policy'svoting.weightsmap, with a vote other thanABSTAIN. - For B: any Decision session whose bound policy sets
commitment.allow_decline_over_approval: trueand whose accepted history contains a negativeCommitmentsealed over a round carrying zero acceptedREJECTballots.
Sessions matching either may fail to replay after the upgrade.
How to run it -- the surfaces that carry each fact. Everything both predicates name is readable on the old binary, before you upgrade, over the ordinary RPC surface:
- The session list and its mode.
ListSessionsenumerates current session metadata; page it withpage_sizeand the returnednext_page_tokenuntil the token comes back empty.GetSessionreturns one session'sSessionMetadata, which carries itsmodeand thepolicy_versionit bound. - The policy's rules.
GetPolicyon thatpolicy_versionreturns thePolicyDefinition; readvoting.algorithm, the key set ofvoting.weights, andcommitment.allow_decline_over_approvalout of itsrules.ListPolicieswith a mode filter narrows the sweep to Decision policies first, which is usually the cheaper order -- a deployment with no matching policy needs no per-session pass at all. - The ballots and the committed outcome.
StreamSessionreplays a session's accepted history: send a frame carryingsubscribe_session_idwithafter_sequence: 0and it emits the accepted envelopes in order, which is where theVotesenders and their values are, and where theCommitment'soutcome_positiveis. Note the one gap: a session whose log has been compacted answers withsession history before ordinal N was compactedinstead of the early entries, so treat a compaction response as "not audited" rather than "clean". - Offline, with the runtime down. The same accepted history is on disk under the file backend at
<MACP_DATA_DIR>/sessions/<session_id>/log.jsonl, one JSON entry per line, filterable withjqwithout starting the runtime. (ASessionResumeentry'sbanked_msfield changed meaning in this release: it now records the remaining TTL banked at suspend rather than the pause's duration -- entries written before this change carry the older quantity under the same field name, with no discriminator between the two.)
Contrast with item 4, which shipped ungated for a reason that does not transfer. Item 4's negative-weighted-total change was ungated because it was "reachable only from a directly-constructed PolicyDefinition, since registration already refuses negative weights". This one is reachable from an ordinary registered policy: a weighted descriptor that simply omits a declared participant from weights passes every admission check, before and after. So do not read item 6 as another item 4.
Why the change is nonetheless right: the old reading let a participant the policy author had explicitly given no voting weight cast ballots that moved the outcome -- in §8's corner, the unilateral power to convert an approving weighted electorate's result into a decline; more commonly, as above, the power to carry a positive round the electorate had rejected. §8 accepts the resulting stored-replay break in writing. The rule itself is in Policy.
Recovery for A. There is no configuration that restores the old reading -- the rule is ungated by design. Correct the policy instead: add to voting.weights, with the weight they should carry, the participants who were always meant to vote, and leave genuine observers omitted. A registered policy is re-resolved from the live registry on full replay, so correcting it there is enough for a session with no checkpoint; a session that carries the old descriptor in a checkpoint keeps that descriptor verbatim, exactly as item 3 describes for the Quorum case.
Be honest about the case the correction does not cover. If the flipped ballots really were cast by intended observers, no weight map makes that session replay: the commitment in its history was authorized by voters the policy had given no weight, which is precisely the outcome §8 calls unsound. The append-only log is untouched either way -- an affected session was not loaded rather than lost -- so the history remains available for audit while you decide.
Recovery for B: there is none, and that is the honest answer. No policy edit makes a B session replay. Clearing commitment.allow_decline_over_approval denies the decline earlier rather than later; setting it changes nothing, because it is the guard and not the knob that refuses. The reject the guard requires was never cast, so there is no descriptor under which that history is authorized -- RFC-MACP-0007 §6.2 held the same before spec #99, which is what makes this a bug fix rather than a semantics change. As with A the append-only log is untouched, so such a session is not loaded rather than lost and its history stays available for audit; but plan for it to stay unloaded, and under MACP_STRICT_RECOVERY=1 plan to clear it before the upgrade rather than after.
7. A Decision SessionStart may now bind an empty participants list
This one is a relaxation and cannot affect an existing deployment. SessionStart for macp.mode.decision.v1 is accepted with participants: [], where earlier releases refused it; every other standards-track mode still requires a non-empty roster. Nothing that used to be accepted is now refused, so no stored session's replay changes and no policy file needs correcting -- and no stored session can already contain an empty roster, because one could not have been accepted.
The resulting session is inert, which is worth knowing before an operator reads one in ListSessions and expects it to progress: with no declared participants, Proposal, Evaluation, Objection and Vote are all refused with FORBIDDEN (the initiator's included), so no proposal can exist and no Commitment can be sealed. Such a session can only expire or be cancelled, and the roster cannot be added to afterwards. See API for the full SessionStart contract and Modes for why Decision is the one carve-out.
8. A Quorum threshold.value of 0 no longer registers
RFC-MACP-0012 1.2.0-draft moved quorum threshold.value from minimum: 0 to exclusiveMinimum: 0 -- the quorum-side twin of the Decision-side floor this release already tightened (item 5). A zero approval bar is trivially satisfied, so a policy that reads as restrictive approved everything. {"threshold": {"type": "n_of_m", "value": 0}} and its percentage equivalent are now refused at registration, which means a MACP_POLICIES_DIR file carrying one aborts startup -- see items 1 and 2.
The floor is keyed on the value key being present. A rules object that omits threshold entirely, or supplies {"threshold": {}} or {"threshold": {"type": "percentage"}}, still registers and still leaves the rule inert, so the ApprovalRequest's own required_approvals stands. Nothing about the built-in policy.default wildcard changes.
9. A percentage Quorum threshold can now resolve one approval lower
RFC-MACP-0012 §4.2 promoted the percentage ceiling rule into normative text and pinned its arithmetic: the effective bar is ceil(value x declared_participant_count / 100), computed with exact integer arithmetic, and implementations "MUST NOT use floating-point division". This runtime already ceiled, but divided by 100 first, and that division is inexact in binary64 for most integer percentages -- the error survived into the ceiling and produced a bar one approval too high. value: 28 over 25 participants required 8 approvals where the rule gives 7; value: 7 over 100 gave 8 instead of 7; value: 68 over 75 gave 52 instead of 51. Thirteen (value, participants) pairs diverge within value 1-100 and participants 1-100, and the divergence grows more common at larger rosters.
Who this reaches. Only sessions bound to a policy whose threshold.type is percentage, and only those whose (value, participant count) pair is one of the affected ones. For them the bar drops by one, so a positive commitment that this runtime used to deny is now allowed at one fewer approval. No stored session's accepted history changes: a commitment the old bar denied was rejected, and rejected messages never enter accepted history (RFC-MACP-0001 §8.3), so replay is unaffected. What changes is the live bar and anything computed from it -- QuorumMode::effective_threshold_for_session, the 1..=participants domain guard on an ApprovalRequest, and the readiness the coordinator polls. RFC-MACP-0012 §8's completion note covers this explicitly: the earlier text stated no rounding direction, so it determined no outcome at a non-integral product, and a runtime that had privately chosen one is the only thing this can perturb.
The denominator is also pinned: it is the participant count declared at SessionStart and does not shrink as ballots, abstentions included, are cast. This runtime already read it that way.
10. An objection-authorized decline is no longer denied by require_vote_quorum or the evaluation gate
RFC-MACP-0007 §6.2 exempts an objection-authorized decline -- a negative Commitment under objection_handling.critical_objection_action: "finalize_decline" with a standing critical Objection -- from the voting tri-state and the decline guard. The guard is a two-conjunct conjunction whose second conjunct is commitment.require_vote_quorum, and §6.2 did not say whether the waiver reached it. This runtime raised that as spec issue #117 and, pending the ruling, kept the quorum gate and the evaluation.* prerequisites applying. Spec PR #126 ruled the waiver covers the guard whole: the quorum condition legitimizes an outcome deriving its authority from the voting result, an objection-authorized decline derives none, so a runtime "MUST NOT deny it for an unmet voting quorum", and evaluation.required_before_voting / evaluation.minimum_confidence "are prerequisites of the same voting pipeline and likewise MUST NOT be applied".
Who this reaches. Only sessions bound to a Decision policy that sets critical_objection_action: "finalize_decline" and at least one of commitment.require_vote_quorum: true or evaluation.required_before_voting: true, and only once a critical objection is standing. For them a negative Commitment that this runtime used to deny with POLICY_DENIED is now allowed. Nothing moves from allowed to denied, and the positive direction is untouched -- a positive commitment under the same veto is still denied, and still reports the unmet quorum and evaluation prerequisite among its reasons.
No stored session's replay changes. The commitment the old gates denied was rejected, and rejected messages never enter accepted history (RFC-MACP-0001 §8.3), so no stored history can contain one. §6.2 states this explicitly for this rule. What changes is live acceptance: a finalize_decline session that was previously reachable only by TTL expiry or an initiator CancelSession can now record a committed negative outcome, so an operator who was working around the strand -- leaving require_vote_quorum false, or soliciting a throwaway ABSTAIN to clear the participation floor -- can drop the workaround. See Policy for the full rule and tests/conformance/decision_finalize_decline_quorum_waiver.json for the canonical discriminator.
Upgrading into suspension-interval recording
1. Expect one-time suspension_intervals mismatch replay warnings on the first boot
A session now records each completed (suspended_at, resumed_at) pair, so that time spent Suspended can be subtracted from the Handoff implicit-accept timeout (RFC-MACP-0010 §5.1). On the first boot after upgrading, startup recovery emits one warning per persisted session that was ever suspended and resumed:
WARN replay/snapshot suspension_intervals mismatch
session_id=... replayed_suspension_cycles=2 snapshot_suspension_cycles=0It is benign, and it does not recur. The warning is the startup consistency check comparing two sources that necessarily disagree exactly once: the snapshot on disk was written by the previous release, which had no such field, so it deserializes as empty; replay rebuilds the pairs correctly from the SessionSuspend/SessionResume entries in the append-only log, which were always there. Replay is the authority -- the recovered session is the correct one, and it is what the runtime serves. Recovery then re-saves each replayed session, so the next boot's snapshot carries the intervals and the comparison agrees. Nothing is lost and no action is required.
Two consequences worth knowing:
- The warning is advisory only. It increments
recovery_replay_mismatches, which is read zero-vs-nonzero, so a nonzero count on this one boot is expected and does not indicate a determinism bug.MACP_STRICT_RECOVERYdoes not turn these into startup failures -- it governs recovery errors, not consistency warnings. - A session whose recovery goes through a mid-session checkpoint written before this release replays with its pre-checkpoint pauses missing for good: the checkpoint fast path replays only the entries after the checkpoint, and a legacy checkpoint carries no intervals. That is deliberately the safe direction -- a short interval list can only make a newly computed implicit-accept deadline land earlier, never later, and it never changes a deadline already recorded in accepted history. There is no setting that avoids this for a checkpoint already on disk:
MACP_CHECKPOINT_INTERVALgates only whether new checkpoints are written, while the recovery fast path keys on aCheckpointentry being present in the log and never consults the variable. Leaving it at0(the default) prevents future legacy-shaped gaps; it does not undo an existing one. The condition is also self-limiting -- every pause recorded from this release forward is in the log after the old checkpoint, so it replays normally.
Environment variables
| Variable | Default | Description |
|---|---|---|
MACP_BIND_ADDR | 127.0.0.1:50051 | gRPC listen address |
MACP_TLS_CERT_PATH | -- | TLS certificate PEM (required unless insecure) |
MACP_TLS_KEY_PATH | -- | TLS private key PEM (required unless insecure) |
MACP_AUTH_TOKENS_FILE | -- | Path to bearer token configuration file |
MACP_AUTH_TOKENS_JSON | -- | Inline bearer token config as JSON string |
MACP_DATA_DIR | .macp-data | Directory for session persistence |
MACP_STORAGE_BACKEND | file | Backend: file, rocksdb, redis |
MACP_ROCKSDB_PATH | .macp-data/rocksdb | RocksDB database path |
MACP_REDIS_URL | redis://127.0.0.1:6379 | Redis connection URL |
MACP_MEMORY_ONLY | off | Set to 1 to disable persistence entirely |
MACP_ALLOW_INSECURE | off | Allow plaintext connections (development only) |
MACP_AUTH_ISSUER | -- | JWT resolver expected iss claim (enables JWT auth) |
MACP_AUTH_AUDIENCE | macp-runtime | JWT resolver expected aud claim |
MACP_AUTH_JWKS_JSON | -- | Inline JWKS document (JSON) for JWT validation |
MACP_AUTH_JWKS_URL | -- | JWKS endpoint URL (fetched + cached) |
MACP_AUTH_JWKS_TTL_SECS | 300 | JWKS cache TTL when fetched from URL |
MACP_AUTH_JWT_ALGS | RS256,ES256 | Comma-separated JWT algorithm allowlist (HS256 requires explicit opt-in) |
MACP_MAX_PAYLOAD_BYTES | 1048576 | Maximum envelope payload size in bytes |
MACP_SESSION_START_LIMIT_PER_MINUTE | 60 | Per-sender session creation rate limit |
MACP_MESSAGE_LIMIT_PER_MINUTE | 600 | Per-sender message rate limit |
MACP_LIST_SESSIONS_DEFAULT_PAGE_SIZE | 100 | ListSessions page size when the request sends page_size = 0 |
MACP_LIST_SESSIONS_MAX_PAGE_SIZE | 1000 | Hard cap a requested ListSessions page_size is clamped to |
MACP_METRICS_ADDR | -- (off) | Prometheus text endpoint bind address, e.g. 127.0.0.1:9464 (src/main.rs:584) |
MACP_CONCURRENCY_LIMIT_PER_CONNECTION | 64 | tonic per-connection concurrency limit (src/main.rs:456) |
MACP_MAX_CONCURRENT_STREAMS | 128 | HTTP/2 max concurrent streams (src/main.rs:460) |
MACP_REQUEST_TIMEOUT_SECS | 30 | Per-request timeout (src/main.rs:464) |
MACP_SHUTDOWN_DRAIN_SECS | 10 | Graceful-shutdown drain deadline (src/main.rs:524) |
MACP_CHECKPOINT_INTERVAL | 0 (disabled) | Log entries between checkpoints |
MACP_CLEANUP_INTERVAL_SECS | 60 | Background maintenance interval in seconds: TTL expiry, memory eviction, disk GC, and eager observation of mode-computed deadlines (the handoff implicit accept, RFC-MACP-0010 §5.1(2)) |
MACP_SESSION_RETENTION_SECS | 3600 | Age (from session start) at which terminal sessions are evicted from memory; their durable data is kept |
MACP_SESSION_DISK_RETENTION_SECS | 0 (keep forever) | Age (from session start) at which terminal sessions' durable data is deleted; 0 disables disk GC entirely |
MACP_STRICT_RECOVERY | off | Set to 1 to fail on any recovery error |
MACP_POLICIES_DIR | -- | Directory of governance policy JSON files preloaded at startup; a file that fails validation aborts startup, and the wire registry becomes read-only |
MACP_POLICIES_DRY_RUN | off | Set to 1 to validate MACP_POLICIES_DIR and exit 0/1 without starting the server |
MACP_POLICY_SCHEMAS_DIR | -- | Development and CI only; the server never reads it. Path to the spec repository's schemas/json/policy directory, used by the enum_lists_match_the_canonical_schemas parity test -- see the warning below |
RUST_LOG | info | Log level filter |
MACP_POLICY_SCHEMAS_DIR is listed here because it is otherwise documented nowhere, and a contributor changing a registration mirror needs it. It is read only by macp-policy's parity unit test, which asserts the hand-written value-domain mirrors in crates/macp-policy/src/registry.rs still match the canonical schemas. Two warnings:
- Point it at a clean
git archiveexport of the spec commit CI reads, never at a sibling working tree. CI checks the spec repo out atSPEC_REV(.github/workflows/ci.yml), so a sibling checkout that is dirty, or ahead of or behind that pin, produces parity failures that do not exist in CI -- and it can move under you mid-session. Export first:git -C <spec-repo> archive <sha> schemas/ | tar -x -C <tmpdir>, then point the variable at<tmpdir>/schemas/json/policy. - A set-but-missing directory panics by design. Setting the variable asserts the canonical schemas are available, so the test refuses to skip silently. Unset it to fall back to a sibling checkout, or to skip the parity check entirely when no checkout exists.
Governance policy files
Validate a policies directory before you roll it out: MACP_POLICIES_DRY_RUN=1 MACP_POLICIES_DIR=/etc/macp/policies macp-runtime reports every file by name and exits 0/1 without starting the server. See Policy.
When a rejected policy file blocks startup, correct the file — do not delete it. Deleting it lets the runtime boot but silently voids governance for every in-flight session bound to that policy_version: the policy resolves to nothing on replay and commitment enforcement then treats the session as having no policy at all. UnregisterPolicy on a policy live sessions are still bound to does the same. The full mechanism, and the four other operational changes in this release, are in Upgrading into registration-time policy validation.
Storage backends
The runtime supports four storage configurations, selected via MACP_STORAGE_BACKEND:
File backend (default) stores each session in its own directory under MACP_DATA_DIR/sessions/<session_id>/. An append-only log.jsonl records every accepted message, and a session.json snapshot is written on each state change. Writes use an atomic tmp-file-then-rename pattern to prevent partial-write corruption.
RocksDB backend uses an embedded key-value store for higher throughput. Enable it by building with the rocksdb-backend Cargo feature and setting MACP_STORAGE_BACKEND=rocksdb. The database path defaults to MACP_ROCKSDB_PATH.
Redis backend stores session data in a remote Redis instance. Enable it with the redis-backend feature and set MACP_STORAGE_BACKEND=redis with MACP_REDIS_URL pointing to your Redis instance.
Durability matrix
The runtime acknowledges a message only after the log append "commit point". What that acknowledgement guarantees differs by backend:
| Backend | Acked ⇒ survives process crash | Acked ⇒ survives host power loss | Notes |
|---|---|---|---|
file | yes | yes | log appends fsync before ack; snapshots are tmp-fsync-rename atomic |
rocksdb | yes | yes | log appends sync the WAL before ack; session snapshots are async (recovered via replay) |
redis | yes (Redis process survives) | no | RPUSH acks in Redis memory; no WAIT/AOF barrier. Cache-tier / single-writer only — the runtime logs a warning at startup |
All backends skip individually corrupt log entries on load (with a warning)
rather than failing the whole session; MACP_STRICT_RECOVERY=1 makes
recovery-level errors fatal.
Single-writer: state authority lives in the runtime process's memory. Two runtimes must never share one data directory or Redis instance.
Observation-surface authorization
GetSession is participant/observer-scoped, while ListSessions and
WatchSessions return metadata for all sessions to any authenticated
identity (RFC-0006 permits this shape). Deployments with confidentiality
requirements between agent groups should front these RPCs with a proxy or
restrict which identities may call them. WatchSignals requires
authentication; ListModes/GetManifest/WatchModeRegistry/WatchRoots
are open discovery surfaces by design. ListSessions is now paged, and the
decision not to sign its opaque page_token rests on exactly the unfiltered
property described here -- a forged cursor can only reposition a caller within
a listing it may already read in full, so if per-caller filtering is ever added
to ListSessions, that no-signature decision must be re-analyzed first.
Memory-only mode disables persistence entirely. Set MACP_MEMORY_ONLY=1 for testing or ephemeral workloads. All session data is lost when the process exits.
Authentication
The runtime applies a pluggable resolver chain assembled at startup:
- JWT bearer (active when
MACP_AUTH_ISSUERis set) -- validates signature, issuer, audience, and expiration against a JWKS. Default algorithm allowlist:RS256,ES256;HS256(shared-secret) requires explicit opt-in viaMACP_AUTH_JWT_ALGS=HS256. Thesubclaim becomes the sender; an optionalmacp_scopesclaim carries capability flags (allowed_modes,can_start_sessions,max_open_sessions,can_manage_mode_registry,is_observer). - Static bearer (active when
MACP_AUTH_TOKENS_FILEorMACP_AUTH_TOKENS_JSONis set) -- looks up opaque tokens in a preloaded identity map. AcceptsAuthorization: Bearer <token>or the alternatex-macp-token: <token>header. - Dev-mode fallback -- activates only when neither JWT nor static bearer is configured. Any
Authorization: Bearer <value>header authenticates the caller as sender<value>with full capabilities. Intended strictly for local development.
If a credential matches a resolver but fails verification (expired JWT, unknown static token), the request is rejected with UNAUTHENTICATED -- the chain does not fall through to a later resolver.
Crash recovery
When persistence is enabled, the runtime rebuilds all sessions from their append-only logs on startup. This process is fully automatic:
- Each session's
log.jsonlis replayed through the mode engine to reconstruct the session state. - If a checkpoint exists, replay starts from the checkpoint and only processes subsequent entries.
- Temporary files (
.tmpsuffixes) left by interrupted atomic writes are cleaned up. - The number of recovered sessions is logged at startup.
Log append failures are treated as fatal: the runtime rejects the message rather than acknowledging it without a durable record. This ensures the log is always the authoritative source of truth.
If MACP_STRICT_RECOVERY=1 is set, the runtime exits on any recovery error. Without it, individual session recovery failures are logged as warnings and the remaining sessions are loaded normally.
Monitoring
The runtime provides operational visibility through several mechanisms:
Logging -- All significant events are logged to stderr: session creation, resolution, expiration, recovery results, persistence failures, and rate limit hits. Set RUST_LOG to debug for detailed request-level logging.
TTL enforcement -- Sessions are expired both lazily (on next access) and proactively by a background task running every MACP_CLEANUP_INTERVAL_SECS. This ensures expired sessions are cleaned up even if no new messages arrive.
Mode deadlines -- The same background task observes deadlines a mode computes, on the same interval. The one in the standards-track modes today is the handoff implicit accept (RFC-MACP-0010 §5.1(2)): an outstanding HandoffOffer past its acceptance.implicit_accept_timeout_ms is accepted by the runtime, which appends the HandoffAccept to accepted history and publishes it to StreamSession subscribers. MACP_CLEANUP_INTERVAL_SECS is therefore the observation-latency bound on that acceptance -- but only the latency: the recorded timestamp_unix_ms is the computed deadline whichever path emits it, so changing the interval never changes permanent history. The deadline is also observed on demand, ahead of the next session-scoped message, so no message is ever evaluated against a stale offer regardless of the interval. The sweep runs after TTL expiry within a pass: a session whose TTL and implicit-accept deadline both lapsed unobserved expires rather than accepting, matching the on-demand path's precedence.
Session eviction -- Terminal sessions (resolved, expired, or cancelled) are evicted from memory once their age exceeds MACP_SESSION_RETENTION_SECS (default one hour), measured from session start rather than from when they became terminal. This is on by default and bounds memory usage. Their data remains on disk and can be replayed if needed. Deleting that durable data is a separate, opt-in step governed by MACP_SESSION_DISK_RETENTION_SECS, which defaults to 0 -- disk GC does not run at all unless you set it.
Log compaction -- When a session reaches a terminal state, the runtime automatically compacts its log into a single checkpoint entry. This reduces storage footprint for completed sessions.
gRPC reflection (dev/debug only) -- Build with the reflection Cargo feature (cargo build --features reflection) to enable the standard gRPC Server Reflection API, so tools like grpcurl/grpcui can call RPCs without a local .proto/.protoset file (e.g. grpcurl -plaintext localhost:50051 list). Off by default and never enabled in the published Docker image -- it exposes the full service/message schema to any client that can reach the port, so treat it the same as MACP_ALLOW_INSECURE: fine for a local instance or an ephemeral CI container, not for a production deployment. The runtime logs a warning at startup when it's active. Without this feature, point grpcurl -proto/-protoset at the published proto instead (@multiagentcoordinationprotocol/proto on npm, macp-proto on crates.io).
Container deployment
Use the repository's Dockerfile (multi-stage, non-root macp user) to build
your own image, or pull the image this repo publishes to GitHub Container
Registry: ghcr.io/multiagentcoordinationprotocol/macp-runtime. The
image does not enable dev mode: the runtime refuses to start without
configured authentication and TLS unless you explicitly opt in for local
development:
# Production: configure auth + TLS. Pin an immutable :X.Y.Z tag (see below) --
# substitute the version you're deploying.
docker run -p 50051:50051 \
-e MACP_AUTH_TOKENS_FILE=/etc/macp/tokens.json \
-e MACP_TLS_CERT_PATH=/etc/macp/tls.crt -e MACP_TLS_KEY_PATH=/etc/macp/tls.key \
-v ./secrets:/etc/macp ghcr.io/multiagentcoordinationprotocol/macp-runtime:X.Y.Z
# Local development ONLY: any bearer token becomes a fully-privileged identity
docker run -p 50051:50051 -e MACP_ALLOW_INSECURE=1 ghcr.io/multiagentcoordinationprotocol/macp-runtime:X.Y.ZWhen deploying in containers:
- Mount a persistent volume at
MACP_DATA_DIRso session logs survive container restarts. - Expose port 50051 (or the port configured via
MACP_BIND_ADDR). - Provide TLS certificates and auth tokens via mounted secrets.
- Set
MACP_BIND_ADDR=0.0.0.0:50051to accept connections from outside the container.
Published image tags
Every image is multi-arch (linux/amd64, linux/arm64) with SLSA provenance
and an SBOM attached, built by .github/workflows/docker.yml. Every push to
main (a "branch build") publishes three of these together, always pointing
at the same digest; a release publishes the other two, also always sharing a
digest with each other but never with a branch build's:
| Tag | Mutability | Meaning |
|---|---|---|
:X.Y.Z (e.g. 0.8.1) | Immutable | Exactly one release. Pin this for production. |
:X.Y (e.g. 0.8) | Moves within the minor | Always the newest X.Y.Z patch published so far. |
:latest | Moves on every push to main | The tip of main — not the newest release. Do not use latest in production; it is not a release channel. |
:main | Moves on every push to main | Synonymous with :latest today (both are branch-build tags on the one push: branches: [main] trigger) — documented separately since nothing in docker.yml guarantees they stay synonymous if a second branch trigger is ever added. |
:<sha> (e.g. :bb45705) | Immutable | One image per commit pushed to main (branch builds only — a release build never gets a SHA tag; see below). |
A release commit is main's tip at the moment it's cut, but main moves
on immediately afterward as later commits land — so by the time you read
this, latest/main/:<sha> are very likely already ahead of every semver
tag. Never assume latest matches the newest release.
A :X.Y.Z tag and its release commit's :<sha> tag do not share a
digest, even though they're built from the same source tree: the release
build and the branch build for that commit run as two separate, concurrent
CI jobs, and each stamps its own org.opencontainers.image.created
timestamp. Same source, two images — don't expect docker inspect to report
identical digests across them.
Every image's org.opencontainers.image.revision label names the exact
commit it was built from — this is reliable even for a manually backfilled
release image (a workflow_dispatch naming an older tag), where
org.opencontainers.image.revision still correctly reports the tag's
commit. The one residual gap on a manual backfill: buildkit's SLSA
provenance attestation records the runner's own GITHUB_SHA (the dispatch
commit, i.e. whatever main was at dispatch time), not the backfilled tag's
commit — only the revision label is authoritative there. Every image
published by the normal, automatic release path (cutting a release via
release-plz) has correct provenance with no such gap.
Do not pin the macp-runtime-v* git tag format as an image tag — GHCR
image tags drop the macp-runtime-v prefix (macp-runtime-v0.8.1 → image
tag 0.8.1). docker.yml's resolve step strips it deliberately, to emit a
bare semver value for docker/metadata-action's {{version}}/
{{major}}.{{minor}} patterns — the prefix is redundant once the image
already lives in the macp-runtime repository path, not a Docker tag
syntax restriction (macp-runtime-v0.8.1 would itself be a perfectly valid
tag string).
Development tools
For development and CI, these additional tools are useful:
cargo-tarpaulinfor coverage reportingcargo-auditfor dependency security auditingbuffor protocol buffer linting and management