Symptom: No decision.finalized canonical event ever lands for the affected commitment
step — that absence, not any error event, is the tell. The run does not hang forever:
the stream goes quiet, STREAM_IDLE_TIMEOUT_MS (default 120s) ends the idle iteration,
the consumer retries/backs off/resubscribes up to STREAM_MAX_RETRIES (default 5), then
degrades to a getSession poll fallback, and once that poll budget is also exhausted the
run is finalized failed with the message polling exhausted without terminal session state (src/runs/stream-consumer.service.ts). Along the way GET /runs/{id}/eventsdoes show session.stream.opened events with status: reconnecting and this terminal
failure — but that message is a red herring: it describes the stream/poll mechanism
timing out, not the real cause, which is the silently rejected commitment described below.
(SESSION_POLL_TIMEOUT_MS cannot be the culprit here — it only governs the one-time wait
for the initiator to open the session, before bindSession/subscribeSession, and is
long past by the time a mid-session commitment is rejected.)
Cause: RFC-MACP-0013 §9 (runtime v0.7.0) tightened the §7.3.1 supersedes check: a
commitment whose supersedes.commitment_hash is not canonical — exactly sha256: +
64 lowercase hex characters — is hard-rejected on accept, with no dual-read
window. Because the control-plane only sees accepted envelopes via its read-only
StreamSession, and the rejection is delivered as an inline MACPError frame on the
sending agent's own bidi stream — not on the observer stream, and not as a negative
ack — the control-plane's view is a silent absence: the commitment
carrying the malformed supersedes ref simply never appears. There is no
message.send_failed or other error event to grep for — the symptom is purely "the run
stopped progressing." Any agent still emitting a pre-0013 placeholder hash ("abc123",
uppercase hex, an uppercase SHA256: prefix, or a whitespace-padded value) against a
≥0.7.0 runtime hits this, and a supersession chain that crosses the version boundary is
permanently severed — the sender must re-emit with a canonical hash; the runtime will
not retroactively accept the old one.
Checks:
Confirm the stalled run has a commitment step expected to carry supersedes
(cross-session commitment reference, RFC-MACP-0001 §7.3.1) and that no
decision.finalized ever landed for it.
Ask the sending agent's own logs for the inline MACPError frame it received for the
rejected CommitmentPayload send — the control-plane never observes this rejection,
only the sending agent does.
If a prior supersedes ref did make it into the projection (e.g. from before the
runtime was upgraded), check GET /runs/{id}/state → decision.current.supersedes:
canonical: false confirms a legacy/malformed hash is in play and flags the agent
code that needs to move to canonical hashes. Note projections written before this
field shipped omit canonical entirely rather than reporting false — it is
derived only when a new decision.finalized is applied, so rebuild the projection
(POST /runs/{id}/projection/rebuild) if the key is absent.
Fix: Update the sending agent to emit a canonical sha256: + 64-lowercase-hex value.
There is no server-side remediation — the control-plane is observer-only and cannot
repair or replay a rejected send.
Symptom: Log line auth_mint_failure reason=... or JWT mint failed; falling back to static bearer.
Explanation:MACP_AUTH_SERVICE_URL is set, but the auth-service is down, returned non-2xx, or its response was unparseable. The credential resolver automatically falls back to RUNTIME_BEARER_TOKEN for this call.
Checks:
Is the auth-service reachable? curl -X POST $MACP_AUTH_SERVICE_URL/tokens -d '{}' -H 'content-type: application/json' (expect a 4xx response, not a connection error).
Is RUNTIME_BEARER_TOKEN set as a fallback? Without it the call proceeds with the deprecated dev bearer (Authorization: Bearer ${RUNTIME_DEV_AGENT_ID}) when RUNTIME_USE_DEV_HEADER=true, which only a dev-mode runtime (MACP_ALLOW_INSECURE=1) accepts — otherwise it fails auth on the runtime side.
If the auth-service is healthy but calls still fail, check MACP_AUTH_SERVICE_TIMEOUT_MS (default 5000 ms) — slow auth-services can time out under load.
Symptom: Log line bindSession no-op for run <uuid>: cannot transition ... (current status=running).
Explanation: Not an error. Two paths can race to bind the same run — RunExecutorService for POST /runs-created runs, and SessionDiscoveryService for runs auto-discovered via WatchSessions. Whichever arrives second sees the run already past binding_session. As of the subscribe-session PR, the second call is a logged no-op; it no longer crashes the process.
When to investigate: only if you see this repeatedly for the same runId — that would indicate a loop somewhere retrying the bind. A single occurrence per run is normal.
Symptom: A run attached to an already-RESOLVED session goes straight to completed, with a decision snapshot but no per-message timeline.
Explanation: Not an error. The runtime compacts a session's log on resolve, so the control plane binds, marks the run running, skips subscribeSession, and lets the poll-only consumer emit the snapshot (see ARCHITECTURE § Request Flow). CANCELLED sessions are still rejected with 409 and EXPIRED with 400. Session lifecycle semantics: macp-runtime/docs/architecture.md.
Symptom: A run has a session.stream.gap canonical event and its projection is flagged historyGap.
Explanation: The resume ordinal was compacted away (FAILED_PRECONDITION). The consumer deliberately degrades to poll-only instead of resubscribing from 0, because the control plane has no message-id dedup. Track via macp_stream_resume_gap_total; details in ARCHITECTURE § Stream Resume.
Symptom: The counter increments with a step label (metrics, publish_event, publish_snapshot, span_events).
Explanation: The events are already committed; only the side effect failed (e.g. Redis StreamHub down). The failure is logged and swallowed on purpose. Investigate the named step's dependency; SSE clients can resume with afterSeq.
Symptom:POST /runs/:id/messages, /signal, or /context returns 410 Gone with errorCode: ENDPOINT_REMOVED.
Explanation: The control-plane is observer-only as of the 2026-04-15 direct-agent-auth refactor. Agents authenticate to the runtime directly and emit their own envelopes via macp-sdk-python / macp-sdk-typescript. See docs/API.md § "Messages & Signals — emission is NOT via the control-plane" for the mapping, and the SDK guides for the new agent flow: macp-sdk-python direct-agent-auth, macp-sdk-typescript agent-framework.
Symptom: Agents call session.send(...) via the SDK but events don't appear in GET /runs/:id/state.
Checks:
Confirm the run's runtimeSessionId matches the session_id the agent is writing to (GET /runs/:id).
Check stream consumer logs for StreamSession reconnection loops — the observer subscribes read-only and must be connected.
Confirm the runtime echoes envelopes back on the stream (some runtimes only echo certain message types). signal.emitted and message.sent canonical events require stream-envelope entries on the observer stream. See macp-runtime/docs/API.md#message-transport for StreamSession semantics and macp-runtime/docs/sdk-guide.md#streaming for the observer lifecycle.
Symptom:docker build fails during npm ci with a 401/404 from npm.pkg.github.com, or cat: /run/secrets/npm_token: No such file or directory.
Explanation: The private proto package installs from GitHub Packages, and the Dockerfile reads the token from a BuildKit secret (npm_token) — not a build-arg.
Fix:
docker build --secret id=npm_token,env=GITHUB_TOKEN -t macp-control-plane .# docker compose: export GITHUB_TOKEN (classic PAT with read:packages) first
Start test postgres: docker compose -f docker-compose.test.yml up -d postgres-test
Test DB uses port 5433 (not 5432) to avoid conflict with dev DB
Real runtime tests fail with InvalidPayload:
Use payloadEnvelope with proto encoding instead of plain payload
Set INTEGRATION_RUNTIME=remote and RUNTIME_ADDRESS=127.0.0.1:50051
Prometheus metric re-registration error:
Tests that create multiple NestJS apps must call promClient.register.clear() between apps
"Test suite failed to run" even though every assertion passed:
Teardown leak — background observation services (StreamConsumerService, SignalConsumerService, SessionDiscoveryService) had in-flight persistRawAndCanonical work when the DB pool closed. Fixed by test/helpers/test-app.ts → drainBackgroundWork() which awaits each service's bounded drain before Nest's own onModuleDestroy sweep. If you see this in a new test, make sure you created the app via createTestApp(...) so the app.close() wrapper is in place.
Some behavior (e.g. ListSessions pagination across a large session store) can only be verified against a real macp-runtime — the mock runtime used by default in npm run test:integration never returns a paginated response. To run one locally:
Build the runtime binary (there is no prebuilt binary; the toolchain is pinned in ../macp-runtime/rust-toolchain.toml):
cd ../macp-runtime && cargo build --bin macp-runtime
MACP_ALLOW_INSECURE=1 is mandatory — without auth configured the runtime refuses to start — and it also waives the TLS requirement. In this mode the runtime accepts any Authorization: Bearer <v> as sender <v>, so the control-plane's existing dev-bearer fallback (RUNTIME_USE_DEV_HEADER) works unchanged.
Omit MACP_MEMORY_ONLY if you want the persisted .macp-data store loaded at boot, but that store is not sufficient by itself for pagination testing: as of this writing it holds 131 persisted sessions, and all of them are terminal (Resolved/Expired/Cancelled). On startup the runtime logs evicted stale sessions from memory (registry + log cache + stream bus) count=131 — it evicts terminal sessions older than a cutoff (../macp-runtime/src/runtime.rs:1200-1237) — and ListSessions/GetSession return nothing for them afterward. Verify this yourself before trusting a persisted-store session count (the runtime does not support gRPC reflection, so use grpcurl with the checked-in proto):
To actually exceed a ListSessions page you must seed live (OPEN) sessions using an agent-role client — the control-plane must never emit envelopes itself (see the observer-only invariant above), so seeding has to go through macp-sdk-python/macp-sdk-typescript or a raw gRPC client acting as an agent, calling SessionStart directly against the runtime. Note the page size that matters here is the control-plane's ownRUNTIME_LIST_SESSIONS_PAGE_SIZE (default 200) — the CP always sends an explicit pageSize on every ListSessions call, so the runtime's own server-side default (100) never applies once the CP is the caller. Gotchas that cost real debugging time when doing this:
The mode id is macp.mode.decision.v1 — notmacp.mode.decision.
SessionStartPayload.configuration_versionandmode_version are both mandatory and must be non-empty strings (../macp-runtime/crates/macp-core/src/session.rs:344-346); an empty string is rejected, it is not treated as "use default".
The ack success field on the response is ok — notaccepted.
MACP_SESSION_START_LIMIT_PER_MINUTE defaults to 60 per sender, so seeding more than 60 sessions in a minute requires spreading the SessionStart calls across distinct sender identities (distinct Bearer values), not just looping as one sender.
Run the control-plane's integration suite against it:
INTEGRATION_RUNTIME=remote RUNTIME_ADDRESS=127.0.0.1:50051 \ RUNTIME_LIST_SESSIONS_PAGE_SIZE=50 npm run test:integration
Specs gated with test/helpers/real-runtime-gate.ts (describeWithRealRuntime) — e.g. test/integration/list-sessions-pagination.integration.spec.ts — only run in this mode; they are skipped, not failed, under the default mock runtime.
The live pagination spec asserts pagesFetched > 1, which requires the seeded store to span more than one page at the control-plane's configured page size. This doc does not prescribe a session count, so check what your store actually holds: if it is smaller than the default RUNTIME_LIST_SESSIONS_PAGE_SIZE (200), the whole store fits in one page and the spec fails with its "only fetched 1 page(s)" guard. There are two remedies, and you only need one: lower RUNTIME_LIST_SESSIONS_PAGE_SIZE (e.g. =50, as above, which spans multiple pages for any store above 50), or seed more than 200 live sessions so the default page size itself spans multiple pages.