Skip to content

Latest commit

 

History

History
571 lines (497 loc) · 33 KB

File metadata and controls

571 lines (497 loc) · 33 KB

MirrorNeuron API Specification

Blueprint review metadata

Blueprint list and detail resources preserve the SDK catalog's skills list as {name, version_constraint} entries and the explicit air-gapped boolean. Skill source locations and other dependency metadata are excluded by the shared projection. An empty list denotes no declared skill dependencies. These read-only resources report requirements, not installed versions or per-run invocation history, and do not prepare runtime resources.

Purpose

mn-api is the HTTP gateway for the local MirrorNeuron runtime. It translates HTTP and Server-Sent Event interactions into calls to the MirrorNeuron Python SDK and returns browser- and desktop-consumable responses. Blueprint run requests apply the SDK-owned input validator before hardware, model, helper-process, or submission work and return a field-level HTTP 422 problem when required configuration is missing. The package installs the mn-api service and the mn-web-ui-server static/proxy service.

This specification covers only this repository. Core runtime semantics and shared Python behavior are external contracts consumed through mirrorneuron-python-sdk.

Compact run monitors bind resource aggregation to the same resolved run directory as their artifacts. Core's explicit storage reference takes precedence over local mappings; an unavailable explicit reference does not fall back to a different directory. Missing or mismatched directory identity leaves resource usage unavailable; no second local-root lookup substitutes empty or unrelated measurements. The SDK owns numeric usage aggregation and its bounds. The dependency minimum is SDK 1.3.58.dev83, including canonical measured call ledgers, stable Job conversation sources and bound container output paths.

Run read endpoints may project a run from a verified SDK local mapping and its replicated submission when Core cannot return the run. A nonterminal shared record is reported as unknown; writable Core run operations still require Core. Human review annotations are separate from Core run control: an authenticated response to a pending request writes to the verified mapped submission ledger when that ledger exists, and otherwise to the local run ledger. Unknown or closed requests return HTTP 409. API-started durable Job runs with a declared host output copy start the SDK host delivery relay, including starts that reuse an unchanged definition. The API loads the committed submission's copy contract without rebuilding the catalog bundle. It also reconciles completed Core-scheduled runs from replicated completion receipts. The relay reports delivery separately from Core execution completion.

Launch keeps the compatibility phase identifier model_install, but it is a blocking preflight: every declared DMR model is selected, installed or reused, and published through the selected node's LiteLLM gateway before a job is submitted. This keeps first downloads outside worker liveness windows. The SDK first-use path remains idempotent for dynamically requested skill models and for a model removed after launch; it must not be the normal path for declared requirements. Explicit model-install routes remain eager.

GET /api/v1/models defaults to the full merged catalog plus discovered installations; installed_only=true preserves the installed-only filter. Model records expose default: true for the configured default and its fallback chain. PUT /api/v1/models/default/installation prepares the first compatible configured model on a local or cluster node without changing default policy. Model installation runs as an API-owned background operation through SDK preparation and gateway synchronization. The existing 202, operation Location, authentication, request fields, and idempotency contract apply to both logical defaults and explicit model IDs.

Managed catalog records expose source: "dmr" or source: "docker". Explicit Docker IDs use the same installation operation, with structured Docker settings and credential environment references resolved by the SDK on the selected native owner. NIM readiness precedes successful registration and gateway synchronization. The API does not accept registry credentials in installation bodies or implement a separate Docker lifecycle.

Catalog blueprint loading applies the shared mn.payloads.v1 contract before agent rendering or validation. Payload Python dependencies participate in HostLocal environment preparation, large assets are staged by reference, and payload model files are packaged locally before normal runtime-model preparation.

Owned Surface

The API owns:

  • the FastAPI application and /api/v1 route surface;
  • the /api/v1 HTTP adaptation for stable-job definitions and one-to-many execution runs;
  • the stateless, URL-bound supervisory MCP adaptation at /api/v1/jobs/{job_id}/mcp for MCP-enabled blueprint jobs;
  • HTTP request parsing, schema validation, authentication, and request limits;
  • JSON/problem response shapes and transport-level status mapping;
  • run/job progress over snapshots and authenticated, resumable SSE;
  • durable group-operation start, status, and resumable SSE event adaptation;
  • API-side launch coordination, artifact access, and process cleanup helpers;
  • API and Web UI server configuration; and
  • injection points that make route behavior deterministic in tests.

It does not own workflow scheduling, gRPC runtime semantics, manifest expansion, model placement algorithms, blueprint domain behavior, or the browser UI.

Interface Contract

Job creation and blueprint-run requests accept the optional observation header X-Launch-Progress-ID. Authenticated GET /api/v1/launch-progress/{progress_id} exposes presentation-only snapshots while the existing synchronous request is pending. The view excludes configuration, validation payloads, and raw failure diagnostics; it retains at most 200 recent events. Blocking preparation stages refresh their current description and elapsed time every ten seconds. Reporting threads end with the wrapped call and never execute, retry, or cancel a job. No artificial completion percentage is inferred from elapsed time.

  • Every REST path is rooted at /api/v1. /api/v2, historical aliases, runtime-run paths, and WebSocket routes are not mounted.
  • GET /api/v1/health advertises api_contract: "mirrorneuron.rest.v1".
  • REST requests and responses do not contain a transport version field.
  • The surface includes runtime/system health, blueprints, bundles, jobs, runs, schedules/events, deployments, models, services, resources, artifacts, and realtime progress.
  • DELETE /api/v1/nodes/{node_id} removes a federated peer through Core and returns its node name and removal status.
  • Existing model-remote and model-proxy REST paths are compatibility adapters over the SDK-owned $MN_HOME/models/registry.json; they do not restore the removed CLI proxy/manual-remote lifecycle or project legacy ledgers.
  • Collection responses contain items and an opaque next_page_token.
  • Streaming endpoints must terminate on completion, error, timeout, or client disconnect and must not leak background tasks.
  • Job workflow definition and latest-run shape endpoints expose public step dependencies and step-agent membership. Workflow-progress polling and run event-stream snapshots expose status and progress without shape fields. Hidden lowered runtime nodes such as start/end/fork/join nodes are transitively projected into source-facing edges and layers at the API boundary.
  • Run workflow-progress uses the saved manifest for that execution ahead of a mutable stable Job definition. It uses the durable workflow ledger for step state when available and replays up to 5,000 recent events for activity; runs without a durable ledger replay the full event history. A root failure is present only while the run is failed; earlier retry errors remain in events.
  • Group operations use fixed Core-owned kinds (cancel_all_jobs, clear_jobs, reconcile_node, and drain_node). Their item events are replayable by sequence cursor. cancellation_pending is accepted durable work, while explicit item failures remain operation failures.
  • Blueprint lifecycle terminology matches the CLI: clients create asynchronous additions with POST /blueprints/{blueprint_id}/additions and removals with POST /blueprints/{blueprint_id}/removals. The removed blueprint installation resource is not mounted.
  • Blueprint additions are API-owned local operations. Their snapshots and SSE events expose monotonic percent, stage, label, and detail fields; terminal failures expose sanitized stable codes, retryability, hints, and bounded prerequisite issues. A successful addition records the blueprint locally so another client-side CLI step is neither required nor permitted as part of the HTTP workflow.
  • job_id is stable and owns configuration, schedules, and shared data; run_id is the execution/control identity. Batch starts create fresh runs. POST /jobs/{job_id}/run-operations accepts an Idempotency-Key and returns a durable API launch operation immediately. Its authenticated operation SSE stream reports progress and terminal success with the created run_id, or a sanitized failure. Replaying the same request after an API restart resumes an interrupted operation and reconciles a Core update that may have committed before its RPC deadline. A changed body with the same key is rejected. A rejected run start caused by temporary runtime resource overload returns HTTP 503 Problem Details with code MN_RESOURCE_EXHAUSTED; no schedulable runtime node returns MN_SCHEDULING_UNAVAILABLE. These run-start responses do not create an accepted run. Other operations keep their existing errors. A blueprint-owned Web UI is optional singular Job state: its durable handle is served only at GET /api/v1/jobs/{job_id}/ui and is shared by every run of that Job. Run UI paths are not mounted. The local mn-web-ui-server may use that authenticated handle to proxy a declared external service UI and its explicitly allowlisted companion ports only while the handle is running; paused, stopped, cancelled, and failed handles are not proxied. The API itself does not expose a general remote-proxy route. Job UI discovery rechecks the exact page and its same-origin code assets through the shared Web UI SDK probe from the proxy host. Registration or a previously passing service check cannot establish readiness. The returned page and assets are checked through the existing iframe proxy at the configured runtime Web UI URL. Only authenticated probe requests read the address/lifecycle descriptor with check_readiness=false, avoiding recursive discovery while retaining lifecycle and port restrictions. Descriptor reads do not establish readiness; ordinary discovery defaults to a fresh check. Probe authentication never reaches worker services. Ordinary iframe and WebSocket requests continue to reject unready handles. The returned web_ui.metadata.readiness receipt is current; an unreachable running handle is projected as starting without rewriting durable lifecycle state. Inactive handles are never promoted. The proxy rejects unready handles and forwards live multipart bytes as they arrive rather than waiting for a full buffer. Discovery carries metadata.load_event (dom-ready by default, optionally did-finish-load) from the current service or its matching SDK claim. This desktop visibility policy never bypasses the page/asset readiness check. Only executable type: service jobs are single-run: ordinary second starts return HTTP 409 Problem Details with code service_run_exists, while explicit replace_existing_run requires a fresh caller-supplied run_id and returns that run plus optional replaced-run and deferred-cleanup metadata. Retry/recovery attempts retain their run ID.
  • Schedule creation accepts Idempotency-Key and replays identical requests. Schedule detail read and update are pending Core/SDK RPC support.
  • Archive retains shared data. Data reset and permanent job deletion are explicit operations; confirmed deletion is rejected while runs are active. Individual run deletion never deletes job data.
  • Stable job_id, execution run_id/execution_id, and internal diagnostic runtime_run_id are separate identities. Clients use only run_id in URLs. Output and saved-event reads use the canonical run_data_ref while work is active, ahead of a final result reference. File access resolves and verifies the exact SDK submission/run binding without requiring a legacy local mapping. Run SSE uses that same output identity for saved events and keeps the execution ID for progress/control. Job MCP context selects the Job's explicit latest-run identity rather than the first collection entry. Its workflow evidence uses the execution ID and its file reads verify the canonical submission/output binding. Bounded MCP actions and service activity watches resolve services under the execution ID, independently of the output ID used for artifacts and saved notices.
  • Blueprint launch creates a stable job plus its first run unless an existing job_id is supplied, and returns both identities. Existing jobs receive the freshly prepared manifest and payloads through atomic bundle replacement before the new run starts. Existing service jobs use a separate replace_existing_run field; blueprint force remains validation bypass and never implies destructive replacement.
  • When owner_node selects a federated Core, launch preserves that owner through asynchronous request normalization and passes the same selected-node handoff into SDK preparation before forwarding Job creation. Distributed workflows prepare HostLocal Python environments on that owner without changing their placement declaration.
  • Local-only blueprints with a host OS requirement prepare HostLocal Python on the native host and prepare a separate Core SDK proxy. Native preparation must return mn.native.host-python.v1 and the proxy environment before submission. Extras and local SCM version identities use shared SDK helpers; API installation includes the SDK's local-source extra. The SDK dependency floor is 1.3.58.dev46, below 2.0.
  • Background output relays poll the execution run ID, which remains separate from the durable job ID used for definition paths and launch responses. Starting a catalog-backed stable Job through /jobs/{job_id}/runs creates the same per-run mapping and relay as /blueprints/{blueprint_id}/runs, including when no configuration override is supplied.
  • Blueprint launch delegates manifest expansion, config application, dependency localization, environment injection, and topology lowering to the same SDK preparation path consumed by the CLI.
  • Blueprint launch accepts a bounded optional secret_environment map. Names must be declared by the blueprint through pass_env; values are injected only into matching executable workers and are excluded from resolved configuration, API responses, progress events, and public monitor manifests.
  • Blueprint launch persists the SDK's sanitized source-facing monitor manifest beside the run identity mapping, matching mn blueprint run and preventing lowered control nodes from appearing as public workflow steps.
  • Blueprint-specific live controls are served by the owning blueprint service. mn-api does not translate product action routes into runtime messages.
  • Stable-job creation may resolve a catalog blueprint_id; only API-trusted catalog sources or uploaded bundle roots are read from the host filesystem. Caller-provided arbitrary host paths are rejected. Creation and executable configuration updates run host-side command input validators before storing the definition. The SDK records those successful checks and removes command rules that Core is intentionally forbidden to run.
  • A legacy MCP-enabled catalog Job exposes the read-only tools get_job_profile, get_latest_run, get_job_context, and watch_job_activity through Streamable HTTP. A response-enabled Job exposes those tools plus ask_job. Tool inputs cannot select another Job. Context responses use mn.mcp.job_context.v1, contain at most 50 evidence records and 256 KiB, and retain the stable profile with warnings when latest-run data cannot be read. Never-run, running, paused, scheduled-waiting, idle, and archived Jobs remain readable; deleted, unknown, and non-enabled Jobs share a sanitized not-found response.
  • ask_job accepts an 8,000-character question, an optional UUID conversation ID, and an optional 128-character request ID. It returns the bounded mn.mcp.job_answer.v1 contract through Core's owner-routed unary query, never creates a Run, and has no REST, SSE, or UI chat equivalent.
  • watch_job_activity performs a bounded active wait. When the blueprint's bounded response agent declares watch_operator_activity, the API routes the watch to that Job's owner node. The Job response agent resolves exactly one passing Run-scoped service and uses the SDK MCP client to receive its activity through MRTR before the API relays it to the chat client through MRTR. Delivery receipt is transport-only and never acknowledges a review notice.
  • The stable supervisory MCP excludes credentials, secret/environment values, raw logs, host paths, arbitrary files, and unrestricted artifact bodies. It cannot mutate job, run, schedule, approval, or configuration state.
  • Legacy Core mn-job-collaboration services remain separate, run-scoped peer surfaces for blueprints that have not opted into the definition response service.

The route definitions and generated OpenAPI document are authoritative for exact fields and paths. tests/test_v1_contract.py protects consumer-visible behavior. Executable replacement uses PUT /jobs/{job_id}/bundle with an opaque bundle_id.

Errors

Validation and application failures use RFC 9457 Problem Details with stable error codes. Responses use application/problem+json and contain type, title, status, detail, instance, code, and request_id, with bounded field issues in errors. SDK and gRPC failures are normalized before reaching clients.

Client responses and logs must not expose secrets, authorization values, raw payloads, tracebacks, or unsanitized internal context. Request and correlation IDs may be returned for diagnosis.

Security and Configuration

  • MN_API_TOKEN enables bearer authentication. Protected HTTP and SSE routes plus stable job MCP accept the Authorization header. Credentials are never accepted in URL query parameters.
  • MN_API_REQUEST_SIZE_LIMIT_BYTES bounds declared request body size.
  • CORS is disabled unless origins are explicitly configured.
  • Artifact and bundle paths must remain within their permitted roots.
  • mn_api.config, mn_api.config_env, and mn_api.config_schema remain compatibility facades over SDK configuration loading and typed shared keys. API-only keys are composed locally. Precedence is environment > .env.<profile> > .env > defaults, including explicit blank environment overrides.
  • MN_MODEL_CATALOG_PATH selects SDK-owned model defaults, entries, and fallback links. The API carries no physical built-in model policy.
  • Sensitive configuration values are redacted based on schema metadata and secret-like key names.

The current supported keys and defaults live in mn_api/config_schema.py and are documented in .env.example and README.md.

Dependency Boundary

Routes call shared clients and helpers rather than duplicating SDK business logic. External services are supplied through configured state/dependencies so tests can use fakes. Importing the package or constructing the app must not require live Core, Redis, Docker, OpenShell, or network access. Public workflow manifests and bounded job activity are projected by SDK helpers; routes only add HTTP transport concerns.

Compatibility

Changes to paths, methods, required fields, response/event shapes, status codes, error codes, authentication, or default behavior require focused contract tests, consumer review, and documentation. This release is an intentional clean break with no compatibility facade.

Verification

The repository acceptance gate is:

python -m ruff check .
python -m pytest
python -m build

The test configuration requires at least 85 percent branch coverage. Live system behavior belongs in cross-repository system tests; this repository's normal suite must remain deterministic and dependency-injected.

Canonical blueprint packages

All blueprint folders and ZIPs use the blueprint/v1 manifest schema and role documents owned by mn_sdk.blueprints. Catalog indexes contain ordered package paths only. Catalog reads are data-only; explicit compilation produces Core's runtime manifest. Default configuration, local overwrites, and invocation values resolve through the SDK. Payload assembly stages one resolved descriptor; launch environment is passed explicitly. Confirmed failure permits owned resource rollback; uncertain submission acknowledgements require reconciliation. See mn-docs/blueprint-standard.md for document ownership and extension schemas.

Blueprint uploads have a separate configurable file-size limit, capped by the SDK format-v1 maximum of 32 GiB. Multipart framing has a 1 MiB allowance. Other requests retain the ordinary HTTP body-size policy. Upload files and extracted contents are independently checked before a package can be launched.

A Job configuration PATCH that repeats the saved resolved configuration reuses the prepared definition and keeps the Job revision unchanged. Other requested attributes still update normally.

Starting a saved Job through POST /api/v1/jobs/{job_id}/runs reuses its prepared definition when overrides are absent or do not change its resolved configuration, matching mn job start. It does not rediscover the catalog, rebuild worker resources, or rerun placement against transient node status. Changed configuration still passes the normal preparation and revision-checked update before start. If another request saves the same resolved configuration during preparation, the run uses that prepared Job and discards its redundant preparation. A different concurrent configuration change remains a conflict.

Streaming read-only Job answers

ask_job on response-enabled read-only Job MCP endpoints accepts stream: bool, default false. Omitted/false returns the existing completed JSON tool result, for agent collaboration. True selects SSE on the same authenticated endpoint and relays actual visible answer text using MCP progress notifications; callers should request progress with a progress token/callback. Each notification message is JSON with schema_version mn.mcp.job_answer_delta.v1, request_id, sequence (starting at 1), and delta. The final result remains mn.mcp.job_answer.v1.

The API reads bounded cursor updates over the existing Job response RPC, at most once per 50ms with a 250ms idle read wait. It cancels the runtime stream when delivery fails or its task is cancelled. An interrupted stream is not converted to an unrelated fallback answer. The relay has a 90-second overall deadline. Bounded action-agent ask_job retains its existing non-streaming turn contract.

The new SSE transport keeps MCP's loopback Host/Origin protection enabled (127.0.0.1, localhost, IPv6 loopback, with explicit ports). Existing JSON transports retain their settings. Non-loopback SSE exposure requires an explicit trusted-host configuration in the transport before deployment; it must not be enabled by disabling rebinding protection.

Shared admission error contract

CLI and API use mn_sdk.error_catalog.ERROR_CATALOG for admission error codes, category, safe message, remediation hint, HTTP status and retryability. Existing codes remain stable. Unknown errors retain MN_EXECUTION_FAILED; raw exception text never becomes an admission explanation. Applications must branch on codes, not message text. AppError.category and AppError.retryable are additive.

Problem code Symbolic code Category HTTP Retryable
1001 MN_MEMORY_REQUIREMENT_UNMET hardware 422 false
1002 MN_CPU_REQUIREMENT_UNMET hardware 422 false
1003 MN_GPU_REQUIREMENT_UNMET hardware 422 false
2001 MN_GPU_MEMORY_UNAVAILABLE capacity 503 true
2002 MN_DISK_UNAVAILABLE capacity 503 true
2003 MN_RESOURCE_EXHAUSTED capacity 503 true
3001 MN_SCHEDULING_UNAVAILABLE scheduling 503 true
4001 MN_PLACEMENT_UNSATISFIED placement 422 false

Hardware errors require a configuration or hardware change. Capacity and scheduling errors may succeed after availability changes; retryable is not a promise of success or authorization to replay a submission automatically. A mixed placement failure reports all observed blockers rather than claiming that every node lacks memory. Unknown/mixed causes require inspection.

Placement errors include bounded details.blockers with code, safe message, one-based node index, and (when measured) required/available amounts, unit, resource, and operator (>= or >). A validated node_label may identify an explicitly reported friendly PC name; raw runtime node IDs and addresses are not labels. Node indices refer to the sorted placement snapshot, not persistent node IDs. Preparation host-memory requirements use total capacity. Core run admission reports memory available for runs after limits and reservations. GPU device memory uses free capacity, preserving a reported zero. Values are shown in GiB (1024 MiB), matching existing placement conversion. Missing measurements are not invented. Public summaries contain no paths, secrets, raw messages, or arbitrary diagnostics; they show up to eight blockers, while structured details contain up to 100. Truncated reports retain the general placement code rather than claiming a cluster-wide resource shortage.

Core appends a bounded mn_admission_v1 JSON marker to placement_failed: gRPC details. The SDK validates the versioned blocker fields, owns the numeric error catalog and messages, and emits the same AppError through CLI and REST. For example: spark has 8.17 GiB of free GPU memory; this work requires 48 GiB. Code MN_GPU_MEMORY_UNAVAILABLE / 2001 suggests stopping other GPU workloads or unloading unused models before retrying. Busy CPU/GPU reservations use MN_RESOURCE_EXHAUSTED / 2003, while insufficient CPU/GPU hardware retains 1002 / 1003. Requirements and scheduling decisions remain unchanged. Malformed, unknown, or legacy generic device errors do not imply a memory cause. Legacy RuntimeError catches still work for preparation placement failures; normalization retains structured identity through launch exception wrappers.

Core's legacy overload and no-schedulable-node markers remain normalized at run-start boundaries. Other operations retain their existing transport error interpretation. No protobuf shape or numeric code changes are required. Deploy the updated Core and SDK together to enable measured run-admission errors; older Core versions retain their existing, less specific error behavior.

Numeric problem codes for automation

problem_code is a stable integer shared across SDK errors, CLI JSON, API Problem Details and individual placement blockers. It is independent of HTTP status and process exit code. Human CLI output also prints the numeric code. Existing symbolic code remains backward compatible. SDK consumers use ProblemCode (IntEnum), PROBLEM_CODES (numeric-to-symbolic lookup), and problem_code(symbol); ERROR_CATALOG holds admission defaults. Assigned values must never be renumbered or reused. Message wording can evolve without changing a problem code. Clients must handle unrecognized values as unknown errors and must not automatically retry them. An unregistered SDK symbol maps to 9000; an unexpected execution failure maps to 9001.

Ranges reserve related problem families: 1xxx hardware, 2xxx capacity, 3xxx scheduling, 4xxx placement, 5xxx input/configuration, 6xxx access/resource state, 7xxx runtime/transport, 8xxx cancellation, 9xxx unknown/internal failures. Exact codes, rather than ranges or message matching, drive automated remediation. category provides a finer label.

Example API problem (HTTP 422):

{
  "problem_code": 1001,
  "code": "MN_MEMORY_REQUIREMENT_UNMET",
  "category": "hardware",
  "status": 422,
  "detail": "Host memory: requires 48 GiB; available 24 GiB.",
  "hint": "Select a node with enough memory or reduce the workflow's memory requirement.",
  "retryable": false
}
from mn_sdk import ProblemCode

if response["problem_code"] == ProblemCode.MEMORY_REQUIREMENT_UNMET:
    # Choose a larger runtime or change requirements before submitting again.
    pass

Job performance

Authenticated GET /api/v1/jobs/{job_id}/analysis returns the SDK job analysis: job_id, snapshot_at, scope: recorded_history, history_complete, runs, running_time, and tokens. Coverage accompanies nullable duration/token values. Unknown jobs use the normal not-found problem response; analysis deadlines return 504. Disconnects cancel further collection. No model calls or runtime starts occur.

Shared assistance adapter

The common dependency floor is 1.3.58.dev0, below 2.0, so both SCM development builds and final releases providing this contract can be installed.

POST /api/v1/assistance/evaluations is authenticated and read-only. Strict request models accept a selected blueprint, optional stable Job/execution, non-secret local setup readiness/mode/revision, dispositions, and an explicit requested help kind. The API verifies identities against its runtime context, projects pending input and declared batch capabilities, and delegates evaluation to SDK common. Strict response models expose mn.assistance.v1, a revision, bounded context, and a typed supported action opportunity. Context-only access does not relax the separate response-service gate for MCP/Job answers. Existing setup, scheduler, lifecycle, and interaction APIs remain authoritative for effects. Optional assistance_task on Job answers and streaming starts has only bounded goal/state/execution/next-question metadata. It is checked against the current execution; it cannot grant permissions or expand the bounded-agent effects.

Durable checkpoint retry

Authenticated v1 run resources expose POST retry-plans and retries. Planning accepts only bounded explicit integer/string setting overrides. Submission requires a positive expected attempt, SHA-256 checkpoint revision and an Idempotency-Key header. Unknown fields and malformed selections are rejected before Core invocation. Core is the sole eligibility/idempotency authority. Submission returns 202 and the accepted attempt identity under the same run ID. Stored history remains inspectable with its record source and missing/unavailable control status; historical presence must not mask Core unavailability for retry.

Job ZIP backup and restore

Authenticated POST /api/v1/jobs/{job_id}/backups returns a private application/zip full offline capsule through the shared SDK. Pause active runs first. Temporary download files are removed after delivery.

Authenticated POST /api/v1/job-restorations?start=true accepts the raw ZIP body with Content-Type: application/zip and returns a 201 new job identity, start status and optional run ID. Omit start=true to restore ready work. ZIP validation, platform and destination hardware admission precede dependency preparation and job creation; failures return HTTP 422 problem responses. This route permits at most 128 GiB and 100,000 capsule entries, while existing route limits remain unchanged. Temporary uploads are removed on completion or failure. Restore does not invoke blueprint additions, validation by catalog ID, or hiring. If start fails after creation, the response retains the new job ID and an actionable start_error.

Requires matching Core ExportJobBackup / RestoreJobBackup streamed RPCs and SDK mn.backup.v3 support. The destination runtime, Python and Docker installation must already be available on a compatible OS/architecture/Python ABI.

Collaboration group catalog contract

Catalog projections preserve the SDK-validated mn.collaboration.group.v1 contract: topology: group, protocol, fixed goalId, mutually accepted blueprint IDs, member capacity (2–16), goalKey/commonGoalKey, groupKey, peersKey, and sameRuntime. Group peer configuration uses stable {jobId, blueprintId} entries. Invalid declarations and private metadata are omitted. Legacy pair declarations remain capacity two. This contract does not launch work or grant approvals.

Preparation timing

Blueprint preparation and job submission write local INFO events worker.preparation.start and worker.preparation.finish, including a static stage, outcome and elapsed milliseconds. Timings cover bundle validation, workflow/dependency resolution, packaged models, sandbox images, host Python environments, payload/runtime staging, native resources and job submission. They are recorded even without a launch-progress ID. Labels, details, configuration, commands and exception contents are excluded. Progress and logging sink failures do not change the submission result or deadlines.