Blueprint execution.json can set runtime.placement.must_run_local: true to
pin its Job owner and all workers to the submitting computer. mn blueprint run
honors that requirement before resource/model preparation, rejects a remote
--node, and cannot disable it through the single-node environment opt-out.
requirements.os: "darwin" additionally requires a Mac, including sample runs.
Install the matching updated SDK and CLI to use this contract.
mn runtime ensure-context-engine prepares the authenticated CPU Membrane package.
Blueprints declare Markdown memory with mn.context / text_memory.enabled.
Preparation uses Markdown storage and one DuckDB index per job; it does not
prepare a GPU compressor. The private context_auth.token is reused across
starts and forwarded through runtime settings. text_memory.enabled=false
disables blueprint preparation, including when its source descriptor is enabled.
The matching v2 service, SDK and persistent Compose template must be installed
before a live run. Optional detailed logs use MN_CONTEXT_OBSERVABILITY=true.
Preparation checks the authenticated v2 protocol even when the container is
already running. An older image fails before submission with a matching-image
instruction. Local development can prepare the current engine with
mn-deploy/install.sh --mode local --build-membrane; a data reset is unnecessary.
mn model list shows every discovered installation of a DMR model, including
local and multiple remote owners of the same unregistered artifact.
Interactive mn run watch keeps completed, failed, and cancelled runs open for inspection until q or Ctrl+C. Terminal snapshots stop polling; redirected output still finishes automatically. Child tasks appear beneath their parent sub-workflow with descriptive labels, stable instance IDs, round, status and elapsed time; [ / ] page tasks and f follows the active or failed task.
Completed cross-node blueprint runs wait up to 120 seconds by default for the
replicated final result and declared output files before completing the local
output copy. Set MN_OUTPUT_COPY_TIMEOUT_SECONDS to adjust this deadline.
Long blueprint launches and mn job create report the current preparation stage,
including dependencies, payload staging, Docker workers, and runtime submission.
mn blueprint run applies the same shared input validator as
mn blueprint validate before checking runtime resources, installing models, or
submitting a job, so missing required inputs fail immediately with the relevant
--set hint.
Validation also accepts runtime bindings for SDK-admitted child-workflow
templates while rejecting unknown step bindings.
Interactive terminals show a spinner with elapsed time. Plain and redirected
human output receive a concise “still waiting” line every ten seconds on stderr.
First-time image builds may take several minutes; an elapsed timer indicates a
pending call, not a measured completion percentage. Existing model-download
progress remains available. --json output stays unchanged.
mn-cli provides the mn command for validating and running blueprints,
inspecting runtime state, managing jobs, exporting artifacts, and starting local
services installed by mn-deploy.
Requires mirrorneuron-python-sdk>1.3,<2.0.
Install locally and run tests:
python3.11 -m venv .venv
. .venv/bin/activate
.venv/bin/python -m pip install -e .
.venv/bin/python -m pytest -qTry the CLI:
mn --version
mn runtime start
mn node list
mn blueprint run message_routing_traceAfter a completed blueprint run copies both its internal run store and a configured user output folder, the human-readable summary prints only the user-facing output destination.
mn runtime start starts a normal, federation-capable Core. There is no
separate worker mode: every Core owns its Redis state, can own jobs, and runs
all agents for those jobs locally. The successful start output includes the
advertised host, gRPC port, node identity, a join token, and the exact
mn node add command to run from another Core. Treat the displayed token as a
credential, keep terminal output private, and use mn node refresh-token when
it must be rotated.
The command omits --grpc-port when the advertised endpoint uses the default
port (55051); it includes the option when a non-default port is required.
With the default Syncthing shared storage enabled, mn node add configures and
verifies both sidecars before registering the Core federation pair. A failed
shared-storage connection creates no new Core federation pair and tells the
operator to correct the sidecars before retrying.
Runtime startup readies the host-native SDK service before Core. Calling
mn runtime start again reuses a healthy native service, preserving warm
definition-scoped response engines while the remaining runtime is reconciled.
If reconciliation recreates Core, the CLI also restarts the local API so its
gRPC client identity matches the recreated runtime.
mn node list shows each node's hostname, health, and role (connection mode
plus job-ownership eligibility) without artifact-only fields such as kind,
owner, or update time. Use mn node show <node> for the complete endpoint and
capability record.
Remove a reciprocal federation registration with explicit confirmation:
mn node remove mirror_neuron@spark --yesThis detaches the peer from the current Core only; it does not delete the peer's owner-local jobs or data.
The CLI uses mn-python-sdk-web-ui for static HTML exports and for the
job-scoped iframe handle that forwards a blueprint-owned web service.
mn model exposes one type-aware workflow: list, add, show, probe,
start, stop, unload, update, remove, and doctor. Add a catalog or arbitrary DMR reference with
mn model add <MODEL>, or register canonical provider JSON with
mn model add --file <definition.json>. Registrations are stored in
$MN_HOME/models/registry.json; provider secrets remain environment-variable
references. Use mn model list --available to include catalog-only choices.
Use mn model show to display the full merged catalog with default markers
for the configured default and its fallbacks; mn model show <MODEL> keeps
the single-model detail view. Both forms show stored facts without live probes.
Run mn model add --default to prepare the first hardware-compatible default
on the selected local or cluster node. It honors catalog overrides and reuses
registrations; choosing a fallback does not replace the logical default policy.
The default list is federation-wide: it merges each connected node's published
Docker Model Runner inventory. A discovered artifact reports ready when its
LiteLLM route is live and installed while route reconciliation is pending;
registry ownership is not presented as a health state.
Add --default with an explicit model or file to make that model the logical
default ahead of the built-in Nemotron/Gemma fallback chain. Provider files
used with --default must contain exactly one model.
When the requested DMR artifact is already installed locally or on a cluster
node, mn model add reuses it and creates the same managed registry record
without pulling a second copy.
The SDK catalog now identifies delivery with source: "dmr" or
source: "docker". NVIDIA Docker models follow the same placement, registry,
gateway, and run cleanup contracts. The first entry is Cosmos3 Nano Reasoner,
using the VSS NIM 1.7 Docker recipe from DGX Spark. Set NGC_API_KEY or
NGC_CLI_API_KEY in the native runtime service environment on its Linux NVIDIA
owner, then use:
mn model add cosmos3 --node spark
mn model stop cosmos3 --node spark
mn model start cosmos3 --node spark
mn model unload cosmos3 --node spark
mn model update cosmos3 --node spark
mn model remove cosmos3 --node spark --yesadd, update, and preparation cache/configure models without starting
inference. stop and unload retain images and cache; explicit start waits
for readiness. An owner-gateway request starts an idle NIM before inference,
recreating a deleted managed container from its image/cache when necessary.
The last run/request owner releases model memory. Docker model containers use
restart=no; doctor reports a cached/stopped installation as idle without
loading it. Lifecycle commands infer a single registered
owner; select --node or --local for replicas. Preparing Cosmos does not
reinstall Nemotron. See the SDK's
Docker model setup and settings,
including NGC authentication and a shared-GPU Spark override.
A managed model may be installed on more than one eligible node. Repeat --node
and combine it with --local to add all requested replicas in one operation:
mn model add small --local --node spark
mn model add medium --node spark --node gpu-2The CLI preflights every target before installation, records successful
replicas if a later target fails, and makes retries idempotent. mn model list --json and mn model doctor report each installation separately. An
untargeted model update updates all recorded replicas. Removing a replicated
model requires --local, one or more --node values, or --all-nodes.
Force a live capability evaluation and save the effective LiteLLM-facing matrix in the SDK model-catalog overlay:
mn model probe
mn model probe gemma4:e2b
mn model probe nemotron-3.5-lightning:latest --json
mn model probe gemma4:e2b --capabilities image,json-schema,stream,thinkingWith no model argument, probe evaluates every model in the same
federation-wide runtime inventory returned by mn model list. It continues
after individual failures, reports every result, and exits unsuccessfully when
one or more probes fail. Supplying a model keeps the single-model behavior. The
default probe covers embeddings, image input, strict JSON Schema output, SSE
streaming, and thinking. When the selected DMR artifact is local, the CLI runs
the identical probe directly against Docker Model Runner and fails if the
LiteLLM result differs. Remote-owner and provider models are tested through the
managed LiteLLM route without exposing or bypassing the owner's direct endpoint.
Model-aware blueprint launch logic is testable without Core, Docker, DMR,
LiteLLM, SSH, or a network. RuntimeModelDependencies supplies the model
catalog, resource report, system summary, BlueprintModelOps, and gateway
effects used by the real run_bundle handler. The reusable
tests/runtime_model_fakes.py cluster records model preparation, remote-route
reconciliation, and LiteLLM synchronization in memory.
Run the focused gate from this workspace:
../mn-system-tests/.venv/bin/python -m pytest -q \
tests/test_run_cmds_models.py tests/test_run_cmds_run.py \
-k "adaptive_model_placement or injected_remote_installed_state or injected_cluster"The runtime-selection scenarios are:
- a local-only 16 GB Apple node validates the portable Gemma fallback policy;
- adding a healthy 128 GB CUDA node validates that Nemotron is feasible;
- already-installed remote models remain usable without a second install;
- blueprint launch selects and prepares the owner node before job submission, then workers use that node's reachable LiteLLM gateway route.
mn blueprint run --debug prints the selected model and preparation result for
blueprint-declared foundational LLMs. RAG and OCR model details are owned by
their skills and appear only in runtime events when those skills first call the
SDK wrapper. Those events report the selected model/node, fallback reason, and
install/reuse state. Debug mode
also prints DockerWorker build commands and complete captured build output,
including builds performed through a remote node's native SDK service.
The run monitor header keeps the workflow, run, and job identity but omits the
blueprint description so progress begins immediately below it.
Its lower section includes a fixed-height, timestamped event tail that follows
the newest workflow events without growing the monitor.
Reattachment uses the manifest projection saved for that exact run; it never
guesses an unrelated catalog blueprint when older run metadata lacks a
blueprint_id.
While the API progress stream is available, mn run watch displays its public
step snapshot directly. The phase counters, current step, worker activity, and
percentages therefore match the API for the same run.
When a remote owner is verifying SDK-staged local inputs, the same monitor
shows Waiting for staged inputs on <node> until Core dispatches the workflow.
The interactive monitor intentionally omits LLM token totals and budgets because
runtime event counters are not authoritative; resource telemetry remains
available to its dedicated commands and structured consumers.
While a model is prepared at launch, the run monitor keeps preparation to a
single compact Preparing <model> on <node>… status below the workflow and
agent progress grid. Model events still carry detailed phase, timing, and byte
telemetry for logs and structured consumers; a lack of byte progress does not
itself fail the job.
Live Spark checks are a separate, opt-in boundary smoke after this injected gate passes; they are not the development loop for placement policy.
CPU-only HostLocal workflows stay on the submitting runtime node by default.
When the local Core runs in Docker, prepared HostLocal Python environments are
reported to submissions through the Core-visible cache mount rather than the
host filesystem path.
Local source requirements such as /workspace/mn-python-sdk[context] are
staged into that cache with their extras and declared source version preserved.
The extras are separate from the source path when checking trusted roots or
rebasing a checkout path for a selected runtime node.
When a local-only workflow also requires a host OS, HostLocal Python workers run
on the native host through the SDK service. Core uses a separate prepared SDK
proxy environment for supervision and cancellation. Source skills retain their
own SCM versions; CLI installation includes the SDK's local-source extra.
The CLI requires SDK >=1.3.58.dev46,<2 for this native execution contract.
Automatic HostLocal service ports use MN_AUTO_PORT_START through
MN_AUTO_PORT_END (62000-62049 by default in the local Docker runtime). That
range is published only on host loopback; the runtime's internal proxy marker
allows the service process to accept Docker forwarding without advertising a
non-loopback endpoint.
Detached runs keep their output relay alive until terminal state unless
MN_RUN_EVENT_RELAY_MAX_SECONDS is explicitly set.
Override blueprint config for one run without changing config/overwrite.json:
mn blueprint run ./vc_assistant \
--set document_sources.folder_path=/path/to/documents \
--set execution.debug=trueRepeat --set for multiple values. Values use JSON types when possible and
otherwise remain strings.
Normal mn blueprint run commands validate declared required inputs from the
merged configuration before submitting a job. For
input_validation.required: ["input_folder"], supply
--set inputs.payload.input_folder=/path/to/source or configure a non-empty
default. --force intentionally bypasses input validation.
For a blueprint-owned web service, override the listener without editing its checked-in config:
mn blueprint run ./cctv_operator --web-ui \
--web-ui-host 0.0.0.0 \
--web-ui-port 61017--web-ui-host and --web-ui-port set web_ui.service.host and
web_ui.service.port for that run. A wildcard host exposes the service to
reachable peers; the blueprint is responsible for its authentication and
network-safety contract.
When a blueprint declares a web_ui service or a deferred job-scoped Web UI,
--web-ui reports the local /jobs/<job_id>/ui dashboard route. Deferred
handles may appear after job submission, but the reported route is stable.
Docker Compose service handles include an allowlist for the dashboard's
declared video and WebSocket companions, so the local Web UI server can proxy
a selected remote node without sending the browser directly to that node's LAN
IP.
mn job create selects an eligible runtime from hardware and workflow
requirements before preparing workers. That selected node owns the submitted
definition. --node explicitly constrains this selection. The blueprint source
is unchanged; mn job start reuses the stored placement and resources.
Create a reusable job once, then start independent runs that share its declared job data:
mn job create ./vc_assistant --job-id vc-diligence
mn job show vc-diligence
mn job start vc-diligence --inputs run-input.json
mn run list --job vc-diligence
mn run show <run-id>
mn run pause <run-id>
mn run resume <run-id>
mn run cancel <run-id>When an older run is no longer visible to Core, run list and run show
recover verified terminal status from its replicated submission. A stored
running marker without a live Core record is shown as unknown.
In an interactive terminal, run pause, resume, and cancel display a spinner
while Core processes the request. JSON and MN_CLI_OUTPUT=plain output remain
free of transient progress so they are safe for automation.
job_id is the durable configuration and data owner. run_id is one
execution and the identity used for control, logs, output, retention, and run
deletion. Starting the same job again creates another run; retrying a run does
not. Use mn blueprint run --job-id <job-id> to run an existing definition.
Without that option, blueprint run creates a durable job and starts
its first run. Human-readable submission, detach, summary, and watch output
labels the durable Job ID and execution Run ID separately.
mn job list shows each definition's canonical Type (service or batch) and
its Node. Node is the Core runtime (owner_node, such as
mirror_neuron@spark) that durably owns the definition and its job data; it is
not a human or account owner.
With --job-id, the CLI prepares the currently installed blueprint revision
and atomically replaces the inactive job's executable bundle before starting
the run. Job data, schedules, and earlier run history are preserved. In
contrast, mn job start and scheduled dispatches are source independent and
reuse the stored definition-scoped submission and Docker services.
Lifecycle commands are deliberately separate:
mn job archive vc-diligence # retains shared data
mn job reset-data vc-diligence # confirms; clears/reseeds and advances generation
mn run delete <run-id> # confirms; cancels an active run, then removes it; never deletes shared data
mn job delete vc-diligence # confirms; deletes all runs, runtime resources, definition, and dataPermanent job deletion automatically cancels and clears attached active runs
before removing the definition and shared data. Run deletion does the same for
an active run before detaching it. If an archive must wait for an
unavailable owner runtime, the command and mn job list report
archive_pending until federation replay settles it.
If a confirmed deletion must wait for an unavailable owner, mn job delete
reports delete_pending and skips submitter-local cleanup; the stale job is
hidden from normal lists while Core replays the owner cleanup on reconnect.
Job and run deletion allow up to five minutes for Core cleanup (with a small client-side forwarding margin), including when the definition belongs to a federated owner node. If any mutating command still times out, the CLI warns that the owner may still be processing it; inspect the current job or run state before retrying. Missing identifiers are reported directly, with the relevant list command instead of a generic execution failure.
Execution status and controls use mn run ...; attached blueprint progress
uses the canonical workflow-progress stream and the same public-step contract as the
launch-time monitor.
mn node reconcile and mn node drain start durable Core operations and
render item updates in completion order. MN_CLI_OUTPUT=plain emits stable
→, ✓, and ! Warning: progress lines; the rich terminal shows live
counters and recent results.
If the owner of a cancelled job is offline, cancellation_pending means the
request was accepted and cleanup is queued for that node's rejoin. It is not a
command failure. Ctrl+C detaches without aborting the operation; reattach with:
mn operation show op-…
mn operation watch op-…Configuration is loaded by mn_cli.config. .env files provide defaults, and
real environment variables always override them. MN_ENV selects the
environment-specific defaults file and defaults to dev when unset.
Precedence:
real environment variables
> .env.${MN_ENV}
> .env
> built-in safe defaults
Development:
export MN_ENV=dev
cp .env.example .env.dev
mn --versionTests:
export MN_ENV=test
mn --versionProduction does not require any .env file. Provide deployment-specific values
through the real environment:
export MN_ENV=production
export MN_HOME=/var/lib/mirrorneuron
export MN_LOG_LEVEL=info
export MN_API_HOST=0.0.0.0
export MN_API_PORT=8080
mn runtime statusmn runtime start --host <LAN-IP> persists the local runtime identity in
$MN_HOME/docker-compose.env. Blueprint launches automatically use that
identity for Compose placement, so MN_NETWORK_ADVERTISE_HOST does not need to
be exported for ordinary local submissions. An explicitly exported value still
overrides the persisted identity.
Keep secrets, credentials, production hostnames, production database URLs, cloud
credentials, and user-specific local paths out of source files. Use
environment variables or uncommitted .env files instead.
After a successful interactive mn runtime start, the CLI checks the newest
stable install_support/v* snapshot in MirrorNeuronLab/mn-deploy. If a newer
release is available, it only prints a reminder to run mn runtime upgrade; it
never prompts for or installs an upgrade during startup or any other command.
mn runtime upgrade is the explicit, confirmation-protected installation
command. Its release plan pins the Core release tag, the SDK/CLI/API Python
package versions, and the Web UI npm version. The upgrader installs the exact
Python package versions from the public GAR agent-skills index and configures
the exact Web UI npm version for Docker Compose; it does not follow a source
branch, package-manager latest tag, or the Core repository's latest-release
endpoint. A component is shown as an upgrade only when the snapshot version is
strictly newer than the installed stable version, so a stale snapshot cannot
offer a downgrade. The former mn runtime update command now points to
mn runtime upgrade.
The Core remains a versioned GitHub Release binary because it is not a Python
or npm package. Its release asset URL is constructed from the same support
snapshot tag. For private mirrors, set MN_DEPLOY_REPO, MN_DEPLOY_REF,
MN_PIP_INDEX_URL, or MN_PIP_EXTRA_INDEX_URL before running the command.
- A running MirrorNeuron core is required for live runtime commands.
- Validation failures render in wrapped terminal tables, so long requirements and recommended fixes remain readable in narrow terminals.
- The default gRPC target comes from
MN_GRPC_TARGET, then local deployment settings, thenlocalhost:55051. - Use
mn blueprint validatebeforemn blueprint run ./folderwhen checking a local bundle. - Validation honors first-use runtime-model preparation, so a compatible declared model need not already be installed.
mn blueprint runvalidates model declarations but does not install models. Workers select, install, and route each managed model on its first actual use.- Docker workers receive a worker-reachable model-control target and use the SDK to select the best cluster node independently for LLM and for model specifications supplied at runtime by RAG and OCR skills.
- Node-local workflows are hard-pinned as a whole after topology lowering. Runtime health rejects nodes whose coordination-store identity differs from the submitting Core or whose Redis endpoint is read-only.
- OpenShell sandbox ownership uses the durable job ID and definition submission ID. Each definition revision gets a separate sandbox, and submission commits its preparation record so resource cleanup preserves active work.
- OpenShell workers that reuse a job-scoped sandbox are prepared before
submission; the submitted node receives the concrete sandbox name and SSH
host instead of asking Core to create host resources.
The submitting host must have an OpenShell CLI matching the gateway version
in
~/.local/binorPATH, even when Core runs in Docker. A missing CLI is reported as a preparation prerequisite before image builds or sandbox registration, rather than as a missing blueprint. Local managed gateways build sandbox images through the Docker CLI, which honors the active Docker context (including Docker Desktop on macOS). OpenShell build contexts use SDK source staging to preserve local package versions and requested extras when Git metadata is absent from the image. defaultis a LiteLLM model group, not a concrete model. Its preferred and fallback entries come from the SDK catalog'sdefaults.llm.modeland per-entryfallback_modellinks. The existing cluster model monitor rebuilds these routes as nodes join, rejoin, or leave; incomplete peer snapshots retain the last safe routes until departure is confirmed. The steady-state inventory check runs once per minute by default. If a federated Core's shared snapshot omits a peer or remains stale, the monitor reads that peer's authoritative Core snapshot directly instead of increasing poll frequency. Gateway route names andfallback_modelare read from the SDK's merged model catalog, including~/.mn/models/catalog.json(or$MN_HOME) and the highest-priorityMN_MODEL_CATALOG_PATHoverride. When a local runtime's DHCP address changes, the monitor rehomes only DMR registrations whose artifact is confirmed on the local host and whose former owner is absent from live membership. It then rebuilds gateway routes from the current live node address; this avoids treating an unverified.localname as a cluster endpoint.- See Cross-node model routing through LiteLLM for the complete proxy topology, node join/departure reconciliation, replica load-balancing, admission limits, and operator diagnostics.
--debugretains complete Docker build diagnostics and prints model preparation results, including selected node and install/reuse state.
Shared configuration parsing and defaults are owned by mn_sdk.config.
mn_cli.config remains a source-compatible facade that composes CLI-only keys
with the SDK schema. Layering is environment > .env.<profile> > .env > defaults, and an explicitly blank environment value overrides dotenv. Set
MN_MODEL_CATALOG_PATH in .env to select an operator catalog containing both
semantic defaults and model entries; there are no separate preferred/fallback
model-name environment variables.
Skill dependencies must declare a version in the manifest or runtime package index, or carry an explicit version constraint in a configured requirement. An unversioned skill fails preparation before installation; no global skill version or SDK-version fallback is applied. References to skills already versioned in the manifest use that declaration. Development staging preserves the SDK's own static project version or, for the running SDK source checkout, its installed distribution version. Missing SDK version metadata is an error.
Blueprint SDK capabilities use the canonical dependencies.json.packages list,
with full distribution names, type: pip, source: gar, and exact versions.
The SDK owns resolution: local development uses source projects and ignores
package/skill release pins; binary mode retains GAR requirements. This applies
to HostLocal and DockerWorker submissions, including blueprint-owned skills.
The workflow monitor event feed prefers explicit activity messages over worker identifiers, including bounded tool-query previews and outcomes. Long entries are ellipsized; recorded events retain the bounded message.
The installed API, native SDK, and Web UI executables are resolved under
$MN_HOME/venv/bin (default ~/.mn/venv/bin), alongside runtime state.
The supervised native runtime checks shared-storage pairing every 30 seconds, independently of model reconciliation. Incomplete pairing is retried every five seconds, including when an already-registered peer starts after the local runtime. Each pass reloads persisted sidecar settings and reads current authenticated Syncthing device identities before restoring reciprocal device and folder registration. Unchanged configurations are not rewritten. Disabled storage and unavailable peers never trigger resets, data deletion, or workflow restarts. The monitor uses existing federation authorization and never joins unknown nodes. Credential changes that invalidate existing federation access require node rejoin.
Install the updated CLI and SDK and restart the native runtime service to enable the monitor. No runtime reset or blueprint modification is required. Registration is distinct from completed file transfer; DockerWorker preparation still verifies the staged context and its readiness marker before building.
mn runtime start persists a private, versioned MN_HOME/node-identity.json
before starting an independently federated Core. Existing names are preserved;
new names use mirror_neuron_<uuid>@127.0.0.1 and do not depend on Wi-Fi or DNS.
The name is an identity, not the remote connection address. Conflicting saved,
configured, or container identities stop startup without changing job ownership.
mn runtime status --json reports identity.expected, identity.actual, and
identity.valid. An unnamed or mismatched Core is critical and exits nonzero.
Use mn runtime start to restore missing startup configuration. Restore the
original identity from backup when configuration conflicts; do not delete the
identity file or reset job data to bypass the check. Existing records owned by
nonode@nohost require explicit repair and are never reassigned automatically.
After a network change, mn runtime reconnect refreshes desktop endpoint
advertisements without restarting a healthy Core. Explicit DNS endpoints remain
configured; automatically detected IPs are refreshed. Peers must resolve DNS
from their containers. If all known endpoints are unreachable, use mn node add
with the peer's current address and federation token. Direct Erlang clusters
retain their existing naming and discovery contract.
The workflow monitor separates fixed phases from runtime-created sub-workflow steps.
The child panel shows task IDs, parent, round, phase, status, elapsed time and failure
reason. It follows active work, with [ / ] paging and f to resume following;
counts include newly discovered tasks. Four child rows and a three-event tail keep
the child view compact. On macOS, attached runs and detached output
relays hold an idle-sleep assertion for their lifetime. Display sleep remains allowed;
explicit sleep is not prevented. An explicit relay time limit also ends its assertion.
Hardware and scheduling failures use shared SDK codes and explain the cause
without requiring debug mode. For example, a 48 GiB memory requirement on a
24 GiB node reports MN_MEMORY_REQUIREMENT_UNMET, the required and available
amounts, and a hint to select a larger node or reduce the requirement.
CLI JSON and API Problem Details include numeric problem_code (for example,
1001 for memory requirements or 3001 for scheduling), category, retryable, and bounded
structured placement details.blockers. See SPEC.md
for codes and retry semantics.
Measured admission blockers also identify a validated friendly PC name and the
available/required resource amounts. GPU memory shortages use 2001 and suggest
stopping other GPU workloads or unloading unused models before retrying. Updated
Core and SDK services are required for measured run-start errors; admission
requirements remain enforced.
Run mn job analysis <job_id> for all recorded execution statistics, or add
--json for the shared SDK/API result in the standard CLI JSON data field. Counts distinguish successful, failed,
cancelled, running, paused, and other unfinished runs. Duration excludes pauses;
missing/partial measurements and estimated tokens are explicit. Plain mode and
NO_COLOR remain supported. This read-only command does not start the job.
mn run retry <run-id> --dry-run --json
mn run retry <run-id> --set catalog_review.walltime_seconds=3600Retry starts a new attempt of the same failed run from a verified durable
checkpoint. Completed steps remain completed. --set path=value is repeatable
and accepts only declared adjustable settings; omitted settings keep their prior
values. Increasing a 20-minute total allowance to 60 minutes leaves 40 minutes
when 20 minutes have already been consumed. Restoring an API or folder may allow
retry without any settings change.
mn run resume continues a paused run. Start a new run for changed inputs,
workflow topology or result-defining configuration. List/show identify runtime
records versus stored history; historical visibility alone does not guarantee
recovery. --dry-run returns eligibility and blocked reasons without dispatch.
The command prints the request key and selected attempt/checkpoint before submission. After a lost response, reuse those original values and identical settings:
mn run retry <run-id> --idempotency-key <key> \
--expected-attempt <attempt> --checkpoint-revision <revision> \
--set catalog_review.walltime_seconds=3600Standard --json output includes structured planning/submission results and
error context. Truly unknown IDs remain not found; an unavailable Core is reported
separately from stored history without its control record.
Managed Markdown context turns may bind a trusted serving-tokenizer integration
with MN_CONTEXT_TOKEN_COUNTER_FACTORY=package.module:create_counter. The factory
receives request, scope and principal keyword arguments and returns a
Membrane VerifiedCounter calibrated against actual provider prompt usage for
that request's serving route, including tools and schema framing. Workers receive
the setting through native/runtime preparation. No factory means counting remains
unavailable; errors or lexical/byte estimates never authorize dispatch. Install
the serving integration in the worker environment before enabling live managed
turns. Tokenization and context processing must remain on CPU.
Context preparation needs no dedicated compression model. Full-runtime optional
compaction uses the normal LiteLLM default route only after CPU preparation
cannot fit a complete request.
# Pause active runs first; offline dependencies are included by default.
mn job backup <job-id> --output /path/to/job-backup.zip
# On the destination, create a new definition and start a fresh run.
mn job restore --input /path/to/job-backup.zip --start
# To restore without starting, omit --start; --job-id selects a new identity.
mn job restore --input /path/to/job-backup.zip --job-id restored-jobThese commands use mn.backup.v3. The ZIP contains the executable bundle,
configuration, job data, available run history/events/artifacts, staged inputs and
outputs, payload model files, transitive Python wheels, and declared Docker images.
--no-air-gapped omits the offline wheel/image capsule. Missing model assets,
unavailable images, unsupported remote service dependencies, or active runs fail
backup instead of producing an incomplete offline capsule. Backup never replaces
an existing destination file.
Restore checks all ZIP paths and hashes, compatible OS/architecture/Python ABI, and destination CPU, RAM, disk, GPU and runner requirements before allocating resources. It creates independent storage and native resources without catalog access or blueprint hiring. Historical executions remain evidence under the new job data directory; schedules are recreated paused with new identities. Restore never replays source executions. A failed start keeps the new job ready for retry.
HostLocal wheels are built for the actual execution Python, including Docker Core's Linux Python, and checked against the destination before installation. Captured dependencies retain the versions installed in the source environments. The native preparation service creates fresh environments from the complete wheel set with package indexes and dependency URL resolution disabled.
The destination must already have compatible MirrorNeuron, Python and Docker /
Docker Model Runner installations. The capsule supplies job dependencies, rather
than operating-system or runtime installers. Keep it private: configuration and
local data may contain sensitive values. Core, SDK, CLI and API must be upgraded
together for the new streamed backup RPCs. mn blueprint export <run-id> remains
a run report export (JSON/Markdown/HTML), with no job restore counterpart.