Skip to content

About

Multi AI Agent at scale

Resources

Stars

3 stars

Watchers

1 watching

Forks

Repository files navigation

MirrorNeuron 🧠

Durable AI workflows on your own machines.
Package the work. Run it. Inspect every step.

Documentation · Get started · Examples · Contribute

License: MIT Elixir/OTP runtime Redis-backed state Beta status

MirrorNeuron is a runtime for AI workflows that need more than a model call: multiple agents, long-running work, reusable configuration, recoverable execution, and results you can inspect. Start on one machine and federate trusted machines when you need more capacity.

This repository contains MirrorNeuron Core, the Elixir/OTP engine that schedules work, supervises agents, persists runtime state in Redis, and exposes gRPC control and observability. The CLI, Python SDK, REST API, and Web UI build on that engine.

Important

MirrorNeuron is in beta. This README follows the current source documentation; installed releases may differ. Check mn --version and command help when following a guide, and use compatible ecosystem releases.

Child-round plan nodes accept an optional non-empty label of at most 1,024 UTF-8 bytes. Core persists it on the child instance and includes it in public topology events; omitted labels retain the template label or instance ID. Labels do not alter task identity, dependencies or execution.

Why MirrorNeuron?

OpenShell workers synchronize shared outputs before cleaning their invocation workspace. Failed transfers preserve the sandbox copy for recovery.

Agent workflows become harder to operate when they span many steps, wait for input, call external services, or need to survive a restart. MirrorNeuron gives that work an explicit execution model and a durable record of progress.

Capability What it gives you
Reusable blueprints Package a workflow with its configuration, input/output contracts, worker code, and dependencies.
Durable jobs and runs Keep a configured definition and its execution history; inspect, pause, resume, or cancel an individual run.
Observable execution Follow workflow steps, events, logs, and output artifacts through the CLI and Web UI.
Bounded adaptive workflows Use branches, parallel steps, scatter–gather, and declared dynamic regions or child workflows under runtime-enforced limits.
Resource-aware execution Select an eligible owner using declared hardware, model, and service requirements, then admit work locally.
Recovery controls Use supervision, durable delivery, leases, and retry policies, with operator review when replay is unsafe.
Local and configured remote models Route model calls through the job owner's LiteLLM gateway, with Docker Model Runner for local inference.

See Why MirrorNeuron for workload fit and the reliability guide for guarantees and limits.

OpenShell buffers command output until completion, including artifact-handoff workers. Its inherited node beacon timeout does not shorten a workflow task deadline. Explicit workflow heartbeat controls remain authoritative; otherwise the declared step timeout governs both fixed and child tasks.

How it works

Blueprint package
  workflow + execution + contracts + configuration + payloads
                             │
                      SDK validation
                       and compilation
                             │
                      Durable job definition
                             │
                        Run on owner Core
                             │
                 Workflow ledger and scheduler
                             │
                 Agents through declared runners
                             │
                     Events, logs, artifacts

A blueprint packages reusable work. The SDK validates and compiles its source into an executable bundle. A job stores the durable configured definition; a run identifies an execution of it. Use the public job ID for definition operations and the public run ID for execution controls, logs, and results.

Batch jobs can retain multiple runs. Service jobs retain at most one attached run, and pause/resume preserves that identity. Retries belong to the same run and have their own attempt identity.

Blueprint authors declare the logical workflow separately from worker execution. Core owns scheduling and completion boundaries; agents and skills own task behavior. Dynamic workflows can instantiate admitted templates within declared bounds. They cannot arbitrarily rewrite running work.

Child plans (execute and stop) are limited to 128 KiB of compact UTF-8 JSON, including task metadata and immutable artifact references. This accommodates the admitted 128-step ceiling; large inputs and results remain in shared files replicated by Syncthing. Older Core builds enforce 32 KiB, so deploy the matching Core before submitting larger plans. Task and round limits are unchanged.

Read Core concepts, Blueprint format, and Runtime architecture for the full contracts.

One owner per job, independent machines

Each federated Core has its own writable Redis. A job and all of its workers stay on one owner Core. Peers exchange authenticated gRPC requests and cached job/run summaries; federation does not require shared Redis or a Distributed Erlang cluster.

CLI / REST API / Web UI
          │
          ▼
   Core A + Redis A ◄── authenticated federation ──► Core B + Redis B
          │                                               │
     A-owned jobs                                    B-owned jobs
     local workers                                   local workers

Scaling means running independent jobs on eligible owners. A peer connection failure does not transfer ownership: the owner can continue its local work, while remote summaries may become stale. Federation does not automatically migrate an unavailable owner's jobs.

Federated peer status uses an authenticated gRPC reachability check every five seconds. Three consecutive failed checks mark a peer unavailable; one successful check restores healthy status. Job and run summary refreshes run separately, so a refresh failure marks cached summaries stale without changing peer status.

Model routing follows the same owner boundary:

Local model:   worker → owner LiteLLM → owner Docker Model Runner
Remote model:  worker → owner LiteLLM → peer LiteLLM → peer Docker Model Runner

See Federation architecture and Model runtime.

Get started

1. Install and start the runtime

Use macOS, Linux, or Windows with WSL2, with Docker installed and running. The checkout-based installation below also needs Git. Docker Model Runner is needed when your chosen blueprint requires local models; there is no universal model prerequisite.

From a directory where you keep projects, clone the deployment tools and review install.sh and its options before running it. The installer installs the selected runtime components and local support files.

git clone https://github.com/MirrorNeuronLab/mn-deploy.git
cd mn-deploy
./install.sh --help
./install.sh

Then, from any directory:

mn --version
mn runtime start
mn runtime status
mn node list

Startup should report Runtime node ready; resolve failed required health checks before submitting work. Startup also prints a federation join credential: keep it private. Use mn runtime doctor for diagnostics.

The default installed endpoints are:

Surface Default
Web UI http://localhost:55173
REST API http://localhost:54001/api/v1
Core gRPC localhost:55051

Use the endpoints reported by your runtime if configured differently. Local state defaults to ~/.mn (MN_HOME). See Installation for platform setup, release selection, upgrades, and editable workspace mode.

2. Validate and launch a blueprint

Choose a package from Examples or a blueprint author. Read its README, inputs, dependencies, and execution policy first. A blueprint can execute code and call external services; validation checks its contract, not whether it is safe to trust.

Run these commands from the directory containing your reviewed blueprint, replacing ./my-blueprint with its actual path:

mn blueprint validate ./my-blueprint
mn blueprint doctor ./my-blueprint
mn blueprint run ./my-blueprint --detached

Resolve the reported model, service, input, and hardware prerequisites before launch. Local launch and doctor paths must begin with ./, ../, or /; other values select catalog IDs. Launch can prepare resources and perform the blueprint's external actions. --detached leaves execution running without the live workflow UI.

3. Inspect the run and its results

Keep both IDs returned by launch. Replace <job-id> and <run-id> below with those public IDs:

mn job show <job-id>
mn run show <run-id>
mn run watch <run-id>
mn run logs <run-id> --channel logs
mn run logs <run-id> --channel events
mn run result <run-id>

Ctrl+C detaches from the watcher. run result downloads outputs to $MN_HOME/outputs/<run-id> by default. Check the terminal state, warnings, artifacts, and required human review before using the result: completion is an execution outcome, not a guarantee of correct conclusions.

If an unfinished run should stop, use mn run cancel <run-id>. Cancellation cannot undo external actions already performed. When no other runs need the local services, stop them with mn runtime stop.

Continue with the full quickstart and CLI reference.

Execution and reliability boundaries

Runner Execution boundary
HostLocal Runs trusted worker code directly in the host execution environment.
DockerWorker Runs prepared commands in Docker containers; image, mounts, environment, and network access remain part of the contract.
OpenShell Runs workers in a sandbox governed by explicit policy, uploads, and network access.

Redis is the durable coordination store. Recovery can replay eligible work; external effects need idempotency or independent deduplication. Arbitrary process-local memory is not checkpointed, and exactly-once external effects are not guaranteed.

Federated nodes share a trust domain. Local deployment does not by itself ensure privacy: a blueprint can call configured remote providers and services. See Security and Reliability for deployment and replay boundaries.

Explore the ecosystem

Component Responsibility
MirrorNeuron Core — this repository Workflow execution, supervision, scheduling, persistence, runners, and gRPC services.
mn-cli Install-facing runtime controls and blueprint, job, run, model, and node commands.
mn-python-sdk Python integration, blueprint validation/compilation, bundle preparation, and runtime clients.
mn-api REST gateway and streaming surfaces.
mn-web-ui Browser-based runtime and workflow inspection.
mn-deploy Installation, Compose services, and release tooling.
mn-agents / mn-skills Reusable agents and Python skill packages.
Membrane Authoritative Markdown runtime memory and CPU DuckDB context retrieval through the SDK.
mn-system-tests Cross-component integration and system validation.

Membrane keeps complete runtime text and durable receipts as Markdown, with one disposable DuckDB index per job. It does not schedule tools or change the workflow DAG. The development Compose template also publishes authenticated Membrane gRPC on 127.0.0.1:${MN_CONTEXT_HOST_PORT:-50052} for host-native SDK responders; Docker clients retain the internal service address. Read Context memory and compression for that contract.

Develop Core

Core contributors need Elixir/Erlang compatible with mix.exs (Elixir ~> 1.16) and Redis for tests that exercise durable state. From a project directory:

git clone https://github.com/MirrorNeuronLab/MirrorNeuron.git
cd MirrorNeuron
mix deps.get
mix format --check-formatted
mix test
mix compile --warnings-as-errors
find scripts -name '*.sh' -print0 | xargs -0 -n1 bash -n

Integration tests may require Docker, OpenShell, Redis, or multiple machines. Use the development guide and testing guide for the relevant setup. To test changes across an editable workspace, follow local-mode installation.

config/                     Runtime configuration
lib/mirror_neuron/          Runtime, persistence, federation, and runners
lib/mirror_neuron_grpc/     gRPC handlers and generated bindings
proto/                      Public protobuf service contracts
tests/                      Unit, E2E, and API tests
scripts/                    Development and release helpers

Read AGENTS.md and SPEC.md before changing Core. Keep domain behavior in blueprints, agents, and skills, and update the canonical documentation when a public contract changes. See Contributing for the contributor workflow and RELEASE.md for Core distribution and release procedures.

Documentation

The detailed mn-docs index is the source for cross-component guides and references.

I want to… Read
Write a workflow Blueprint standard · DAG flow patterns
Integrate Python or HTTP Python SDK · REST API
Operate recurring or service work Schedules and events · Deployments
Add machines or model capacity Federation guide · Resources and devices
Diagnose a failure Troubleshooting · Environment variables

Security and license

Please use the repository's security reporting page for vulnerabilities; do not disclose them in public issues.

MirrorNeuron Core is released under the MIT License.

Response tool descriptions

Bounded response-agent tools accept an optional non-empty description of at most 1,200 Unicode code points after trimming. Core preserves this planner guidance without changing effects, argument validation, or execution authority. Deploy this validator with the matching SDK common and job-response packages before loading description-bearing blueprint declarations. Managed desktop Core requires its configured node identity at startup. The CLI persists that identity independently of the network address and reports an unnamed or mismatched Core as unready. After Wi-Fi changes, mn runtime reconnect refreshes endpoint advertisements without restarting healthy Core processes. Direct distributed-Erlang deployments retain their existing naming policy.

OpenShell artifact handoff v1

Set artifact_handoff: {"version": "mn.artifact_handoff/v1"} on an OpenShell executor to use owner-node immutable artifact commits. The SDK's mn_sdk.artifact_handoff helpers declare outputs and resolve verified inputs. This mode rejects sync_shared_storage: true; that flag explicitly selects the legacy whole-tree contract for unmigrated blueprints. Deploy the matching SDK in worker images when upgrading Core.

Core records dispatch before execution, captures diagnostics separately, seals and verifies declared files, and atomically publishes a receipt with their bytes before releasing step results. Transfer retries never dispatch the worker. A missing trustworthy execution record blocks replay; an operator must explicitly start a new run to retry an unknown outcome. No exactly-once model-call guarantee is claimed. Job sandbox caches are disposable and verified on every hit.

Receipts and transaction journals live under outputs/runs/<actual-runtime-run-id>/.handoff/ on the owning node. Startup maintenance reconciles published receipts and independently retries cleanup. Workflow replay can recover a committed executor result without the sandbox. Owner storage loss is outside this durability guarantee. Replica consumers must verify the receipt and every referenced file; filesystem replication alone is not a completion signal.

Child workflow start, plan-commit and round-complete events publish bounded topology_delta step/edge metadata for the newly active planner or task DAG. The projection excludes task inputs, outputs and artifact contents, allowing live monitors to discover runtime-created steps without transferring large blobs.

Running-time measurements

Run persistence maintains running_time (accumulated_ms, active_since, complete) in durable records and live projections. Co-worker analysis exposes these fields through the existing run JSON contract. Queue and pause intervals are excluded; an interrupted coordinator session retains observed time and marks unobserved recovery gaps incomplete. Historical records are not backfilled with wall-clock estimates. No new gRPC method or protobuf fields are required. Run projections expose run_data_ref with only the prepared submission ID, physical workflow run ID and syncthing storage identity. This binds numeric usage to the public execution while it is active, without exposing manifests, host paths, inputs or output content. Missing/invalid identities omit the reference; consumers must verify replicated run identity before reading usage.

Retry failed runs

PlanRunRetry verifies recovery eligibility; RetryRun submits the verified attempt/checkpoint selection with an idempotency key. These commands recover a failed run under its existing run and stable job IDs, with a new attempt and fenced lease epoch. ResumeRun continues paused work; it does not retry failures. The CLI and OtterDesk use these same Core commands.

Logical checkpoints retain completed outputs, unfinished inputs, child-workflow state and dynamic graph revisions. Recovery restores the ledger rather than seeding the entire workflow or restoring process memory. Missing inputs, changed bundles, corrupt artifacts, reset job data, active cleanup and uncertain external effects block dispatch with a reason. Executor/module replay requires a declared idempotency contract or a verified artifact handoff receipt; uncertain dispatches remain blocked. Logical effect keys survive retry while delivery IDs change. Dependency skips caused by a failed predecessor reopen on retry. Their reasons are read from verified checkpoint output artifacts as well as inline outputs; intentional branch skips and mapped-work placeholders remain preserved.

Only explicitly supplied fields declared in manifest metadata run_retry.configuration_fields may change. Fields declare type: integer with minimum/maximum, or type: string with allowed_values. Original inputs, workflow topology and result-defining configuration remain immutable. SDK contexts receive MN_RUN_RETRY_JSON with effective overrides and preserved active usage. An allowance of 60 minutes after 20 minutes of execution leaves 40 minutes; queued, paused and failed intervals do not add to the durable run clock. Executor invocations refresh their retry context from the fenced Core record after queue admission, so preparation and paused intervals are not charged by a stale startup context. A missing running lease blocks external invocation.

Failed control records, checkpoints and referenced bundles/artifacts are retained until explicit run/job deletion or job-data reset. Logs and delivery history keep their bounded retention. Resource status reports retry_checkpoint_storage. Historical files alone do not establish eligibility: planning must verify the remaining Core record and supported checkpoint. Run retry is manual only.

Oversized prompts use Membrane CPU preparation first, then optional original-ID selection on the existing LiteLLM default route. No dedicated compressor model is prepared. Managed SDK turns also retain source-backed working checkpoints.

Durable job backup transport

ExportJobBackup(JobRequest) streams bounded JobBackupChunk files from a quiescent stable job. RestoreJobBackup(stream JobBackupChunk) checks the complete mn.backup.v3 inventory and creates a fresh job definition with copied data. Both RPCs require v1 requests, client identity, and runtime identity, and are denied in network-only mode. Export follows the authoritative federated owner.

Snapshots include the archived bundle, persistent job data, retained run/event history, artifacts, payload blobs, current staging and retained historical staging. Start/resume is serialized with export. Source executions are preserved as evidence under .mn-restore/<source-job-id> rather than replayed; schedules receive fresh identities and remain paused. Restoring never overwrites an existing job. The SDK owns ZIP64 transport, full offline wheels/images/models, host compatibility and hardware preflight, and fresh native resource preparation.

The file protocol bounds chunks to 1 MiB, entries to 100,000 and uncompressed content to 128 GiB; path, symlink, duplicate, inventory and checksum failures reject restore. Earlier internal mn.backup.v2 snapshots are not accepted by these RPCs. Upgrade the generated SDK bindings and API/CLI adapters with Core.

HostLocal wheel capture and platform inspection use the existing native PrepareRuntimeModel transport with purpose=job_backup. Core forwards these Python requests without selecting a default model; the native SDK/CLI owns execution-platform wheel builds and fresh offline virtual environments.

Execution-node context service binding

DockerWorker binds Membrane endpoint, authentication, observability and serving counter settings from the execution node as one service configuration before launch. A submitter's credential cannot authenticate a different node's service. Missing execution-node endpoint/authentication fails before dispatch. Secrets stay in the trusted runner environment and are excluded from diagnostic output.

Run placement failures expose bounded measured resource blockers to the SDK. With updated Core and SDK services, callers receive the existing problem codes (for example MN_GPU_MEMORY_UNAVAILABLE / 2001), a friendly PC name, free memory versus required memory, and safe remediation. Admission requirements are unchanged. See SPEC.md for the versioned admission detail contract.

About

Multi AI Agent at scale

Resources

Stars

3 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages