Policy-gated control between an agent's intent and a real-world effect.
An agent proposes a tool call. REMORA is a governance overlay between that proposal and the protected tool: it evaluates the exact action before it runs, and returns one of four decisions:
| Decision | Meaning |
|---|---|
| ACCEPT | The action may proceed under the current policy and context. |
| VERIFY | A defined check or review is required before execution. |
| ABSTAIN | Required information is missing or insufficient. |
| ESCALATE | The action requires a higher-authority decision. |
The enforcing surface is POST /v1/execution/*. An approval is bound to the exact proposal it was given for and is consumed once, at the policy-enforcement point, before governed dispatch.
Status: research and shadow-mode software with no production certification. External replication is pending; every deployment still needs its own validation.
python -m pip install -e ".[dev]"
python -m remora tryOne ordered reading path. Every other document points back here.
- Developer handoff: the shortest technical path through the repository
- Architecture: components, data flow, module stability
- Execution quickstart: configure and run the enforcing path
- API reference: public interfaces with a
curlround-trip; wire contract inschemas/openapi.json - Python SDK: the one namespace with a backward-compatibility guarantee
- Evidence and claims: what each result establishes, and what it does not
- NEGATIVE_RESULTS.md: failed hypotheses and limitations, kept permanently
The documentation index lists the complete registered set.
Five steps, deliberately few.
- Authoritative context. Tool meaning, target, risk and approved intent come from deployment-owned sources such as a Signed ToolSpec, never from the calling agent.
- Policy decision. Deterministic hard guards outrank model-derived signals.
- Review or grant.
VERIFYandESCALATEenter bounded review;ACCEPTreceives a short-lived, single-use grant bound to the proposal. - PEP and dispatch. The grant is consumed, an
ExecutionLeasebinds policy identity to call identity, andGovernedToolDispatcherinvokes the deployment-owned callable. - Lifecycle and evidence. Dispatch intent, outcome, effect verification and audit records stay joinable by proposal identity.
What sits outside that path
The separate /v1/assess research surface can use oracle, evidence and uncertainty components. They are not prerequisites for the execution kernel and cannot override its deterministic hard-guard floor.
DEVELOPER_OVERVIEW.md draws the CORE / OPTIONAL / EXPERIMENTAL / HISTORICAL boundary; the machine-checked product truth contract classifies every capability; the execution quickstart walks a deployment-shaped example.
Headline values are governed by the claim register and stay tied to committed result artifacts and scope caveats. Every number below is bounded by documented assumptions, benchmark populations and evaluation protocols. None is a general safety guarantee.
| # | Benchmark | Result | Scope |
|---|---|---|---|
| 1 | AgentHarm | 0.0% wrongly allowed (0/208); 95% upper bound 1.81% | Intent classification; harmless-twin refusal 100.0% |
| 2 | Adversarial simulator | 0.0% unsafe runs (0/70 templates; 700 tasks); 95% cluster-level Wilson upper bound 5.2%; utility +0.456 | Simulated; effective N = 70 (70 templates × 10 cosmetic variants); unsafe-rate gap Δ=0.0143 vs. baseline, not statistically significant |
| 3 | BFCL v4 (C-ext3) | Native wrong-call acceptance 0.0% (0/500; Wilson 95% upper bound 0.76%); irrelevant-tool refusal 100.0% (300/300); required-input guessing 0.0% (0/398) | Sealed once, 2,799 episodes, frozen semantic bundle + authority floor; utility targets missed; see NEGATIVE_RESULTS §39 |
| 4 | Historical regression | 0.0% wrongly allowed (0/167) | Previously observed failures only |
How to read those numbers
- The deterministic simulator has 70 independent templates; cosmetic variants do not increase the effective sample size.
- BFCL v4 C-ext3 measures the safety axis only: declared semantic authority takes native wrong-call acceptance from 24/500 to 6 to 0 across the ablation arms. The utility side of the same run (read autonomy, argument routing) missed its pre-registered targets; those findings and the permanent 10.9% baseline live in NEGATIVE_RESULTS §39 and CLAIM-019's caveat, not here.
- Internal reproducibility is not external replication or field validation.
Replaced claims are kept in superseded claims. Failed hypotheses and limitations are kept in NEGATIVE_RESULTS.md.
python -m remora assess drop_database
python -m remora whatif drop_database # what would it take to ACCEPT? nothing a model can say
python -m pytest tests/ -q
python scripts/demo_industrial_maintenance.py
python -m remora doctorLibrary call, and why it is advisory
assess_tool_call(...) evaluates the context the caller supplies. It does not control downstream credentials. For enforcement, route the call through /v1/execution/* and make the governed dispatcher the only path to the protected tool.
from remora import assess_tool_call
assessment = assess_tool_call(
"drop_database",
{"db": "prod-main"},
risk_tier="critical",
action_type="destructive_write",
)For external review or pilot work, set REMORA_RUNTIME_PROFILE=review or controlled_pilot. Those profiles require the Signed ToolSpec and durable-state prerequisites described in the execution quickstart.
It cannot enforce against a credential path that bypasses its dispatcher. Deployment identity, credential custody, downstream authorization and operational controls stay with the deployment. Capability maturity is tracked, per capability, in the capability register.
Repository map and current engineering state
| Path | Purpose |
|---|---|
remora/policy/ |
Policy observation and decision logic |
remora/toolcall/ |
Tool authority, ToolSpec and routing contracts |
remora/enforcement/ |
PDP-to-PEP grant, lease, dispatcher and outbox |
remora/governance/ |
Review, lifecycle, audit and effect verification |
servers/ |
HTTP surfaces and deployment wiring |
tests/ |
Deterministic regression and contract tests |
docs/assurance/ |
Claims, capabilities, release gates and review records |
docs/research/ |
Active research material and benchmark design |
docs/archive/ |
Superseded or historical material |
paper/ |
Research paper and supporting publication artifacts |
Research modules, AROMER, the older statistical-physics work and historical design documents stay in the repository for reproducibility and audit history. Their presence does not make them part of the enforcing runtime path; the capability register says what is wired.
The state of the servers/execution_api.py decomposition (issue #241, closed) and of the decoupled dispatch worker behind REMORA_ASYNC_DISPATCH (issue #82, closed; not enabled in any reference profile) is stated once, in DEVELOPER_OVERVIEW.md.
For product-oriented integration see Assured Agent Execution, which consumes pinned REMORA artifacts rather than copying the governance core.
The research-to-control mapping lives in docs/research/research_control_matrix.generated.md; related work in docs/09-related-work.md.
Generative-AI tools were used during development. AI-generated text or code is not evidence by itself; claims must resolve to committed artifacts, tests or verified sources. Disclosure: docs/AI_USE.md.
Citation metadata is in CITATION.cff; GitHub renders it as BibTeX from the "Cite this repository" panel.
REMORA versions from v0.10.0 are source-available under the Business Source License 1.1, with commercial licensing under the REMORA Commercial License. Research, benchmarking and reproducibility work are permitted within the terms in LICENSING.md.
Contribution requirements, branch lifecycle, documentation style and claim hygiene are defined in CONTRIBUTING.md and docs/10-contributing.md.
