The MongoDB version of SQLite.
The real MongoDB engine, embedded directly in your application.
One process. One local directory. MongoDB's own queries, aggregations, commands, and WiredTiger storage.
No server deployment. No ports. No connection string.
Warning
Experimental. This project was created by Jeroen Vervaeke. It is not production-ready or supported by MongoDB.
Not a MongoDB-compatible reimplementation. MongoDB's actual server code runs inside your process.
- π Real MongoDB execution β queries, cursors, aggregation pipelines, commands, and storage run through MongoDB's own engine, backed by WiredTiger.
- π¦ SQLite-like deployment β open a local directory; no database server to install or manage.
- π¦ Rust-native, with bindings β work with typed Serde collections or raw MongoDB
documents, or reach the same engine from another language.
- π Python β
pymongo_embedded.MongoClient("mongodb_embedded://./data"), PyMongo's own client over the in-process engine. - π¨ Node.js β
@0q/embedded-mongodb, the MongoDB Node.js driver over the same engine, and the URI scheme that mongosh understands.
- π Python β
- πΎ Persistent storage β clean close and reopen cycles preserve data in the supplied directory.
- π§΅ Parallel access β share one client across threads; commands run in parallel over a pool of sessions, fixed at a chosen size or grown and shrunk on demand between a floor and a ceiling.
- π Automatic IDs β missing
_idfields receive anObjectId, matching the official drivers.
The directory uses MongoDB's native on-disk format. After a clean shutdown, hand it between
embedded-mongodb and the matching pinned mongod build; only one process may own it at a time.
Technical implementation
Commands travel from bson::Document through the safe Rust API, CXX bridge, DBDirectClient,
ServiceEntryPointShardRole, MongoDB command/query/catalog code, and finally WiredTiger.
MongoDB is pinned as the shallow mongo/ submodule. The embedded-mongodb-sys crate owns the
native implementation, CXX bridge, and build script; the safe Rust crate builds BSON helpers on
top. Startup uses a 256 MB WiredTiger cache plus a 64 MB spill cache.
Important
Directories written by a build published before this fix carry damaged indexes. Opening one with a current build repairs it automatically, moving rather than deleting any document it has to evict.
The defect. Starting the storage engine filled MongoDB's collection catalog and nothing
else β DatabaseHolder::openDb was never called for a database that already existed on disk,
so every collection loaded from such a directory came back with an empty in-memory index
catalog. Writes after that reopen went into the record store and into no index at all, _id_
included: a duplicate _id was accepted and both copies stayed, and documents written that way
are invisible to any query answered from an index while a collection scan still returns them.
Two counts of the same collection could disagree depending on the plan chosen.
Which directories. Any directory that was written to after being reopened, by a build from before this fix. A directory that was only ever written to in the session that created it is sound, and so is every directory a current build creates. Only the collections written to after a reopen are affected; the rest of the directory is untouched.
The repair. Client::new checks a directory it did not create for missing index entries and
runs the engine's own validate {repair: true} over any collection that has them. It happens
once: a directory that has been through the pass carries a .embedded-mongodb-index-repair
marker and is not checked again, and a directory created by a current build is marked without
being checked at all. Everything it repairs is reported through tracing at WARN, naming the
collection, how many index entries were inserted, how many documents moved, and where they went.
Every binding opens through that same Client::new β MongoClient("mongodb_embedded://β¦") in
Python and NativeBridge.open on Android both go through it β so the pass, the marker and the
skip variable below behave identically whichever language opens the directory. Neither binding
installs a tracing subscriber, though, so the WARN records go nowhere unless the host
application has one: there, the marker and the repaired collections are what to look at.
The check is a full validation of every collection in the directory, so the first open after upgrading is slower than the ones after it. That is the trade: one scan against silently wrong query results.
Evicted duplicates are moved, not deleted. Where two documents ended up sharing an _id,
the index can only hold one of them. The other is moved into
local.lost_and_found.<collection UUID>, and the WARN record names that collection. Read it
back like any other:
let evicted = client.database("local").run_command(&doc! {
"listCollections": 1,
"filter": { "name": { "$regex": "^lost_and_found\\." } },
})?;One thing the repair can delete. validate {repair: true} is the engine's general-purpose
repair, not one written for this defect alone. If it meets a record whose BSON cannot be read β
unrelated corruption, not anything this defect produces β it removes that record, and there is no
lost and found for those. The pass reports any such deletion in a WARN of its own, naming the
collection and the count. To look before anything is touched, set the variable below for one
open and run validate without repair yourself.
Skipping it. Set EMBEDDED_MONGODB_SKIP_INDEX_REPAIR to 1, true, yes or on to leave
the check out. Any other value leaves it on β no and off included, deliberately, so that a
value nobody meant as yes cannot quietly switch off a repair. Skipping does not write the marker,
so it suppresses the pass rather than cancelling it: the next open without the variable set still
checks the directory.
Forcing it. Delete .embedded-mongodb-index-repair from the directory and the next open
checks it again. Worth knowing if a directory has been back to an older build since β the marker
records that a check happened, not which engine wrote the data afterwards.
Doing it by hand. The pass runs nothing you cannot run yourself. Per collection:
let report = client.database("shop").run_command(&doc! { "validate": "orders" })?;
// report.valid, report.errors, report.missingIndexEntries
let repaired = client.database("shop").run_command(&doc! { "validate": "orders", "repair": true })?;
// repaired.numInsertedMissingIndexEntries, repaired.numDocumentsMovedToLostAndFoundvalidate {repair: true} is idempotent, so running it against a sound collection changes
nothing. Note that its reply reports the state it found: a collection that is sound afterwards
can still come back with valid: false in the same reply that says repaired: true. Validate
again to see the result.
Open a directory, insert a document, and query it back. The crate-root API is async β every
command is dispatched to a pool of dedicated engine threads (one per session, eight by
default), so an .await parks a task, never a runtime thread:
use embedded_mongodb::bson::doc;
let client = embedded_mongodb::Client::new("./data").await?;
let items = client.database("app").collection("items");
let inserted = items.insert_one(doc! { "name": "embedded" }).await?;
let item = items.find_one(doc! { "_id": inserted.inserted_id }).await?;
println!("{item:?}");The same API exists synchronously as embedded_mongodb::blocking for callers without an async
runtime β the Python and Android bindings go through it:
let client = embedded_mongodb::blocking::Client::new("./data")?;mongod sizes its journal, its cache and its free-space floors for a server. The journal is the one that is badly wrong inside a phone application: at mongod's settings an Ireland-scale directory holding 2.25 MiB of documents and indexes occupies 202 MiB, because two journal files are allocated in full whether or not anything is written to them. This engine defaults to one 8 MiB journal file and no pre-allocated spare, which takes the same directory to 10.25 MiB. Cold open gets faster too, because recovery scans the journal at startup and there is 25x less of it to scan.
Journalling itself is untouched: the files are smaller and there is one rather than two, but
every write is logged and fsynced exactly as before. tests/durability passes in full, and the
two properties it exists to pin are unchanged β a write acknowledged under {w:1, j:true}
survives SIGKILL (zero lost, measured), and recovery replays a strict prefix of the write
history rather than a torn one (no gaps, measured). Writes acknowledged without j: true
still lose the tail written since the last journal flush, which mongod performs every 100 ms.
That window did not widen: over ten killed runs each, the tail lost was 43-511 writes at
mongod's journal settings and 114-464 at these β overlapping ranges, with the worst single
run belonging to mongod's settings.
The cache and the free-space floors are left where they were, and all five are settable:
| Limit | Default | Set through |
|---|---|---|
| Journal file size | 8 MiB (mongod: 100 MiB) | Client::with_options |
| Journal pre-allocation | off (mongod: on) | Client::with_options |
| WiredTiger cache | 256 MB | Client::with_options |
| Free disk to start an index build or spill a query | 500 MB, as mongod | Client::with_options, or Client::process_limits at any time |
| Parallel commands (session pool) | 8, fixed | Client::with_options |
The cache figure is the value this engine has always used, and is also the floor mongod will not go below on a server; mongod's default is half of system memory above the first gigabyte, which is not a number that means anything on a phone. It is a ceiling WiredTiger grows into rather than memory it takes, and a cold read-only process at Ireland scale peaks well under it, so it is exposed for tuning rather than because the default is wrong.
use embedded_mongodb::{Client, Concurrency, FreeDiskFloor, JournalFileSize, OpenOptions};
let options = OpenOptions::new()
.journal_file_size(JournalFileSize::from_kibibytes(2048)?)
.free_disk_floor(FreeDiskFloor::from_mebibytes(32)?)
.concurrency(Concurrency::from_count(4)?);
let client = Client::with_options("./data", options).await?;Concurrency::from_count fixes the pool at a size chosen up front. A caller that cannot predict
its own concurrency can name a range instead, and let the pool find it:
use std::time::Duration;
use embedded_mongodb::{Client, Concurrency, OpenOptions};
let elastic = Concurrency::dynamic(1, 32, Duration::from_secs(60))?;
let client = Client::with_options("./data", OpenOptions::new().concurrency(elastic)).await?;That opens one session, opens more β up to 32 β whenever a command arrives with every session busy, and closes the extras once they have sat idle for a minute. The ceiling is then a limit rather than a cost: nothing is paid for concurrency that never happens. The async client's worker threads follow the same policy, one per session, so the two scale together.
Anything left unset keeps the engine's own default, so Client::new(path) and
Client::with_options(path, OpenOptions::new()) open identically.
Keeping that promise for the free-space floor takes work rather than nothing, and the reason is
worth knowing. Both floors are MongoDB server parameters, and this engine keeps one runtime
for the whole life of the process, so a floor belongs to the process rather than to the
Client that named it: it outlives that client's close, and left alone it would still be in
force at the next open. Every open therefore establishes the floor β the caller's if named,
MongoDB's own otherwise, read from the engine before anything moved them. An application that
opens one database on a lowered floor, closes it and opens another gets the defaults it asked
for rather than the first database's floor. Two consequences: a floor moved with
client.process_limits().set_free_disk_floor(..) on a running client lasts only until the next
open, and while a client is open the floor is shared by every database name it serves.
The handle is where the scope is named, because that is the scope the two floors have:
let limits = client.process_limits();
limits.set_free_disk_floor(FreeDiskFloor::from_mebibytes(32)?)?;
let now = limits.free_disk_floors()?;The free-space floor is the one worth thinking about before lowering. It is what stops an index build or a spilling query from starting when the device is nearly full β and nothing stops one that runs out part-way: WiredTiger answers a full disk by panicking, which takes the host process down without an error reaching the caller. How much headroom is enough depends on how much data is about to be indexed, which is why the default is left where MongoDB put it.
Explore the complete runnable examples:
The pymongo-embedded package keeps PyMongo's API for remote servers and routes embedded URIs to
the in-process engine:
from pymongo_embedded import MongoClient
remote = MongoClient("mongodb://localhost:27017/")
local = MongoClient("mongodb_embedded://./data")
local.app.items.insert_one({"name": "embedded"})mongodb+embedded://./data is the URI-valid spelling and behaves identically.
AsyncMongoClient makes the same substitution over PyMongo's asynchronous classes:
import asyncio
from pymongo_embedded import AsyncMongoClient
async def main():
client = AsyncMongoClient("mongodb_embedded://./data")
try:
await client.app.items.insert_one({"name": "embedded"})
print(await client.app.items.count_documents({}))
finally:
await client.close()
asyncio.run(main())No command ever occupies the event loop: the engine's own worker threads do the blocking, and
awaiting costs a parked task. Commands run in parallel over the same eight sessions, so
asyncio.gather over eight of them takes about as long as one.
Cancelling really cancels. Dropping the future stops the engine rather than only stopping the wait, so a timeout bounds the work and not merely your patience, and the session goes back to the pool at once rather than when the query would have ended:
try:
await asyncio.wait_for(cursor.to_list(), timeout=0.5)
except TimeoutError:
... # the engine has stopped; the session is already freeThe constructor is the one thing that blocks, deliberately. Opening does storage recovery and the
one-time index repair scan, the longest block this library ever does and not something an event
loop should be running. Construct before the loop starts, or reach for
await asyncio.to_thread(AsyncMongoClient, uri) inside one.
Only one embedded engine may be open per process, synchronous or asynchronous. Opening a second
says only one embedded MongoDB runtime may be open per process; closing the first releases it.
Run the included examples -- synchronous and asyncio -- with:
./scripts/python
./scripts/python examples/python/asynchronous.pyThe runner creates a clean environment under .cache/python, builds incrementally, and uses the
sibling mongo-python-driver checkout. It also accepts normal Python arguments:
./scripts/python -i # Open a Python shell.
./scripts/python your_script.py
./scripts/python -m pip install another-packageSet PYMONGO_SOURCE if the PyMongo checkout is elsewhere.
The binding's tests come in two kinds. test_*.py are behaviour -- CRUD and cursors through
both clients, close and reopen, the one-engine-per-process rule -- and CI runs them against the
wheel it builds. measure_*.py assert on wall-clock ratios instead: how much faster eight
commands are than one, how much of the interpreter another thread got while a close waited,
whether a cancelled task really freed its session. Those want a machine with cores to spare, so
they are run by hand:
./scripts/python -m unittest discover -s python/tests -p 'test_*.py'
./scripts/python -m unittest discover -s python/tests -p 'measure_*.py'To build a distributable wheel containing the native engine:
python -m pip install "maturin[patchelf]"
maturin build --release
python -m pip install target/wheels/pymongo_embedded-*.whlThe first build downloads the published engine; see Build and test for the alternatives.
This package is not published to PyPI. The wheel exists so the engine can be vendored into one,
not as a distribution -- but it does install, and CI installs it: PyMongo 4.18 is on PyPI, so
nothing but pymongo-embedded itself is missing from there. ./scripts/python wires up a
mongo-python-driver checkout instead, which is what you want while changing the binding.
The binding supports PyMongo 4.18 commands through both clients, including normal CRUD, cursors, aggregations, and bulk document sequences. Authentication, TLS, compression, sessions, transactions, change streams and exhaust cursors are not supported.
Threads and tasks both run in parallel. Every connection PyMongo hands out reaches the same
engine, which runs commands over a pool of eight sessions, so eight of them issue eight
commands at once rather than queueing behind one another. PyMongo's own maxPoolSize is the
other ceiling, and a fan-out gets the smaller of the two.
@0q/embedded-mongodb runs the MongoDB Node.js driver against the in-process engine, the way
pymongo-embedded does for PyMongo:
const { MongoClient } = require('@0q/embedded-mongodb');
const client = new MongoClient('mongodb_embedded://./data');
const items = client.db('app').collection('items');
await items.insertOne({ name: 'embedded' });
console.log(await items.findOne());
await client.close();MongoClient is the driver's own class: an address it does not recognise is handed to the
driver untouched, so one class serves both a server and a directory on disk. Both spellings of
the scheme work, mongodb_embedded:// and mongodb+embedded://, with a relative or an absolute
directory after them, which is created if it does not exist.
The driver is not modified and nothing of its internals is reached into. It speaks the wire protocol to a socket, so the package gives it one: a listener on a Unix socket in a private temporary directory, inside the same process, that hands every message it reads to the engine and writes the reply back. The lower layer is available on its own for a caller that already has a driver -- mongosh is one:
const { open } = require('@0q/embedded-mongodb');
const embedded = await open('./data');
const client = new MongoClient(embedded.uri); // The driver's MongoClient, unchanged.
// ...
await client.close();
await embedded.close();open returns a promise because opening does storage recovery and the one-time index repair
scan; openSync blocks instead, for callers with no loop to await on, and is what the
MongoClient constructor uses. Only one embedded engine may be open per process; opening a
second says only one embedded MongoDB runtime may be open per process, and closing the first
releases it.
The handshake is answered by the package rather than the engine, as the Python binding does,
and for the same two reasons: the engine's own hello advertises sessions, which a direct
client cannot use, and a topologyVersion, which would have the driver's monitor park an
engine strand on an awaitable hello every ten seconds. Everything else goes to the engine as
sent. Connections run their commands in parallel over the engine's strand pool, up to the
driver's maxPoolSize and the engine's strand count, whichever is smaller.
The package ships the addon and the engine in platform packages, the way napi-rs packages do,
so npm install @0q/embedded-mongodb fetches only the one for the machine it runs on. To build
and test it from a checkout:
cd embedded-mongodb-node
npm install
npm run build # cargo build through napi-rs, then the engine copied beside the addon
npm testPublishing is the publish-node workflow, dispatched by hand: it builds and tests on the three
platforms, then publishes the platform packages and the root through npm's trusted publishing,
with a provenance attestation and no token. The same steps by hand are npm publish --access public in each npm/<platform> directory and then in the package root.
Authentication, TLS, compression, sessions, transactions, change streams and exhaust cursors are not supported, and neither is Windows.
A fork of mongosh accepts the same URIs, and is
published as @0q/mongosh so nothing needs installing:
npx @0q/mongosh mongodb_embedded://./dataIt opens the directory through this package and talks to it over the driver it already
carries, so show dbs, db.items.find() and the rest work as they do against a server.
The embedded deployment model creates a path toward:
- π₯οΈ Local-first applications with MongoDB data stored beside the app.
- π§° Self-contained developer tools and CLIs without a database service to provision.
βοΈ Offline and edge workloads that keep working without network access.- π§ͺ Tests and demos that start with the application instead of waiting for infrastructure.
- π One engine across ecosystems, with Rust today and potential Python and JavaScript bindings.
- Lifecycle and persistence β open, clean close, reopen, and persistence on disk.
- Documents and typing β BSON and Serde-backed values, generated
ObjectIds,insert_one, andinsert_many. - Queries and cursors β filtered
find,find_one, array matching, comparison operators, and batched cursorgetMore. - Commands β
pingthrough the public BSONrun_commandAPI. - Aggregation β a multi-stage product report using
$match,$unwind,$group, arithmetic,$sort, and$project. - Concurrency and errors β
Client: Send + Sync, concurrent inserts, and structured duplicate-key errors. - Observability β MongoDB-to-
tracingseverity mapping. - Engine identity β
buildInfo, theserverStatussection and theembeddedMongodbcommand all report the embedded build and agree with one another. - Index repair β a data directory a pre-fix engine damaged is checked in as a fixture, and the repair pass is held to repairing it, to running once, to leaving a healthy directory alone, and to moving rather than deleting a duplicate.
One end-to-end integration test covers the operations demonstrated by all three examples.
- Wider command set β commands beyond those exercised by the covered helpers and
ping. - Cursor cancellation β early cursor drop and its
killCursorscleanup path. - Advanced MongoDB features β authentication, replication, transactions, change streams, TTL, backup, and encryption.
- Failure and scale β crash recovery, stress tests, and large-data workloads.
- Portability and upgrades β automated
mongodhandoff, MongoDB upgrades, and Windows.
MongoDB logs are emitted through tracing under the embedded_mongodb::mongo target. Events carry
the MongoDB ID, component, context, severity, and lossless JSON record; open, command, and close
operations add spans. Without a subscriber the library remains silent. The basic example installs
tracing-subscriber.
Nothing can attach a shell or Compass to an in-process engine, so the engine says what it is
through three ordinary commands. Use any of them to assert in a test that you are running
against the embedded engine rather than a real mongod:
let build_info = client.run_command("admin", &doc! { "buildInfo": 1 })?;
// modules = ["embedded"]
// buildEnvironment = { β¦, embedded: "true", embeddedAuthor: "Jeroen Vervaeke" }
let status = client.run_command("admin", &doc! { "serverStatus": 1 })?;
// status.embedded = { embedded: true, author, repository, mongoVersion }
let about = client.run_command("admin", &doc! { "embeddedMongodb": 1 })?;
// the same payload, as a command of its ownmodules is the cheapest check; a real mongod never reports an embedded module.
gitVersion is deliberately left alone β it reports the MongoDB commit the engine was built
from, and the pinned commit is recorded in NOTICE.
Criterion measures open, insert_one, find_one, and close separately:
cargo bench --bench operationsThe engine is not compiled locally. cargo build downloads the library published for the
current target, checks it against a SHA-256 committed in
embedded-mongodb-sys/prebuilt.rs, and links that:
cargo test --all-targetsNo submodule, no Bazel, no C++ toolchain. The download is cached outside the target
directory β under $XDG_CACHE_HOME/embedded-mongodb, or ~/Library/Caches/embedded-mongodb
on macOS β so cargo clean does not throw it away.
There are three modes, and the first match wins:
EMBEDDED_MONGODB_NATIVE_LIB_DIR=<dir>β use thelibembedded_mongodb_native.soin<dir>, unconditionally. This is the answer for hermetic builds, air-gapped machines and distribution packaging.EMBEDDED_MONGODB_BUILD_FROM_SOURCE=1β compile the engine from the pinned submodule. Hours, and about 13 GB of disk.- Otherwise, download the library published for this target.
EMBEDDED_MONGODB_CACHE_DIR moves the download cache, and BAZEL and
EMBEDDED_MONGODB_BAZEL_JOBS apply to a source build. There is no way to skip the checksum.
Two things to know before putting this behind a firewall. A plain cargo build now reaches
both github.com and release-assets.githubusercontent.com, so a proxy allowlist naming
only the first still fails. And cargo build --offline does not suppress the download β
cargo does not pass that flag to build scripts β while CARGO_NET_OFFLINE=1 does, and turns
a missing cache entry into an error naming the file it wanted.
The published Linux libraries are built on Ubuntu 24.04 and therefore need glibc 2.39 or
newer; a source build has no such floor. Rather than failing at load time, build.rs
compares the requirement against the host and stops the build with the remedy.
Prebuilt libraries are published for x86_64-unknown-linux-gnu,
aarch64-unknown-linux-gnu, aarch64-apple-darwin, aarch64-linux-android and
x86_64-linux-android. Anything else β an Intel Mac, musl, a BSD β builds from source, which
a later section covers. Intel macOS is absent because that runner could not finish a build
inside GitHub's six-hour job limit, and GitHub retires the image in August 2027 regardless.
Both 64-bit ABIs are published, compiled against bionic at API level 24 β Android 7.0 β and verified to load, open a database and answer commands on an API 24 device. 32-bit Android is not supported: MongoDB builds only for 64-bit platforms.
The embedded-mongodb AAR sets minSdk 26 even so, because org.bson reaches for
java.time and that arrived in API 26; android/README.md has the detail. Rust callers
linking the engine directly are not bound by that and can target 24.
Cargo has to be told which toolchain to use. cc, which the cxx bridge and
link-cplusplus run, carries no NDK of its own and looks for a <triple>-clang++ the NDK
does not ship:
ndk=$ANDROID_NDK_HOME/toolchains/llvm/prebuilt/linux-x86_64/bin
export CC_aarch64_linux_android=$ndk/aarch64-linux-android24-clang
export CXX_aarch64_linux_android=$ndk/aarch64-linux-android24-clang++
export AR_aarch64_linux_android=$ndk/llvm-ar
export CARGO_TARGET_AARCH64_LINUX_ANDROID_LINKER=$ndk/aarch64-linux-android24-clang
cargo build --release --target aarch64-linux-androidcargo-ndk sets the same variables, if you would rather not.
Ship two files with the application: libembedded_mongodb_native.so and the NDK's
libc++_shared.so. The engine links its own C++ runtime statically and exports only the six
extern "C" entry points, but the bridge compiled into the Rust crate uses the NDK's default
shared runtime, as any other NDK library in the same application does.
Neither Android library gets link-time optimization β those flags are GCC- and ld.bfd-
specific, and the NDK ships neither β so both land near 47 MB against the x86_64 Linux
build's 34 MB. They do get --gc-sections, identical code folding and the version script.
A source build needs ANDROID_NDK_HOME or ANDROID_NDK_ROOT pointing at an NDK r27 or
newer, and EMBEDDED_MONGODB_ANDROID_API overrides the API level. The NDK's clang
cross-compiles both ABIs from any host, so no Android hardware is involved in building one.
Needed only to change the engine itself. MongoDB's pinned build requires Python 3.13, Bazel, C++20, and a supported compiler and linker. Its build documentation currently lists GCC 14.2 or Clang 19.1 and roughly 13 GB of free space.
git submodule update --init --depth 1
./scripts/apply-mongo-patches
cd mongo
python3.13 buildscripts/install_bazel.py
export PATH="$HOME/.local/bin:$PATH"
cd ..
EMBEDDED_MONGODB_BUILD_FROM_SOURCE=1 cargo test --all-targets --release--release is not optional once the patches are applied. Patches 0003, 0004 and 0006 leave
dangling references to the code they remove, and only the release build's link-time
optimization and --gc-sections eliminate them. A debug build links β nothing checks for
undefined symbols there β and then fails to load with undefined symbol: mongo::executor::makeNetworkInterface or similar.
embedded-mongodb-sys/build_native.rs holds the Bazel invocation and rebuilds incrementally
when the native sources change. Cargo-run tests and examples find the library through the sys
crate's build output; standalone binaries still need it in the platform loader path.
Changing anything the published library was built from β the submodule pin, patches/,
embedded-mongodb-sys/native/ or build_native.rs β makes the published library stale.
build.rs detects that and refuses to use it, rather than linking an engine that no longer
matches the source beside it. Publish a new one with gh workflow run native.yml --ref <branch> -f publish=true, which builds every target and commits the regenerated manifest.
That comparison is against a commit, so it needs the history that holds it: clone this
repository in full, or git fetch --unshallow a shallow one. A shallow checkout that cannot
reach that commit is refused outright rather than built against a library nothing has
checked.
Sizes of the published libraries, as recorded in the manifest:
| target | bytes |
|---|---|
x86_64-unknown-linux-gnu |
34,525,944 |
aarch64-unknown-linux-gnu |
37,387,448 |
aarch64-apple-darwin |
72,745,856 |
macOS is roughly twice the size because native/BUILD.bazel gates link-time optimization and
the version script on @platforms//os:linux, so it gets neither.
The release build is size-optimized rather than speed-optimized: -Os, link-time optimization,
per-function and per-data sections with --gc-sections, packed relative relocations, only the
six extern "C" entry points exported, and no TLS, gRPC, OpenTelemetry or enterprise modules.
Run ./scripts/apply-mongo-patches before building. It trims the embedded ICU collation tables
(2.6 MB), removes the slot-based execution engine so queries run on the classic one (4.7 MB), the
replication implementation the embedded server never uses (1.1 MB), the sharding runtime (2.5 MB)
and the network stack (0.9 MB), and fixes an assertion that aborted the host process on the first
hello a driver sends. See
docs/native-size-reduction.md for the measurements and what
further reduction would cost. Packed relative relocations require glibc 2.36 or newer.
LTO is linked with ld.bfd, which must be on PATH; lld cannot read GCC's IR. The C++ runtime
is linked statically, which costs a couple of megabytes and keeps a published library's
GLIBCXX requirement from following whichever toolchain happened to build it.
- MongoDB server internals are private and change frequently. Each submodule update can require lifecycle and Bazel dependency work.
- Many server components assume one global runtime. Multiple simultaneous
Clientvalues or different active database directories are rejected. - Commands issued through one
Clientare thread-safe and run in parallel up to the session pool's ceiling (OpenOptions::concurrency, default 8 fixed sessions); commands beyond it wait for a session to come free. - There is no process isolation: a MongoDB fatal invariant, memory fault, or abort terminates the Rust host.
- Authentication, replication, transactions, change streams, TTL, backup, encryption, and the wider command set are outside the current supported scope.
- Prebuilt libraries cover x86_64 and aarch64 Linux and aarch64 macOS. Intel macOS builds from source: that runner could not finish inside GitHub's six-hour job limit, and the image retires in August 2027. Windows is untested; see issue #9.
- This project embeds MongoDB Community Server as a modified work first published on 2026-07-27
and licensed as a whole under SSPL-1.0. Distribution or service use requires a license review;
see
LICENSE.
Production use would require owning a MongoDB fork, a narrow supported command matrix, crash testing, and upgrade work.
