Skip to content

bug(sync): engram sync hangs indefinitely on Windows with a large multi-project database β€” reproduces on 1.20.0 and 2.0.0-rc.11Β #1175

Description

@hjagar

πŸ“ Bug Description

engram sync (local export, no --cloud) hangs indefinitely on a machine with a large, long-lived engram.db shared across many projects. The process prints only the initial "Exporting memories..." line and never returns. It is not a classic I/O deadlock: Get-Process shows CPU time actively climbing throughout, but the engram.db-wal file size stays static for minutes at a time β€” active computation with no visible forward progress.

Reproduced independently on both 1.20.0 and 2.0.0-rc.11. The 2.0.0-rc.11 build already contains the projectsΓ—rows repair-scaling fix from #858 / PR #912 (verified via gh api repos/.../compare/v2.0.0-rc.11...<PR#912 merge commit> β†’ "behind", i.e. the fix commit is an ancestor of the tag), so this does not appear to be fully explained by that fix β€” either a residual scaling issue in the same area, or a separate bottleneck elsewhere in the sync export path.

πŸ”„ Steps to Reproduce

  1. Accumulate a ~/.engram/engram.db over an extended period across many projects (in our case: ~30 enrolled/foreign cloud sync targets β€” confirmed via engram doctor's sync_target_closed_space check β€” and several hundred thousand total observations rows; the live DB was ~1.13 GB before a schema migration and ~2.5 GB after).
  2. cd into one project's directory β€” in our reproduction, a project that was never cloud-enrolled (engram cloud upgrade doctor --project <name> returns reason_code: cloud_config_error, "cloud configuration is required before upgrade bootstrap").
  3. Run engram sync (no flags).
  4. Observe: it prints Exporting memories for project "<project>"... and never returns.

βœ… Expected Behavior

engram sync for one project's local export should complete in a reasonable, bounded time (seconds), scaling with that project's own new/changed rows β€” not with the total size or project count of the shared engram.db, and not with cloud-enrollment state of unrelated projects.

❌ Actual Behavior

  • On 1.20.0: reproduced twice in a row in the same session. Each attempt hung ~6 minutes with zero output before being killed manually (TaskStop/taskkill). engram sync --status (read-only) responded instantly and correctly both before and after each hang, so only the export path (sync with no flags) is affected.
  • On 2.0.0-rc.11: after upgrading, the first invocation of any command triggered a one-time schema migration (~8 minutes, engram.db-wal growing to ~1.67 GB, CPU actively climbing β€” this part completed successfully, exit code 0, and is a separate expected cost, not part of this report). Immediately after, a fresh engram sync attempt for the same previously-hanging project ran for 16+ minutes with Get-Process CPU time continuously increasing (92s β†’ 421s β†’ 608s β†’ 952s across the observation window) while engram.db-wal size did not change at all for multi-minute stretches, before being manually terminated.
  • No error, stack trace, or partial progress indicator is ever printed. The only way to stop it is to kill the process externally.

πŸ–₯️ Environment

Operating System: Windows

Engram Version: Reproduced on both 1.20.0 and 2.0.0-rc.11

Agent / Client: Claude Code

πŸ“‹ Relevant Logs

$ engram sync
Exporting memories for project "<project>"...
[hangs β€” no further output; process killed externally after 6-16+ minutes across multiple attempts]

$ engram sync --status
Sync status:
  Local chunks:    17
  Remote chunks:   15
  Pending import:  0
[returns instantly, both before and after each hung attempt β€” unaffected]

# CPU time (PowerShell Get-Process -Id <pid>).CPU, seconds, on v2.0.0-rc.11 attempt:
t+0s:    106.20
t+~5min: 421.86
t+~6min: 607.86
t+~11min: 937.13
t+~11.5min: 952.42
# engram.db-wal size stayed at exactly 1665184552 bytes across the entire t+5min..t+11.5min window

πŸ’‘ Additional Context

  • Closely related to bug(sync): enrolled repair still scales by projects x rows before first syncΒ #858 / PR fix(store): optimize enrolled mutation repair (#858)Β #912 ("enrolled repair still scales by projects Γ— rows before first sync"), but that fix's commit is confirmed present in the 2.0.0-rc.11 build we tested, and the hang still reproduces β€” at a scale (~30 projects, hundreds of thousands of rows) beyond what that issue's own benchmark matrix covered (worst case there was 100 projects / 100K rows). This suggests either a residual multiplicative cost not covered by the fix(store): optimize enrolled mutation repair (#858)Β #912 fix, or a different bottleneck in the sync export path itself (as opposed to the pre-sync repair step bug(sync): enrolled repair still scales by projects x rows before first syncΒ #858 targeted).
  • The affected project in our reproduction was never cloud-enrolled, ruling out per-project cloud state as the direct cause β€” the cost appears tied to the shared database's overall size/project count rather than the invoked project's own data.
  • We did not have access to sync.go's source to pinpoint the exact loop; happy to provide the full engram doctor --json output, exact DB statistics, or run further instrumented reproduction steps if useful.
  • Both the database and the previous binary were backed up before any of this investigation, so we have a clean before/after state if a maintainer wants specific before/after comparisons run.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions