Skip to content

fix: reconnect stored WhatsApp sessions after object restarts - #5

Closed
jlucaso1 wants to merge 2 commits into
mainfrom
fm/cloudflare-auto-reconnect-free-tier-r1
Closed

jlucaso1 wants to merge 2 commits into
mainfrom
fm/cloudflare-auto-reconnect-free-tier-r1

Conversation

@jlucaso1

@jlucaso1 jlucaso1 commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

What Changed

  • Restore registered WhatsApp sessions when a new Durable Object instance starts. Unpaired sessions still require /start.
  • Retry closed connections through /status or /start after a 30-second in-memory cooldown, and share concurrent connection attempts. Add regression coverage for restoration, retries, and failed starts.
  • Document request-triggered recovery, cooldown reset after restart, and credential cleanup failures.

Instance restarts explain how the previous code returned to stopped, but do not establish the production trigger. The observed 30-minute graph change cannot distinguish metric sampling or inactivity from eviction, WebSocket closure, or deployment/restart. Cloudflare samples memory at invocations; its documented lifecycle does not establish a 30-minute inactivity timer.

Periodic authenticated /status requests could trigger recovery without manual starts, subject to Free plan limits. This branch adds no scheduler and does not guarantee uninterrupted operation. The actual production cause remains unverified.

Cloudflare lifecycle, dashboard evidence, and Free-tier limits

Why an instance restart resets the old bot, and what the 30-minute graph means

Before this change, Bot.socket and Bot.status lived only in instance memory and the constructor initialized status as stopped. A new Durable Object instance therefore had no socket and reported stopped; that is a code-path explanation, not evidence of which event restarted the production object. Cloudflare documents that Durable Objects may shut down after lifecycle inactivity/eviction, deployments, runtime updates, or placement decisions, and that shutdown hooks are not provided. Its lifecycle documentation says an outbound WebSocket can delay eviction for at most 15 minutes; after that it no longer prevents eviction, and the normal 70–140 second inactivity window applies when there are no further requests/events. A Durable Object shutdown terminates WebSocket connections. Those documented mechanisms make an in-memory socket loss possible, but do not prove the production trigger or a fixed 30-minute timer.

Cloudflare's Durable Objects memory chart is periodic V8-isolate memory sampling while objects are active. Even when filtered to one object, each sample is for the whole hosting isolate, which can include other objects and Worker code. The Worker memory chart instead reports invocation-time isolate-memory samples and reservoir-sampled percentiles. Neither chart is a WhatsApp connection-health, per-instance restart, or eviction-cause log. The observed change around 30 minutes cannot by itself distinguish a memory-sample/inactivity pattern from isolate/object eviction, WhatsApp WebSocket closure, or a deployment/runtime restart. Production cause remains unverified; no authenticated production queries were made.

What Free can support—and why this is not a 24/7 guarantee

The current Durable Objects pricing and Free limits list, per day, 100,000 requests and 13,000 GB-seconds, plus 5 million SQLite rows read and 100,000 rows written; daily limits reset at 00:00 UTC, and operations over a Free limit fail. Free Durable Objects require SQLite storage. Because Cloudflare bills a running DO using its allocated 128 MB, an idealized continuously billable single DO would consume about 0.125 GB × 86,400 seconds = 10,800 GB-s/day before other activity—nominally below 13,000 GB-s/day, but not a safe headroom or uptime promise. Other work, account usage, SQLite row operations, restarts, and connection handling still count.

One possible unattended-wake design is a recurring Durable Object alarm. Alarms can wake an object and are guaranteed at least once, but each DO has only one scheduled alarm; thrown handlers receive up to six automatic retries. Alarm invocations count toward DO requests, and pricing counts each setAlarm() as a row write. For scale, a 10-minute recurrence would be about 144 alarm invocations and 144 schedule writes per day for this object, before other work. This is plausible within the nominal Free counters in isolation, but the alarms themselves consume quota, Free-limit excess fails, the docs do not promise exact-time execution or uninterrupted WhatsApp availability, and a wake/reconnect necessarily leaves a gap after a lost socket. Paid DO usage currently includes 400,000 GB-s and 1 million requests per month; overage is $12.50 per million GB-s and $0.15 per million requests, with a $5/month minimum. Paying changes the billing envelope, not the Cloudflare/WhatsApp uptime guarantees.

This PR deliberately implements only bounded request-driven recovery: persisted credentials reconnect when a new instance receives a request, and /status retries a closed socket no more than once every 30 seconds. It does not add a recurring alarm. A guarantee of 24/7 continuity on the Free plan cannot be proven from these features, and this PR makes no such promise. With no incoming request, this implementation makes no new retry attempt; even a scheduled alarm would improve wake-up opportunities, not prove an uninterrupted external WhatsApp connection.

Risk Assessment

✅ Low: The change adds bounded, request-driven recovery with shared startup serialization and preserves authenticated access and explicit pairing.

Testing

Five focused mocked tests passed. The real local Worker passed authentication, unpaired-state, and QR-start checks, recorded as HTTP evidence. No production access or deployment occurred. Paired-session recovery could not be exercised.

  • Live validation: ⚠️ inconclusive - 3 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Missing or incorrect admin credentials cannot access bot status ✅ pass live Local Worker HTTP responses
Repeated status requests leave an unpaired bot stopped ✅ pass live Local Worker HTTP responses
Explicit start opens a WhatsApp connection and produces a pairing QR ✅ pass live Local Worker HTTP responses
First request after an instance restart restores the registered session without another start ⏸️ untested no No paired WhatsApp test account was available. Complete local pairing with a dedicated account, restart the runtime with persisted storage, and request status.
Status requests recover a disconnected session while cooldown and concurrent requests bound retries ⏸️ untested no Requires a paired test account and a controlled interruption of its live connection. Provide that account and complete local pairing.
Repeated start requests remain bounded when socket creation fails ⏸️ untested no The live runtime did not encounter socket creation failure. A controlled transport failure setup is needed to verify this boundary without mocking.
Revoking the linked session clears credentials and permits explicit pairing again ⏸️ untested no Requires a paired test account whose linked device can be revoked. Provide the account and complete local pairing.
Evidence: Local Worker HTTP responses, pairing credential redacted
[
  {
    "method": "GET",
    "path": "/status",
    "auth": "invalid or absent",
    "status": 401,
    "body": "Unauthorized"
  },
  {
    "method": "GET",
    "path": "/status",
    "auth": "invalid or absent",
    "status": 401,
    "body": "Unauthorized"
  },
  {
    "method": "GET",
    "path": "/status",
    "auth": "valid",
    "status": 200,
    "body": "{\"state\":\"stopped\"}"
  },
  {
    "method": "GET",
    "path": "/status",
    "auth": "valid",
    "status": 200,
    "body": "{\"state\":\"stopped\"}"
  },
  {
    "method": "POST",
    "path": "/start",
    "auth": "valid",
    "status": 200,
    "body": "{\"state\":\"connecting\"}"
  },
  {
    "method": "GET",
    "path": "/status",
    "auth": "valid",
    "status": 200,
    "body": {
      "state": "waiting_for_qr",
      "qr": "[redacted pairing credential]"
    }
  }
]
- Outcome: ⚠️ 2 warnings across 1 run (2m7s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

⚠️ **Test** - 2 warnings
  • ⚠️ Live evidence does not establish authenticated reconnection or logout recovery. Provide a dedicated WhatsApp test account and complete local QR pairing to exercise persisted-session restart, disconnect cooldown, concurrent retries, and credential revocation without production access.
  • ⚠️ live validation verdict: inconclusive (3 of 7 scenarios were driven live against the product); untested: First request after an instance restart restores the registered session without another start, Status requests recover a disconnected session while cooldown and concurrent requests bound retries, Repeated start requests remain bounded when socket creation fails, Revoking the linked session clears credentials and permits explicit pairing again
  • Live validation: ⚠️ inconclusive - 3 of 7 scenarios driven live against the product
Scenario Result Live Evidence
Missing or incorrect admin credentials cannot access bot status ✅ pass live Local Worker HTTP responses
Repeated status requests leave an unpaired bot stopped ✅ pass live Local Worker HTTP responses
Explicit start opens a WhatsApp connection and produces a pairing QR ✅ pass live Local Worker HTTP responses
First request after an instance restart restores the registered session without another start ⏸️ untested no No paired WhatsApp test account was available. Complete local pairing with a dedicated account, restart the runtime with persisted storage, and request status.
Status requests recover a disconnected session while cooldown and concurrent requests bound retries ⏸️ untested no Requires a paired test account and a controlled interruption of its live connection. Provide that account and complete local pairing.
Repeated start requests remain bounded when socket creation fails ⏸️ untested no The live runtime did not encounter socket creation failure. A controlled transport failure setup is needed to verify this boundary without mocking.
Revoking the linked session clears credentials and permits explicit pairing again ⏸️ untested no Requires a paired test account whose linked device can be revoked. Provide the account and complete local pairing.
  • npm ci --no-audit --no-fund
  • npm test -- test/index.test.ts
  • node node_modules/wrangler/bin/wrangler.js dev --local --ip 127.0.0.1 --port 8787 --var ADMIN_TOKEN:local-test-token --persist-to .wrangler/test-state --log-level error
  • Python HTTP requests exercised unauthorized status, repeated authorized status, explicit start, and subsequent QR status.
  • Removed generated dependencies, runtime state, and WASM copy; verified a clean worktree.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.


Summary by cubic

Reconnects stored WhatsApp sessions after a Durable Object restart so the bot no longer requires a manual start to resume.

Instances previously started in the stopped state and kept the socket only in memory, so production restarts made the bot idle until someone called POST /start. Now the first request on a new instance restores registered credentials from storage, and /status//start retry closed connections after a 30-second in-memory cooldown while sharing concurrent attempts.

  • Unpaired sessions still require /start for QR pairing.
  • Reconnect is request-triggered only, with no scheduler, so unattended uninterrupted operation is not guaranteed.
  • The production restart cause remains unverified: the observed 30-minute memory graph change cannot be attributed to a specific Cloudflare trigger.

Written for commit c101271. Summary will update on new commits.

Review in cubic

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-23T13:42:33.099793Z c101271 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c10127155c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/index.ts
if (this.socket) return
private start(allowPairing: boolean): Promise<void> {
if (this.socket) return Promise.resolve()
if (this.starting) return this.starting

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve pairing intent when sharing a start

If a pre-registration socket closes and, after the cooldown, a /status retry overlaps with POST /start while authentication is pending, the status request installs a non-pairing starting promise and this branch makes the explicit start share it. Because the stored credentials are still unregistered, openSocket(false) exits with stopped, so the POST succeeds without creating a socket or QR; preserve or upgrade allowPairing when a pairing request joins an in-progress automatic attempt.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant