Repository navigation
fix(segments): exclusão de contato invalida o cache de deletados em todos os processos (CRM-245) - #129
Conversation
…sos via Redis (CRM-245) O sinal de contato excluído saía só pelo EventEmitter2 do processo que recebe o /events/identify. No split api/worker, o segment-worker, que recalcula os segmentos, nunca o recebia e seguia com a lista velha até o TTL de 5 minutos. O DeletedContactsSignalRelay publica o sinal num canal Redis (evo-flow:db<REDIS_DB>:segments:contact-deleted; pub/sub ignora o índice do banco) e reemite como evento local, marcado como relayed, o que chega dos outros processos. Cada processo ignora a própria origem. Redis fora do ar não segura o boot nem a ingestão: os caches voltam ao TTL. Co-Authored-By: Claude Code <[EMAIL_REDACTED]>
… velho (CRM-245) Na janela de 15 s depois de uma exclusão o cache não responde, e cada chamada ia ao ClickHouse: um recálculo de N segmentos disparava N buscas iguais. Chamadas concorrentes agora compartilham a busca em andamento. Cada invalidação incrementa uma geração e descarta a busca em voo, então um resultado que começou antes da exclusão volta para quem pediu mas não é gravado por 5 minutos. O expiresAt passa a ser calculado quando a busca termina. Co-Authored-By: Claude Code <[EMAIL_REDACTED]>
…de deletados (CRM-245) O spec do relay simula dois processos sobre um Redis em memória: a exclusão ingerida na api invalida o cache do worker, a origem não recebe o sinal duas vezes, o que chega de fora não é republicado, mensagem malformada ou de outro canal é ignorada e falha de publish não lança. No cache: chamadas concorrentes fazem uma só busca, dentro e fora da janela de bypass, e uma busca iniciada antes do sinal não grava o conjunto velho. Co-Authored-By: Claude Code <[EMAIL_REDACTED]>
… fim (CRM-245) Achado da review: com o expiresAt calculado ao fim da busca, uma busca que começava dentro da janela de 15 s e respondia depois dela gravava por 5 minutos um conjunto lido dentro da janela, possivelmente sem a exclusão. É a corrida que a janela existe para evitar, e a develop já decidia pelo início. Co-Authored-By: Claude Code <[EMAIL_REDACTED]>
…dis (CRM-245) Com o Redis fora, o quit() rejeitava e os clientes seguiam reconectando depois do shutdown; disconnect() encerra na hora. O erro de conexão era logado a cada tentativa de reconexão (40 linhas em 10 s): agora sai uma linha por queda e outra quando a conexão volta. Co-Authored-By: Claude Code <[EMAIL_REDACTED]>
…y (CRM-245) Cobre os achados da review: a busca que começa dentro da janela e responde depois dela não é servida do cache; um TestingModule com o EventEmitterModule real prova que o @onevent publica a exclusão local uma vez e não republica a reemitida; shutdown usa disconnect; erro de conexão sai uma vez por queda. Co-Authored-By: Claude Code <[EMAIL_REDACTED]>
Reviewer's GuideFixes stale deleted-contact counts across separate API and segment-worker processes by relaying invalidation signals over Redis pub/sub, while hardening cache refresh behavior against concurrent requests and races with invalidation. The relay is deliberately best effort—Redis outages preserve the existing TTL fallback—and the new tests cover cross-process delivery, loop prevention, malformed input, lifecycle behavior, and cache concurrency/generation semantics. Flow diagram for generation-safe deleted-contacts cache refreshflowchart TD
Request["getDeletedContacts()"] --> Hit{"Valid cached set?"}
Hit -- Yes --> ReturnCache["Return cached set"]
Hit -- No --> InFlight{"inFlight fetch exists?"}
InFlight -- Yes --> Share["Share existing Promise"]
InFlight -- No --> Fetch["fetchDeletedContactsFromClickHouse()"]
Invalidate["invalidateCache()"] --> Generation["Increment generation"]
Generation --> Clear["Clear cached set and expiry"]
Fetch --> Compare{"Fetch generation matches?"}
Compare -- Yes --> Store["Store result with fetch-start expiry"]
Compare -- No --> ReturnFresh["Return result without caching"]
Store --> ReturnFresh
Share --> ReturnFresh
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
There was a problem hiding this comment.
Hey - I've found 3 issues
Prompt for AI Agents
Please address the comments from this code review:
## Individual Comments
### Comment 1
<location path="src/modules/segments/services/deleted-contacts-cache.service.ts" line_range="71" />
<code_context>
+ : startedAt + this.CACHE_TTL;
+ this.logger.debug(`Cached ${deletedContacts.size} deleted contacts`);
+ }
return deletedContacts;
} catch (error) {
this.logger.error('Failed to fetch deleted contacts:', error);
</code_context>
<issue_to_address>
**Deleted contacts remain in segments**
When a deletion signal invalidates the cache while a segment computation is awaiting a ClickHouse fetch, `loadDeletedContacts` returns the fetched set after its generation becomes stale, so the segment computation uses the pre-deletion set and can retain the deleted contact for that run.
Check the generation before returning the set and refetch or discard it when it has become stale.
Also at `src/modules/segments/services/deleted-contacts-cache.service.ts:72`.
</issue_to_address>
### Comment 2
<location path="src/modules/segments/services/deleted-contacts-signal.relay.ts" line_range="69" />
<code_context>
+ });
+
+ for (const client of [this.publisher, subscriber]) {
+ client.connect().catch(() => undefined);
+ }
+ }
</code_context>
<issue_to_address>
**Deletion signals are dropped**
When a deletion event arrives while the publisher is connecting or reconnecting, `onContactDeletedIngested` publishes only once, and `enableOfflineQueue: false` makes `publish` reject before the publisher is ready; the relay logs and drops the signal, leaving other processes with stale deleted-contact data until their cache TTL expires.
Wait for publisher readiness or retry/queue failed publishes so each deletion signal is delivered.
Also at `src/modules/segments/services/deleted-contacts-signal.relay.ts:131`.
</issue_to_address>
### Comment 3
<location path="src/modules/segments/services/deleted-contacts-cache.service.ts" line_range="40-41" />
<code_context>
return this.cached;
}
+ if (this.inFlight) {
+ return this.inFlight;
+ }
+
</code_context>
<issue_to_address>
**Bypass fetch is reused**
When a fetch started during the 15-second bypass window remains pending after the window closes and another caller requests deleted contacts, `getDeletedContacts` returns the existing `inFlight` promise without checking when that fetch started, so the caller receives a result that may not include the deletion and can treat the deleted contact as active.
Check the in-flight fetch's start time against the bypass window and start a fresh fetch for callers arriving after the window closes.
</issue_to_address>Sourcery assessment
Needs a human reviewer. 3 findings to address first, and if the relay or generation handling is wrong, deleted-contact results can remain stale or cause extra ClickHouse queries across processes, but the impact is bounded to in-memory cache state and can be repaired by clearing or allowing the cache to expire. Reverting removes the new behavior, although cache state already created before the revert does not disappear immediately.
Blocking findings: src/modules/segments/services/deleted-contacts-cache.service.ts:71, src/modules/segments/services/deleted-contacts-signal.relay.ts:69, src/modules/segments/services/deleted-contacts-cache.service.ts:41
… fetch The bypass-window comment still said every call queries ClickHouse; concurrent calls now share one fetch. Also drops a history note from the relay header.
Summary
EventEmitter2do processo que recebe o/events/identify. Com API e workers em processos separados (RUN_MODE=apiesegment-worker), o worker que recalcula os segmentos nunca o recebia: o contato excluído seguia contando por até 5 minutos, até o TTL do cache vencer.DeletedContactsSignalRelay, com uma instância por processo:evo-flow:db<REDIS_DB>:segments:contact-deleted(pub/sub ignora o índice do banco);EventEmitter2local, marcado comorelayed, o que chega de outros processos, e ignora a própria origem;disconnect()no shutdown.DeletedContactsCacheService:Security
contactIde um id de processo. O canal apenas invalida cache, não transporta dado.Test plan
npx jest --ci --maxWorkers=2→ 1150 passed, 27 skipped.TestingModulee oEventEmitterModulereal: o@OnEventpublica a exclusão local uma vez e não republica a reemitida.RUN_MODE=segment-workerextra:PUBSUB NUMSUB= 2);contact.deletedemPOST /api/v1/events/identifypublicou{origin, contactId}no canal.Changed Files
src/modules/segments/services/deleted-contacts-signal.relay.ts(novo)src/modules/segments/services/deleted-contacts-cache.service.tssrc/modules/segments/segments-cache.module.tssrc/modules/segments/services/deleted-contacts-signal.relay.spec.ts(novo)src/modules/segments/services/segment-canonical-event-names.spec.tsLinked Issue
🤖 Generated with Claude Code
Summary by Sourcery
Synchronize deleted-contact cache invalidation across processes and harden cache refresh behavior against concurrent and in-flight queries.
New Features:
Bug Fixes:
Enhancements:
Tests: