Repository navigation
Kid Mode: blocked-words content-output filter (planned) #138
Description
Activity
- addedsafetySafety / correctness / child-safety bugSafety / correctness / child-safety bug
on Jun 3, 2026 - added a commit that references this issue
on Jun 3, 2026 Progress note — partial coverage from the 2026-06 audit pass; core scope still open
Two of the items split out of this issue have landed, and one adjacent sub-gap is now closed, but the headline deliverable — a post-generation blocked-words filter on the live voice path — remains unimplemented. Recording where things stand so this doesn't read as done:
✅ C3 (split out → #53) — closed. PR #153 added
AND m.namespace NOT LIKE 'person_pending:%'to the FTS query inbrain_db.ts, so unreviewedperson_pending:<id>facts no longer leak into livememory_lookuprecall. (Needs adotty-piimage rebuild to deploy.)✅ Dashboard say/start-story ingress — closed (PR #146). The operator-facing
/actions/sayand/actions/start-storyhandlers previously sanitised control chars + length but never calledcontent_filter(), so an operator (or any LAN client when dashboard auth is unset) could make Dotty speak arbitrary unfiltered text with kid-mode on.bridge/dashboard.py:_kid_blocked()now routes both through the shared filter, fail-safe ON when kid-mode is unconfigured. This is a real sub-gap, but it is not this issue's scope — it guards an admin ingress, not LLM output.🔶 H4 (
/xiaozhi/admin/*unauthenticated) — foundation merged, not yet enforced. The admin-auth epic (#149 server middleware, #150 behaviour, #151 pi-ext, #152 bridge) is complete across all four services but permissive — everything is a no-op until the deliberate enforcement flip (setDOTTY_ADMIN_TOKENin all four container envs, deploy callers first, xiaozhi last). Tracked outside the issue list.❌ Still open — this issue's actual scope: a post-generation blocked-words filter wired into the live voice path's TTS-bound text, gated on
kid_mode, importable by bothPiVoiceLLMandOpenAICompat.Verified in code today:
content_filter()lives inbridge/text.pyand is bridge-only — neithercustom-providers/pi_voice/pi_voice.pynorcustom-providers/openai_compat/openai_compat.pyimports it. Both live providers still apply prompt-level steering only (build_turn_suffix(kid_mode)), exactly as this issue's original body describes. The audit pass did not touch this path.Remaining work unchanged:
- Lift
content_filter()(or its_BLOCKED_WORDS_REcore) into a module both live LLM providers can import. - Wire it into TTS-bound text on the live path, gated on
kid_mode. - Red-team verify on-device — must not ship unverified.
🤖 Drafted with Claude Code (Opus 4.8); reviewed by a human before posting.
- Lift
Status: planned / future work. Acknowledged in the README as not-yet-shipped, not a commitment to a near-term deadline.
What's missing
Kid Mode today applies prompt-level steering only — the persona prompt plus the per-turn sandwich (
build_turn_suffix(kid_mode)→_wrap_with_sandwich,custom-providers/pi_voice/pi_voice.py). That keeps responses age-appropriate and on-topic by instruction, but there is no post-generation content filter on any live path. A blocked-words output filter (content_filter()/_BLOCKED_WORDS_RE) existed only in the retired ZeroClawbridge.py/api/messagehandler and exists in no live code today.docs/faq.mdalready states this honestly ("If the LLM says something inappropriate, the stack passes it through").Scope when picked up
PiVoiceLLM,OpenAICompat) can import.kid_mode.Out of scope here (split from the old red-team issue)
_voice_memory_search_blocking(memory_lookupFTS) has no namespace filter;person_pending:<id>review-queue rows can leak into recall. (Privacy, not content.)/xiaozhi/admin/*routes are unauthenticated. (LAN auth, not content.)These should get their own issues if still valid.
Supersedes the content-filter portion of #22 (red-team pass), now closed.