Skip to content

Kid Mode: blocked-words content-output filter (planned) #138

Description

@BrettKinny

Status: planned / future work. Acknowledged in the README as not-yet-shipped, not a commitment to a near-term deadline.

What's missing

Kid Mode today applies prompt-level steering only — the persona prompt plus the per-turn sandwich (build_turn_suffix(kid_mode) → _wrap_with_sandwich, custom-providers/pi_voice/pi_voice.py). That keeps responses age-appropriate and on-topic by instruction, but there is no post-generation content filter on any live path. A blocked-words output filter (content_filter() / _BLOCKED_WORDS_RE) existed only in the retired ZeroClaw bridge.py /api/message handler and exists in no live code today.

docs/faq.md already states this honestly ("If the LLM says something inappropriate, the stack passes it through").

Scope when picked up

  • Add a post-generation blocked-words filter to a module both live LLM providers (PiVoiceLLM, OpenAICompat) can import.
  • Wire it into the TTS-bound text on the live voice path, gated on kid_mode.
  • Verify on-device (red-team session) — must not ship unverified.

Out of scope here (split from the old red-team issue)

  • C3 — _voice_memory_search_blocking (memory_lookup FTS) has no namespace filter; person_pending:<id> review-queue rows can leak into recall. (Privacy, not content.)
  • H4 — /xiaozhi/admin/* routes are unauthenticated. (LAN auth, not content.)

These should get their own issues if still valid.

Supersedes the content-filter portion of #22 (red-team pass), now closed.

Activity

  1. added
    safetySafety / correctness / child-safety bug
    on Jun 3, 2026
  2. BrettKinny commented on Jun 6, 2026

    @BrettKinny
    OwnerAuthor

    Progress note — partial coverage from the 2026-06 audit pass; core scope still open

    Two of the items split out of this issue have landed, and one adjacent sub-gap is now closed, but the headline deliverable — a post-generation blocked-words filter on the live voice path — remains unimplemented. Recording where things stand so this doesn't read as done:

    ✅ C3 (split out → #53) — closed. PR #153 added AND m.namespace NOT LIKE 'person_pending:%' to the FTS query in brain_db.ts, so unreviewed person_pending:<id> facts no longer leak into live memory_lookup recall. (Needs a dotty-pi image rebuild to deploy.)

    ✅ Dashboard say/start-story ingress — closed (PR #146). The operator-facing /actions/say and /actions/start-story handlers previously sanitised control chars + length but never called content_filter(), so an operator (or any LAN client when dashboard auth is unset) could make Dotty speak arbitrary unfiltered text with kid-mode on. bridge/dashboard.py:_kid_blocked() now routes both through the shared filter, fail-safe ON when kid-mode is unconfigured. This is a real sub-gap, but it is not this issue's scope — it guards an admin ingress, not LLM output.

    🔶 H4 (/xiaozhi/admin/* unauthenticated) — foundation merged, not yet enforced. The admin-auth epic (#149 server middleware, #150 behaviour, #151 pi-ext, #152 bridge) is complete across all four services but permissive — everything is a no-op until the deliberate enforcement flip (set DOTTY_ADMIN_TOKEN in all four container envs, deploy callers first, xiaozhi last). Tracked outside the issue list.

    ❌ Still open — this issue's actual scope: a post-generation blocked-words filter wired into the live voice path's TTS-bound text, gated on kid_mode, importable by both PiVoiceLLM and OpenAICompat.

    Verified in code today: content_filter() lives in bridge/text.py and is bridge-only — neither custom-providers/pi_voice/pi_voice.py nor custom-providers/openai_compat/openai_compat.py imports it. Both live providers still apply prompt-level steering only (build_turn_suffix(kid_mode)), exactly as this issue's original body describes. The audit pass did not touch this path.

    Remaining work unchanged:

    • Lift content_filter() (or its _BLOCKED_WORDS_RE core) into a module both live LLM providers can import.
    • Wire it into TTS-bound text on the live path, gated on kid_mode.
    • Red-team verify on-device — must not ship unverified.

    🤖 Drafted with Claude Code (Opus 4.8); reviewed by a human before posting.

  3. BrettKinny commented on Jun 6, 2026

    @BrettKinny
    OwnerAuthor

    Scoped the remaining voice-path filter work into #157 (implementation sub-task). This issue stays as the umbrella; #157 carries the concrete plan + acceptance criteria.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    safetySafety / correctness / child-safety bug

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions