An AI-powered robotic companion that supports aphasia recovery through structured, conversational speech therapy sessions.
- Overview
- Features
- Architecture
- Repository Layout
- Tech Stack
- Getting Started
- Environment Variables
- Data Flow
- Supabase Setup
Aphasia is a language disorder most commonly caused by stroke or brain injury, affecting a person's ability to speak, understand, read, and write. Waabi is a robotic therapy assistant designed to deliver consistent, engaging, and data-driven speech therapy exercises — bridging the gap between clinical appointments and at-home rehabilitation.
Waabi combines a LangGraph-based agentic therapy pipeline, a FastAPI communication bridge, a React web portal for therapists and patients, and a touchscreen-capable Kivy GUI that runs directly on the robot's hardware (e.g. a Raspberry Pi).
- Agentic therapy sessions — A LangGraph pipeline guides patients through structured speech exercises (perception → analysis → feedback → execution).
- Real-time voice interaction — Microphone capture, text-to-speech via Piper, and optional ElevenLabs Conversational AI integration.
- Session reporting — Completed sessions are automatically persisted to Supabase for therapist review.
- Dual interfaces — A Kivy touchscreen GUI for bedside/robot use and a React web app for therapist and patient portals.
- Hardware–cloud bridge — A FastAPI server mediates commands, status updates, UI events, and audio uploads between the robot and the cloud.
- Progress tracking — Recharts-powered dashboards in the web app let therapists monitor patient progress over time.
- Role-based access — Supabase Auth with protected routes for therapist and patient roles.
┌─────────────────────────────────────────────────────────┐
│ Web App (React) │
│ Therapist Portal │ Patient Portal │
└──────────────────┬──────────────────────────────────────┘
│ Supabase Auth (anon key)
▼
┌─────────────────────────────────────────────────────────┐
│ Supabase (PostgreSQL) │
│ profiles │ sessions │ session_reports │ robot-audio │
└───────────────┬─────────────────────────────────────────┘
│ service role key
▼
┌─────────────────────────────────────────────────────────┐
│ FastAPI Bridge (backend/main.py) │
│ GET /commands/{device_id} │ POST /status │
│ POST /audio/{device_id} │ POST /ui-events │
└──────────────────┬──────────────────────────────────────┘
│ HTTP polling
▼
┌─────────────────────────────────────────────────────────┐
│ Hardware Execution Agent (Raspberry Pi) │
│ hardware/execution_agent/main.py │ run_gui.py │
│ Kivy Touchscreen GUI + Drivers (speak, listen, UI) │
└──────────────────┬──────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────┐
│ LangGraph Therapy Agent (agentic/) │
│ Perception → Analysis → Feedback → Execution │
└─────────────────────────────────────────────────────────┘
Aphasia_Therapy_Assistant/
├── agentic/ # LangGraph speech-therapy pipeline + Supabase helpers
│ └── graph.py # Main session entrypoint (run_session_for_patient)
├── backend/
│ ├── main.py # FastAPI bridge: commands, status, audio, UI events
│ ├── requirements.txt
│ └── .env.example
├── frontend/ # React + Vite web app (therapist / patient portals)
├── hardware/
│ └── execution_agent/ # Robot / Raspberry Pi software
│ ├── main.py # Polling loop + hardware drivers
│ ├── run_gui.py # Kivy launcher (therapy GUI + LangGraph)
│ ├── env_bootstrap.py # .env merge order utility
│ ├── config.json
│ └── .env.example
├── ML/ # Speech analysis experiments and research
├── integrations/
│ └── elevenlabs_convai/ # ElevenLabs Conversational AI integration
├── pipecat/ # Optional real-time voice assistant stack
├── supabase/
│ └── migrations/ # SQL migrations (robot_ui_events, pipeline tables, etc.)
├── scripts/ # Helper shell scripts
├── tests/ # Test suite (pytest)
├── laptop_server.py # Convenience server launcher
├── requirements.txt # Root Python dependencies
├── requirements-dev.txt # Dev/test dependencies
└── netlify.toml # Frontend deployment config
| Layer | Technology |
|---|---|
| Frontend | React (Vite), TypeScript, Tailwind CSS, Recharts, React Router v6 |
| Backend / Bridge | Python, FastAPI, Uvicorn |
| Therapy Agent | LangGraph, Groq (LLM inference) |
| Database & Auth | Supabase (PostgreSQL), Supabase Storage |
| Robot GUI | Python, Kivy, kvlang |
| TTS | Piper (local), ElevenLabs (cloud) |
| CI/CD | GitHub Actions, Netlify |
| Testing | Pytest |
- Python 3.10+
- Node.js 18+ (for the frontend)
- A Supabase project with the required tables and storage bucket (see Supabase Setup)
- A Groq API key
- PortAudio installed on the robot device (
sudo apt-get install portaudio19-dev)
Run from hardware/execution_agent/:
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env # Fill in SERVER_URL, DEVICE_ID, Supabase & Groq keysNo microphone / headless testing? Set
RUN_GUI_SKIP_PORTAUDIO_CHECK=1in your.env.
python run_gui.pyThe GUI will prompt for the patient's full name (must match a profiles.full_name row with role = 'patient' in Supabase). The LangGraph pipeline then runs the session and writes a report to session_reports on completion.
.env files are merged in this order (later files override earlier ones):
repo root .env → backend/.env → agentic/db/.env → execution_agent/.env → hardware/.env
cd /path/to/Aphasia_Therapy_Assistant
python -m venv .venv
source .venv/bin/activate
pip install -r backend/requirements.txt
# Set SUPABASE_URL and SUPABASE_SERVICE_KEY in .env (repo root or backend/.env)
uvicorn backend.main:app --reload --app-dir .The --app-dir . flag ensures agentic package imports resolve correctly from the repo root.
cd frontend
npm install
cp .env.example .env.local # Set VITE_SUPABASE_URL and VITE_SUPABASE_ANON_KEY
npm run devOpen http://localhost:5173 in your browser.
| Variable | Description |
|---|---|
SUPABASE_URL |
Your Supabase project URL |
SUPABASE_SERVICE_KEY |
Supabase service role key (server-side only) |
GROQ_API_KEY |
Groq API key for LLM inference |
| Variable | Description |
|---|---|
SERVER_URL |
URL of the FastAPI bridge |
DEVICE_ID |
Stable identifier for the robot (e.g. robot-pi-001 or a patient UUID) |
| Piper paths | Path to local TTS model files |
| Variable | Description |
|---|---|
SUPABASE_URL |
Supabase project URL |
SUPABASE_SERVICE_KEY |
Service role key |
BRIDGE_PATIENT_PROFILE_ID |
(Optional) Patient profiles.id UUID. Use when DEVICE_ID is a label rather than a UUID; the bridge loads the patient's name for session greetings. |
Robot ID vs Patient ID:
DEVICE_IDidentifies the robot in API paths. For the bridge greeting, either setBRIDGE_PATIENT_PROFILE_IDon the server, or setDEVICE_IDto the patient'sprofiles.idUUID directly (legacy behaviour).
- Launch
run_gui.pyfromhardware/execution_agent/. - Patient enters their name; the GUI looks up the matching
profilesrow in Supabase. agentic/graph.pyrunsrun_session_for_patientthrough the full LangGraph pipeline.- On completion,
persist_session_statewrites results tosession_reports. - The FastAPI bridge is not required for session persistence in this path.
- Start the bridge:
uvicorn backend.main:app --reload --app-dir . - Start the polling agent:
hardware/execution_agent/main.py(withpollingenabled inconfig.json). - The robot calls
GET /commands/{DEVICE_ID}, executes actions, then posts toPOST /status. POST /statuswrites tosession_reportsonly whenack.resultincludespatient_idand the full expected state fields. For complete session persistence, prefer the Kivy/LangGraph path.
frontend/authenticates via Supabase Auth using the anon key (set in Vite env).- Therapists can view patient session reports, assign target words, and monitor progress.
- Patients can access their session history and exercises.
-
Apply migrations from
supabase/migrations/to your Supabase project (includesrobot_ui_events,agent_pipeline_steps, and others). -
Ensure these tables exist:
profiles— user accounts with arolecolumn ('therapist'or'patient') andfull_namesessions— therapy session definitions includingtarget_wordssession_reports— completed session results written by the agent
-
Create a storage bucket named
robot-audioif you use robot audio upload functionality. -
Configure RLS (Row Level Security) policies as appropriate for your deployment — therapists should read all patient records; patients should read only their own.