A microservices WhatsApp AI assistant built to solve a personal problem - remembering important things and storing files that may be needed in the future.
Mindraft is designed to handle four core functions, all accessible through WhatsApp:
- Task Management - Add and track tasks
- Reminders - Set reminders so nothing gets forgotten
- Expense Tracking - Log and monitor expenses
- Document Storage - Store and retrieve documents using RAG with a vector database
Each function is built as its own MCP (Model Context Protocol) server, keeping concerns separated and independently scalable.
The system is split into two processes connected by a Redis Stream:
WhatsApp ──webhook──▸ FastAPI (app.py) --XADD──▸ Redis Stream
│
XREAD
│
Worker (worker.py)
│
┌───────┴───────┐
│ Groq / LLM │
└───────┬───────┘
│
WhatsApp reply
Why the decoupling? Meta's webhook expects a fast HTTP response. If the server called Groq directly, the LLM inference latency could cause the webhook request to time out. By pushing incoming messages into Redis and returning immediately, the server always responds within Meta's timeout window. A separate worker picks messages off the stream at its own pace and handles the LLM call + reply.
This also adds resilience - if the worker crashes, the LLM session expires, or there's a transient API error, messages are safely persisted in Redis and will be processed once the worker recovers.
| Layer | Technology | Rationale |
|---|---|---|
| API Framework | FastAPI | High request throughput, async-native |
| Message Queue | Redis Streams (Docker) | Decouples webhook response from LLM processing |
| LLM Provider | Groq | Fast inference for real-time conversations |
| LLM Model | Llama 3.1 8B Instant | Low latency with a large context window |
| Messaging | WhatsApp Business API (Meta) | Direct integration via webhooks |
| Tunneling | Cloudflare Tunnel | Free tier, exposes local server to webhooks |
-
WhatsApp Integration - Meta API and webhook setup is complete. The bot can send and receive messages. The access token is currently valid for 24 hours; this will be changed during deployment.
-
Redis Decoupling - Incoming messages are pushed to a Redis Stream (
XADD) by the FastAPI server. A background worker reads from the stream (XREAD), calls Groq, and sends the reply back via WhatsApp. Redis runs in Docker (required on Windows). -
LLM Integration - Groq is set up and working. The worker forwards messages to Groq and relays the response. Llama 3.1 8B Instant was chosen for its speed and large context window.
-
Local Development Tunnel - Cloudflare Tunnel on the free tier exposes the local server to WhatsApp webhooks.
pip install -r requirements.txtCreate a .env file in the project root:
VERIFY_TOKEN=<your-webhook-verify-token>
PHONE_NUMBER_ID=<your-whatsapp-phone-number-id>
ACCESS_TOKEN=<your-meta-access-token>
GROQ_API_KEY=<your-groq-api-key>
You need three processes running in separate terminals:
1. Start Redis via Docker(for Windows)
docker run -d --name redis -p 6379:6379 redis2. Start the FastAPI server
uvicorn app:app --reload --port 50003. Start the worker
python worker.py4. Start the Cloudflare tunnel (separate terminal)
.\cloudflared.exe tunnel --url http://localhost:5000- fastapi
- uvicorn
- httpx
- groq
- python-dotenv
- redis