A desktop voice assistant with a physical OLED face built using Python and Arduino UNO .
The assistant listens for voice commands, transcribes speech locally using Whisper, performs system automation, answers questions using a local LLM, and displays real-time expressions on an OLED display.
Created by Divyanshu Verma , a Class 9 student .
- Wake word activation
- Offline speech recognition using Faster Whisper
- Voice Activity Detection using Silero VAD
- Text-to-Speech responses
- Physical animated OLED face
- System automation
- Local AI conversation using Ollama
- Serial communication between Python and Arduino
- Command execution
- Continuous conversation mode
- Arduino UNO
- 0.96" SSD1306 OLED Display (I2C)
- USB Cable
- Computer running Python
- faster-whisper
- silero-vad
- sounddevice
- numpy
- pyttsx3
- requests
- pyserial
- Adafruit GFX Library
- Adafruit SSD1306 Library
- Wire Library
The OLED display changes its facial expression depending on Jarvis' current state.
| State | Expression |
|---|---|
| Idle | Calm face |
| Listening | Waiting for speech |
| Thinking | Animated processing face |
| Speaking | Talking animation |
Examples include:
- Open Notepad
- Open Browser
- Open GitHub
- Open YouTube
- Open VS Code
- Open Linux (WSL)
- Shutdown Computer
- Cancel Shutdown
Any command that is not recognized is automatically sent to the local AI model.
Jarvis/
│
├── jarvis1.1.py
├── arduino/
│ └── jarvis_face.ino
├── requirements.txt
├── LICENSE
├── README.md
└── .gitignore
Clone the repository
git clone https://github.com/Python-devloper-student/Jarvis-Physical-V1.gitInstall dependencies
pip install -r requirements.txtUpload the Arduino sketch to the Arduino UNO.
Start Ollama and make sure your custom model is available.
Update the COM port if necessary.
Run
python jarvis.py- Jarvis waits for the wake word.
- Voice is detected using Silero VAD.
- Audio is transcribed locally using Faster Whisper.
- If the command matches a predefined action, it is executed.
- Otherwise, the request is sent to a local LLM through Ollama.
- Responses are spoken aloud while the OLED face animates according to Jarvis' state.
- ESP32 wireless version
- Battery powered hardware
- Camera integration
- Face recognition
- Home automation
- Mobile control
- GUI dashboard
- Memory system
- Vision capabilities
This project is licensed under the MIT License.
You are free to use, modify, and distribute this project, but proper credit to the original author is appreciated.
Divyanshu Verma a class 9 student from india
GitHub: https://github.com/Python-devloper-student
Created as a personal learning project to explore embedded systems, AI, computer vision, and desktop automation.